Skip to content
AI.info

Research

AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

Overview Research area: Natural language processing, LLM agents, and data-centric machine learning. Technical level: Intermediate. Scope: The paper introduces a controlled benchmark for testing whethe

AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
arXiv
2609.40097
Published
2026-09-30
Authors
Ruifeng Yuan, Yizhi Li, Yaxin Du, Fengyu Cai, Yiqi Liu, Hou Pong Chan, Chenghua Lin, Yun Chen, Jian Yang, Bryan Dai, Pinyan Lu, Chenghao Xiao

AI summary

Overview

  • Research area: Natural language processing, LLM agents, and data-centric machine learning.
  • Technical level: Intermediate.
  • Scope: The paper introduces a controlled benchmark for testing whether LLMs can improve training data, predict the effects of their changes, and use feedback to refine their decisions.

What This Paper Is About

Many AI research benchmarks let agents change several things at once, so it can be hard to tell whether progress came from better data, different training settings, or more resources. AutoDataBench focuses on data: it tests whether LLMs can find problems in training examples, organize useful training material, or create new examples that improve a model. It also examines whether their data strategies work beyond the evaluation tasks they receive feedback on.

Key Contributions

  1. A data-focused benchmark: AutoDataBench evaluates three aspects of data intelligence—data diagnosis and repair, data organization, and data construction—through tool-use, retrieval, and knowledge-injection tasks.
  2. A controlled experimental setup: Agents work within fixed training procedures and task-specific resource limits, while hidden evaluation sets test whether improvements generalize beyond the feedback tasks.
  3. A behavioral probe of data-effect reasoning: The researchers record agents’ predictions about model performance before training, then compare those predictions with results.
  4. A test of trajectories as training data: The paper reuses agents’ benchmark interactions during mid-training and finds improvements on downstream coding evaluations.

Main Findings

  • Tool-use data repair can approach a clean-data reference. The LLMs’ mean in-distribution scores range from 80.82 to 81.98, compared with 70.45 for the random baseline and 81.65 for the expert reference. Their scores on BFCL are above the baseline but below the expert reference.
  • Strong retrieval performance does not ensure transfer. GPT-5.6-Sol has the highest mean in-distribution retrieval score, 40.46, close to the expert reference of 40.62. However, all LLMs score below the expert on the held-out retrieval tasks. Qwen-3.7-Max has the highest mean out-of-distribution score among the LLMs, at 30.18.
  • LLM-created data helps inject new knowledge. Every evaluated LLM matches or exceeds the expert’s Novel accuracy of 48.40. Kimi-3 achieves the highest mean Novel accuracy, 60.60. Knowledge acquisition and retention do not necessarily move together: Claude-4.7 leads on Retention, while GLM-5.2’s Retention score is below the unadapted model’s.
  • Most standard runs improve through experimentation. Of 63 standard runs, 61 improve over their initial datasets. The average gains are 8.11 points for tool use, 6.46 for retrieval, and 4.91 for knowledge injection.
  • Forecasting performance can track results, but may add a burden. On retrieval, GPT-5.6-Sol and Kimi-3’s forecasts have mean absolute errors of 0.64 and 1.55 points. In four of the six prediction-enabled model–task comparisons, the best score is below the corresponding standard-run mean.
  • Benchmark trajectories improve coding results in a separate training setup. Adding auto-research trajectories improves all five reported coding scores. The largest gains are on CRUXEval input prediction, from 72.00 to 74.12, and SWE-bench Multilingual, from 27.67 to 33.33.

Methodology in Plain English

The researchers built three tasks around different ways training data can matter:

  • Tool use: Agents work with a function-calling dataset that has deliberately introduced errors. They can filter, repair, deduplicate, or otherwise revise examples, then train a tool-use model on the result.
  • Retrieval: Agents organize training data for an embedding model. They are given query–positive pairs and can mine hard negatives—similar-looking passages that should not be treated as relevant—while choosing how to mix data sources.
  • Knowledge injection: Agents turn information newer than the target model’s knowledge cutoff into context, question, and answer examples. A fixed distillation process then trains the model to answer without seeing the supporting context.

Across tasks, the training pipeline and other non-data components are held fixed. Agents inspect data, implement changes, train a model, receive task-specific evaluation feedback, and can revise their approach. Hidden evaluation data are not available during this optimization, so the researchers can check whether a strategy transfers beyond the target tasks.

The study evaluates seven frontier LLMs, with three standard runs for each LLM–task pairing. It also runs a separate forecasting study with GPT-5.6-Sol and Kimi-3: after seeing an evaluation result, each model predicts the score for its next data intervention before that next training evaluation is run.

For the data-engine experiment, the researchers add benchmark trajectories to a coding model’s mid-training data and compare it with a matched control. Both models then receive the same supervised fine-tuning.

Why This Matters

For research, AutoDataBench provides a way to study data-improvement ability separately from changes to models, training recipes, or compute. Its results also show why measuring only performance on the feedback tasks can miss weak generalization.

Potential real-world applications include:

  • Function-calling systems: finding inconsistent labels or examples that teach the wrong relationship between a user request and a tool call.
  • Search and retrieval systems: selecting informative training examples and negatives to improve passage retrieval.
  • Knowledge updates: turning new documents into grounded training examples for models that need to learn new facts.
  • AI-assisted software development: using research trajectories as training material for agents that investigate problems, test hypotheses, and revise their actions.

For industry, the work addresses a practical bottleneck: data choices can strongly affect model quality, but testing those choices through training experiments can be costly. A benchmark for data intelligence could help identify which agents are useful for curation workflows. The paper’s coding results also suggest that the records of data-focused experimentation may themselves be reusable training material.

Future Directions

  • Test whether the benchmark’s findings hold across more tasks, data sources, model sizes, and training paradigms.
  • Investigate how to improve out-of-distribution transfer, especially when an agent receives strong feedback on a limited set of target tasks.
  • Study whether asking agents to forecast outcomes consistently helps them learn from feedback or instead distracts from data optimization.
  • Separate the effects of individual data changes from the combined effects of experimentation, checkpoint selection, and training variability.

Target Audience

This paper is most useful for researchers and practitioners working on LLM agents, data curation, retrieval and tool-use systems, model training, and AI research automation. It is also relevant to teams that want to evaluate whether an agent can make data changes that improve a trained model—not merely produce plausible data-processing code.

Authors’ abstract

Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.

Read the original paper