Skip to content
AI.info

Research

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

Overview Research area: Large language model agents, specifically self-improving agents and test-time scaling for long-horizon tasks. Technical level: Advanced. Scope: The paper introduces AREX-2, a 2

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
arXiv
2609.38288
Published
2026-09-29
Authors
Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li, Qiwei Ye, Zhicheng Dou, Defu Lian, Zheng Liu

AI summary

Overview

Research area: Large language model agents, specifically self-improving agents and test-time scaling for long-horizon tasks.

Technical level: Advanced.

Scope: The paper introduces AREX-2, a 27B-parameter agent trained on synthetic long-horizon improvement trajectories from machine learning engineering and algorithmic programming, and evaluates whether the resulting "long-horizon reflection" capability transfers to deep research and general agentic reasoning benchmarks.

What This Paper Is About

Most agent systems implement the try-measure-revise loop in the scaffolding around the model, not inside the model itself: the harness decides when to retry and what to keep, while the model produces one attempt at a time. The authors define self-improvement as the ability of an agent, given more rounds on a task, to convert that extra budget into a better solution through its own judgment. The goal is to build training data that teaches this loop and to test whether the capability is domain-agnostic.

Key Contributions

  1. A formalization of self-improvement as test-time scaling. The paper decomposes total improvement over a round budget into two quantities: reflection, which sets the mean gain per productive round, and long-horizon execution, which sets the number of rounds that remain productive. The authors hypothesize that both are domain-agnostic meta-skills.

  2. A three-step data-construction pipeline. A teacher model converts GitHub repositories and online-judge problems into executable environments consisting of a task and a scoring function; a teacher agent then runs many rounds in each environment; trajectories are selected as a whole by final score and process validity, with no condition on individual rounds.

  3. Whole-trajectory supervision with step-level loss masking. Failed runs, regressions, and abandoned approaches are deliberately retained in the context, and the training loss is applied only to the decisions that follow them and make progress, rather than to the failures themselves. Steps such as repeated polling and near-duplicate turns receive no loss.

  4. AREX-2, trained from Qwen3.8-27B. The model is fine-tuned on the new trajectories plus the unchanged deep-research data of the previous AREX recipe, making any deep-research gain attributable to transfer rather than new research data.

Main Findings

  • Leading score on MLE-bench Lite. AREX-2 reaches 81.8 Any Medal (mean over three seeds), the highest in the comparison table: 8.1 points above Naive-N0.5-Flash (73.7) and 9.1 points above the strongest closed-weight baseline, GPT-5.6 Sol (72.7).

  • Competitive algorithmic programming at 27B parameters. On Frontier-CS (188-task Agent Track), AREX-2 scores 70.7, exceeding the strongest reported open-weight baseline by 16.0 points and placing within 5.7 points of the strongest closed-weight system.

  • Transfer to deep research with no new research data. AREX-2 scores 84.0 on BrowseComp, 52.6 on text-only HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, surpassing both AREX (4B) and AREX (122B) on all four benchmarks. It leads reported models of at most 40B parameters on BrowseComp, text-only HLE, and DeepSearchQA; on GAIA it trails XYZ-Aquila-mini (97.1) and Agents-A1 (96.0).

  • A larger round budget keeps paying off with feedback. On Frontier-CS over a five-hour budget per problem, AREX-2 reaches 54.4 after one hour, 65.9 after two hours, and 70.7 after five hours, gaining 4.8 points between the second and fifth hour, of which 2.2 points come in the final hour. DeepSeek-V4-Pro stops improving at 44.7 after two hours and DeepSeek-V4-Flash at 39.1 after three hours.

  • Improvement without a correctness signal. On BrowseComp, where no score is given during a run, AREX-2 rises from 64.8 accuracy at 47 turns to 84.0 at 143 turns. At roughly 140 turns it is more than 12 points ahead of AREX (122B), exceeds that model's final accuracy in less than half the turns, and gains about three times as much per turn, with less than a quarter of its parameters.

  • Each component contributes measurably. In the five-stage MLE-bench Lite ablation, the base Qwen3.8-27B scores 28.8 with no skills, 41.2 with general skills, and 68.2 with task-specific skills and more rounds. AREX-2 reaches 75.8 with the base number of rounds and 81.8 with more rounds. Matching skills and rounds, training adds 13.6 points over the base model; with fewer rounds, AREX-2 still exceeds the fully equipped base model by 7.6 points; and the trained model gains 6.0 points from more rounds.

  • A remaining gap. The authors state that AREX-2 still trails the strongest reported BrowseComp and HLE results.

Methodology in Plain English

The researchers first built training environments rather than training examples. A teacher model read GitHub repositories and online-judge problems and turned each one into a sandbox with a task and an automatic scoring function, such as a validation metric on a held-out split or a graded fraction of hidden tests passed. An environment was accepted only if its reference solution scored and a naive baseline scored well below it, guaranteeing room for improvement.

A teacher agent then worked on each accepted task under three deliberate conditions: a round budget far larger than a single attempt requires (hours of wall-clock time and hundreds of tool calls), encouragement to acquire operational knowledge on the job by reading source and documentation, searching for dependencies, and running small experiments, and feedback after every submission so the agent could compare results with expectations, keep or revert changes, and form a next hypothesis. Compact "skills" documents were also placed in the agent's context for this data-generation stage.

Trajectories were kept or discarded as wholes based on final score relative to the reference and on process validity, not on whether each round succeeded. During fine-tuning on Qwen3.8-27B, the full trajectory stayed in context while the loss was applied only to the decisions that made progress after failures. The new trajectories were the only change from the previous AREX recipe, whose deep-research data was left untouched, so deep-research results measure transfer. Evaluation used each benchmark's official agent-based protocol: a maximum of 300 inner turns and 1500 total turns for non-coding benchmarks, five hours per task on the Frontier-CS Agent Track, and the OpenMLE protocol with a twelve-hour budget per task on MLE-bench Lite, where AREX-2 was evaluated with skills in its context.

Why This Matters

Impact on research. The paper argues that self-improvement can be trained as a meta-skill rather than engineered into a scaffold, and it provides a data-construction recipe plus ablations separating skills, training, and round budget. The cross-domain results support the hypothesis that reflection and long-horizon execution transfer, offering a route to improve deep-research agents without collecting new deep-research trajectories.

Real-world applications (as suggested by the paper's own setting):

  • Automated machine learning engineering, where data analysis, experimentation, and model evaluation are iterated toward a validation metric.
  • Algorithmic programming assistance on open-ended problems with graded rather than binary scoring.
  • Deep research and information gathering, including tool-augmented reasoning and multi-step question answering.
  • Long-running agentic workflows that must recover from failed attempts and abandoned strategies rather than reporting a single answer.

Industry relevance. The result that a 27B model matches or exceeds much larger systems on several benchmarks bears directly on deployment cost, and the finding that accuracy keeps rising with turns on BrowseComp shows that inference-time compute can be traded for quality on non-coding tasks, including when no correctness signal is available.

Future Directions

  • Widening the training domains beyond machine learning engineering and algorithmic programming, since the paper's hypothesis predicts that other verifiable domains could also supervise long-horizon reflection.
  • Lengthening the horizons in training data, given that the frontier experiments show AREX-2 still improving at the end of its budget on Frontier-CS and approaching diminishing returns only gradually on BrowseComp.
  • Closing the loop by letting the model's own trajectories become its next round of training data, which the conclusion names directly.
  • Reducing the remaining gap to the strongest BrowseComp and HLE results, along with understanding when whole-trajectory selection with step-level loss masking is preferable to per-round filtering.

Target Audience

Researchers and engineers working on LLM agents, agentic reinforcement learning and imitation learning, test-time scaling, and training-data synthesis for tool-using models. It is also relevant to practitioners who need agents that sustain long execution loops under a fixed compute budget, and to readers interested in whether capabilities learned in one verifiable domain carry over to another. The formalized round-gain analysis makes the paper accessible to those approaching it from a machine learning evaluation background, though familiarity with agent benchmarks and fine-tuning pipelines helps.

Authors’ abstract

We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-27B, achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its budget of rounds grows. These results show that long-horizon reflective data is an effective route toward self-improving agents.

Read the original paper