Research
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Overview Research area: Natural Language Processing, specifically self-evolving search agents trained with reinforcement learning (cs.CL; arXiv:2609.39102v1, 30 Sep 2026). Technical level: Advanced. T

- arXiv
- 2609.39102
- Published
- 2026-09-30
- Authors
- Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu, Zijun Tian, Yufan Zou, Shuhan Sun, Hanxin Chen, Zeyu Zhang, Weizhi Du, Yueting Li, Tianyu Shi, Alaa Khamis
AI summary
Overview
- Research area: Natural Language Processing, specifically self-evolving search agents trained with reinforcement learning (cs.CL; arXiv:2609.39102v1, 30 Sep 2026).
- Technical level: Advanced. The paper assumes familiarity with proposer–solver self-evolution loops, policy-gradient training, pseudo-labels, and reward hacking.
- Scope: The paper diagnoses a failure mode it names co-cheating in self-evolving search agents, where a question proposer and an answer solver increasingly agree on the same wrong answers so that internal reward rises without external correctness rising, and it proposes CrossFit, a source-level cross-fitting fix to proposer feedback.
What This Paper Is About
Self-evolving search agents build their own training curricula: a proposer turns source documents into questions and pseudo-labels, a solver is trained on the admitted pairs, and the solver's performance on new proposals becomes the proposer's reward. The paper shows that repeating this closed loop lets proposer and solver converge on shared errors, so agreement becomes an unreliable proxy for correctness and the "reward" improves while real answer quality does not. The goal is to diagnose that failure with an independent audit and to intervene in the loop so that agreement stops being rewarded when it merely reflects a shared mistake.
Key Contributions
- Identifies and names co-cheating, an optimization outcome in which the proposer and solver self-reinforce the same incorrect labels, and defines a measurable quantity for it: false-agreement mass, the fraction of evaluated pairs that agree on the same incorrect answer.
- Builds a post-hoc, evidence-backed audit using
gpt-6-astra/highthat reconstructs a reference from the source document and judges the saved pseudo-label and the five solver responses used for proposer reward. The auditor never affects admission, model updates, or reward, so it measures the exact examples behind the in-loop signal. - Introduces two interventions: MSV (multi-sample verification), an admission-time test that queries the same model three times with the source and three times without it, and CrossFit, the main method, which splits source documents into two folds and scores questions from one fold with an auxiliary solver trained only on the other.
- Validates the mechanism through fixed-bank replay and controls on 3,000 saved questions, isolating source ancestry of the feedback solver from curriculum selection, evaluator duplication, arbitrary partitioning, and extra auxiliary optimization.
Main Findings
- Co-cheating grows with training. Under standard coupled feedback, false-agreement mass in round 1 is only 0.004 (Qwen3.5-4B) and 0.003 (Qwen3.5-9B), with incorrect labels appearing mostly as lost credit. By round 3, false-agreement mass reaches 0.061 and 0.088 while lost credit falls, showing disagreement being replaced by shared mistakes rather than by better answers.
- MSV helps only modestly. After self-evolution, MSV reduces false-agreement mass from 6.1% to 5.7% on Qwen3.5-4B and from 8.8% to 7.2% on Qwen3.5-9B, and it adds six labeler generations per candidate.
- CrossFit cuts false agreement roughly in half. It lowers false-agreement mass to 3.0% (4B) and 3.7% (9B), and the combined MSV + CrossFit treatment reaches 0.020 and 0.017 in the round-3 audit.
- Source-excluded replay nearly eliminates the failure. Replaying identical proposals with source-excluded feedback reduces false agreement to 0.4% (4B) and 0.1% (9B), isolating feedback ancestry from changes in the generated curriculum.
- CrossFit raises label truth. Round-3 adopted-label truth rises from 0.747 to 0.819 on Qwen3.5-4B and from 0.737 to 0.851 on Qwen3.5-9B relative to Dr. Zero.
- Downstream search improves at both scales. Evaluated on a fixed 1,325-question suite (200 questions each from Natural Questions, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA and MuSiQue, plus all 125 Bamboogle examples), CrossFit reaches 48.8% at 4B and 51.2% at 9B, improving over coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points.
- Multi-hop tasks gain most. Gains average 10.0/10.9 points across HotpotQA, 2WikiMQA, MuSiQue and Bamboogle, versus 7.3/5.2 across the single-hop datasets.
- The advantage accumulates across rounds, not within a fixed set. CrossFit's lead over Dr. Zero grows from 4.2 and 4.3 points after round 2 to 8.8 and 8.4 points after round 3.
- Verification is not a substitute for changing feedback provenance. MSV adds only 0.7–0.8 points over Dr. Zero, and combining it with CrossFit adds only 0.3 points beyond CrossFit alone at each scale (0.491 vs 0.488 at 4B; 0.515 vs 0.512 at 9B).
- A separate evaluator is not enough if it saw the same source. Same-source auxiliary solvers yield false-agreement mass of 0.064 and 0.087, and full-data auxiliary solvers 0.058 and 0.069, both close to the coupled control. Randomly splitting individual questions only reduces it to 0.050 and 0.062, whereas the source-ID split gives 0.004 and 0.001.
- Source exclusion also raises fixed-bank quality. Fixed-bank solver truth rises from 0.687 to 0.770 at 4B and from 0.717 to 0.868 at 9B, with accuracy rising from 88.1% to 91.5% and from 87.0% to 91.7%.
- Extra auxiliary optimization is not the cause. A half-budget control reaches false-agreement mass of 0.005/0.002 and replay accuracy of 91.6%/91.8%, matching the full source-ID result.
- Cost is real. MSV raises the reserved budget by about 90% at both scales (379 to 719 H200-hours on Qwen3.5-4B and 476 to 903 on Qwen3.5-9B); CrossFit at 25 updates per fold adds 72% and 79%, dropping to 36% and 40% if halved to 25 updates per round. Combining both interventions costs 2.6 and 2.7 times the Dr. Zero budget.
Methodology in Plain English
The authors start from an existing self-evolving search loop (Dr. Zero) in which a proposer writes questions from source documents, a solver answers them, and the match between the solver's five responses and the adopted pseudo-label becomes the proposer's reward. Rather than trusting that in-loop agreement, they run an independent auditor afterwards: for every audited step they save the source document, the adopted pseudo-label and the five solver responses, then have gpt-6-astra/high build an evidence-backed reference from the source and judge those same outputs. The auditor never touches training, so it gives a clean external measure of whether agreement was actually correct.
That audit motivates two fixes at different points in the loop. MSV intervenes before training: it samples three answers with the source and three without it, admits the task only if both sides produce a majority answer that agrees, and uses that compatible majority as the label. CrossFit intervenes on the feedback path: source documents are assigned once to fold 0 or fold 1, and questions inherit their source's fold. Two auxiliary feedback solvers each train only on one fold, and in the next round they are crossed — questions from fold 0 are scored by the solver trained on fold 1 and vice versa. The reward formula itself is unchanged (the original frontier reward f(k) = (5-k)/4), and the main solver still trains on all admitted questions; only the solver supplying the proposer's reward differs. Because a question is evaluated by a solver that never saw its source's pseudo-labels, a same-source error cannot be replayed back as apparent progress.
To separate the effect of feedback provenance from the effect of a changing curriculum, the authors also replay a fixed bank of 3,000 saved questions and labels while varying only the training provenance of the feedback solver, and they run controls that keep the auxiliary-solver architecture but change its data (same-source, full-data, random question-level split, half budget).
Why This Matters
The paper reframes a familiar-sounding reward-hacking story as a concrete, measurable pathology of self-generated curricula, and it shows that the fix is about who evaluates a proposal, not just how good the label is. It gives the field an audit protocol that can be applied to any loop where the same model family supplies both questions and feedback.
- Search and retrieval agents: Teams training agents to interleave reasoning with search can use source-level cross-fitting to avoid rewarding self-confirming errors, especially on multi-hop questions where the paper's gains are largest.
- Synthetic data and pseudo-labeling pipelines: The result that a separate judge is insufficient unless its training data exclude the evaluated source applies to any pipeline that reuses model-generated labels.
- Evaluation and auditing: The post-hoc, evidence-backed audit with coverage accounting (reported around 0.853–0.861 depending on treatment) offers a template for measuring internal training signals against external truth.
- Model and agent operations: Reporting reserved budget in H200-hours, tokens and judge requests per run lets practitioners weigh the reliability gain against roughly 1.4x to 2.7x compute costs.
Industry relevance: any organization that trains search-augmented assistants without human-annotated QA training data — the paper notes none of the self-evolution treatments uses human-annotated QA training data — faces exactly this failure mode, and the interventions are implementable by changing data routing rather than adding a new reward model.
Future Directions
- Extend exclusion to connected sources. The authors state that shared pretraining, overlapping web evidence, semantically related sources, and an adaptive proposer can still induce correlated mistakes, so source-level folds are only a partial fix.
- Measure end-to-end efficiency. The paper explicitly says its results do not yet establish lower end-to-end cost, and it reports auxiliary overhead of 72% and 79% for CrossFit at 25 updates per fold, or 36% and 40% when halved.
- Distinguish "fewer errors" from "easier tasks." Lower false-agreement mass can come from rejecting difficult proposals rather than improving learning, so coverage, task difficulty, fixed-probe performance and downstream capability must accompany that number.
- Establish statistical guarantees or independent structure. The authors borrow cross-fitting's exclusion principle, not its asymptotic guarantees, noting the adaptive curriculum lacks a demonstrated orthogonal score or independent sample structure; they also flag judge bias as motivation for human validation.
Target Audience
Researchers and engineers working on reinforcement-learned search agents, self-play and self-generated curricula, pseudo-label training, and reward design will get the most from this paper. It is also relevant to evaluation and trust-and-safety practitioners who need to detect when an agent's internal training signal has decoupled from real correctness, and to anyone building retrieval-augmented assistants without human-annotated supervision. The heavy reliance on reward formulas, audit statistics and multi-round training schedules makes it best suited to readers with some background in reinforcement learning for language models rather than complete beginners.
Authors’ abstract
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.