Skip to content
AI.info

Research

Learning to Learn from Context: Synthetic Training from Perturbed Public Documents

Overview Research area: Natural Language Processing — synthetic training data for large language models, specifically the capability of "context learning" (learning a task's rules, facts, and procedur

Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
arXiv
2609.33642
Published
2026-09-27
Authors
Haoyi Wu, Yang Xiao, Yusong Sun, Wenyang Hui, Zhaokai Luo, Chengyue Jiang, Mu Chuan

AI summary

Overview

  • Research area: Natural Language Processing — synthetic training data for large language models, specifically the capability of "context learning" (learning a task's rules, facts, and procedures from a document supplied at inference time rather than from pretrained memory).
  • Technical level: Intermediate. The high-level idea is accessible, but reproducing the work requires familiarity with supervised fine-tuning, reinforcement learning from rubric rewards, and long-context evaluation.
  • Scope: The paper presents a human-annotation-free pipeline that perturbs public documents to create ~10k long-context question-answer samples, and shows that training on them improves a 35B student model on CL-bench and transfers to several other general capabilities.

What This Paper Is About

Real tasks increasingly require an LLM to read a long document and reason over it, but models remain weak at this: even frontier models solve fewer than 30% of the 1,899 tasks in CL-bench, which spans 18 domains. Using public documents for such training is risky because the public web has largely been consumed during pretraining, so naive training rewards memorization rather than genuine document use. The paper's goal is to turn already-seen public documents into effective context-learning supervision without any human annotators, by perturbing the documents so answers can no longer come from parametric memory.

Key Contributions

  1. A synthesis pipeline over perturbed public documents. Source documents are deep-rewritten (proper-noun renaming, numeric perturbation, section renumbering, list-order shuffling), then paired with generated questions and rubrics, answered by a teacher with its native reasoning trace preserved.
  2. A document-dependence "gap check." Each question is answered with and without the document; a sample is admitted only if the rubric pass rate with the document is at least 0.60, the pass rate without it is at most 0.30, and the difference is at least 0.40.
  3. A no-annotation dataset of about 10k samples. The pipeline produces 9,625 training samples from 3,515 public documents, covering 16 of CL-bench's 18 subcategories, with each sample's rubric list reusable as an RL reward prompt.
  4. Empirical validation through SFT and rubric-reward RL, including ablations on the gap check, data scaling, reward design, and generalization to fictional (non-public) contexts.

Main Findings

  • Large gains on the target benchmark: SFT raises Qwen3.6-35B-A3B from 13.7 to 22.8 on CL-bench, and a subsequent rubric-reward RL stage reaches 24.6, comparable to HY3 (23.5) and Qwen3.8-2.4T (23.9). On CL-bench Life the same models move from 8.4 (baseline) to 12.6 (SFT) to 13.8 (SFT + RL).
  • RL needs SFT initialization: Applying RL directly to the base model improves CL-bench only from 13.7 to 16.9, substantially worse than initializing from SFT.
  • Results are strongly teacher-dependent, and not in the expected direction: With identical questions, students trained on different teachers' answers span 19.7 to 22.8 on CL-bench. Kimi-K3, the strongest teacher (27.0), yields the weakest student (19.7); Qwen3.8-2.4T, a mid-ranked teacher (23.9), yields the best student (22.8). Students inherit the teacher's reasoning length more than its performance.
  • The gap check matters: Training on 3,000 samples that pass the gap check gives 19.17% on CL-bench versus 17.75% for 3,000 unfiltered samples; 32.5% of generated samples fail the check before retry.
  • Data scaling has not saturated: Accuracy rises monotonically with sample count — from 13.7 untrained to 22.8 at the full 9,625-sample pool, with a rapid gain from 16.0 to 20.2 when the count reaches 1,000 — and shows no saturation within the tested range.
  • Reward design is a bottleneck: Standard RL gives 24.6 with median think length 10,481 tokens. Adding a length penalty cuts length to 4,012 but drops accuracy to 20.91; mixing in the task-level pass rate (alpha = 0.5) leaves accuracy at 22.80 while raising think length to 14,398. The RL model also has a higher failure rate (about 2%) on anti-hallucination rubrics.
  • Broad transfer, with knowledge flat or slightly down: SFT improves long-context understanding (LongBench v2 +2.9, AALCR +5.4, MRCR-4needle +2.1), instruction following (IFBench +14.0, AdvancedIF +3.3), and reasoning (ARC-AGI-1 +16.4, ARC-AGI-2 +11.8, GPQA-Diamond +5.1, HMMT-2025 +4.1). Code is essentially unchanged (LiveCodeBench-v6 +0.5, OJBench -0.9), and knowledge is mostly flat with small declines (CEval -1.0, SimpleQA -2.2), against MMLU-Pro +0.7 and SuperGPQA +0.6.
  • RL does not uniformly transfer: It improves AALCR, GPQA-Diamond, and HMMT-2025 over SFT, but ARC-AGI-1 falls by 12.0 and ARC-AGI-2 by 9.4.
  • Gains concentrate on fictional contexts, not public documents: In an unofficial audit of the 500 CL-bench contexts, 277 (55.4%) are public documents, 36 (7.2%) modified public documents, 184 (36.8%) fictional, and 3 (0.6%) unclear; among the 1,899 tasks, 836 (44.0%) are public-document based and 1,063 (56.0%) fictional. Accuracy on fictional/unclear contexts rises from 17.5 (base) to 31.1 (SFT) to 32.9 (SFT + RL), while on public documents it rises from 9.0 to 12.2 to 14.1.
  • Teacher identity affects generalization differently from the target benchmark: After RL, the GLM-5.2-initialized and Qwen3.8-2.4T-initialized models reach similar CL-bench scores (24.5 vs 24.6), but on CL-bench Life they reach 9.1 and 13.8 respectively.

Methodology in Plain English

The researchers start from existing public documents rather than writing new ones. To stop the model from answering out of memory, they have GLM-5.2 rewrite each document — renaming entities, changing numbers, renumbering sections, shuffling list order. For each rewritten document, a generator proposes up to three questions with rubric lists of at least seven items each, drawn from 9 persona archetypes, 7 question types, and 4 rubric types, with at least five rubrics required to cite specific values from the document.

Each question is then answered twice by a teacher model: once with the rewritten document as context and once with the document removed (with an explicit instruction to answer from memory). A judge scores both answers against the rubrics, and only samples where the document clearly helps are kept. Rejected samples are regenerated once using the rejection reason as extra context, then discarded if they fail again.

The student (Qwen3.6-35B-A3B) is trained on the admitted samples using the teacher's verbatim reasoning trace followed by the answer — no post-hoc polishing of reasoning. Training runs for 3 epochs with a global batch size of 32 and a maximum sequence length of 131K, using AdamW with beta1 = 0.9 and beta2 = 0.95, cosine decay from a peak learning rate of 5e-6 to 1e-7, 10% warmup, weight decay 0.1, and gradient clipping 1.0. A second stage applies GSPO reinforcement learning on the same prompts, with each sample's rubrics providing the reward: the mean per-rubric pass/fail judged by Qwen3.5-397B-A17B, using the CL-bench judge prompt verbatim, with a KL penalty coefficient of 0.001 and a constant learning rate of 1e-6. Evaluation uses a GPT-5.1 judge, with low reasoning effort for CL-bench and high for CL-bench Life.

Why This Matters

The work shows that documents a model has already seen during pretraining can still be turned into useful supervision, provided they are perturbed enough that memory no longer supplies the answer. That removes the main cost barrier to context-learning data — expert human annotation of long documents — and makes the pipeline reproducible at scale.

Real-world applications:

  • Document question answering over internal manuals and specifications, where a model must apply rules from a provided document rather than recall general knowledge.
  • Retrieval-augmented reasoning, where the model must reason over retrieved passages rather than paraphrase them.
  • Agentic workflows, where the model must interpret tool outputs and logs it has never seen before.
  • Compliance checking, verifying whether a plan conforms to a specification supplied in context.

Industry relevance: the approach converts public corpora into training data without paid annotation, and a 35B model reaching CL-bench scores comparable to frontier models reported as 10–100× larger points to a cheaper route to strong context learning. The observed trade-off — gains in context use alongside small drops in knowledge benchmarks such as SimpleQA (20.9 to 18.7) and CEval (92.0 to 91.0) — is directly relevant to teams deciding what to train on.

Future Directions

  • Scaling the data. The scaling curve shows no saturation up to the full 9,625-sample pool, so a larger pool may yield further gains, though the exact ceiling is not reported.
  • Understanding the teacher–student relationship. Why the strongest teacher (Kimi-K3, 27.0) produces the weakest student (19.7) while a mid-ranked teacher (Qwen3.8-2.4T, 23.9) produces the best student (22.8) is left to future work, as is the failure of the instruction-following data patch, which lowered CL-bench performance further.
  • Better RL rewards. Both tested alternatives were worse than the standard rubric reward (alpha = 0.5 gave 22.80, length penalty gave 20.91 versus 24.60), so avoiding reward hacking without losing task accuracy remains open.
  • Covering the full benchmark taxonomy. The corpus covers 16 of CL-bench's 18 subcategories; Workflow Orchestration and Operational Procedures were omitted because they are dominated by agent-style, conversation-form traces that are hard to source from public documents.

Target Audience

Researchers and engineers working on long-context language models, synthetic data generation, and post-training (SFT and RL) will benefit most. It is also relevant to practitioners building RAG, document-analysis, and agent systems who need models that follow instructions grounded in a supplied context, and to anyone weighing the trade-off between context-learning ability and retained world knowledge. Readers looking for a fully solved recipe should note the paper's own caveats: results are teacher-dependent, reward design is unresolved, and the provenance analysis of CL-bench contexts is described by the authors as an unofficial rough estimate.

Authors’ abstract

Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality documents are an abundant alternative, but much of the public web has already been consumed during pretraining: training on such documents naively would reward memorization rather than context learning. In this work, we attempt to make use of high-quality public documents with small perturbations and empirically find that LLMs can successfully generate context-dependent reasoning traces and answers, which are then used to train a student model. Specifically, we construct a synthesis pipeline that (i) rewrites source documents to reduce memorization risk, (ii) generates questions and rubrics that require reasoning over the document, (iii) answers the questions with the document as context, and (iv) admits only samples that genuinely depend on the document. Without any human annotators, our pipeline generates about 10k samples from 3.5k documents, and the resulting student model substantially improves the performance on CL-bench. SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%). We also observe a broad transfer of improvements to long-context understanding, instruction following, and reasoning, while code generation and knowledge remain mostly flat. We hope this work provides a reproducible and scalable way to improve the ability of LLMs to learn from context, and to facilitate further research on context-grounded reasoning.

Read the original paper