Skip to content
AI.info

Research

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

Overview Research area: Natural Language Processing, specifically synthetic training-data generation for LLM "working agents" that read files, coordinate tools, and produce deliverable artifacts. Tech

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
arXiv
2609.38923
Published
2026-09-30
Authors
Qisheng Su, Hanchen Wang, Guanru Zhu, Huicheng Jiang, Qiuyinzhe Zhang, Kou Shi, Zhen Fang, Ziao Zhang, Qingnan Ren, Zehui Chen, Tao Gui, Feng Zhao

AI summary

Overview

Research area: Natural Language Processing, specifically synthetic training-data generation for LLM "working agents" that read files, coordinate tools, and produce deliverable artifacts.

Technical level: Advanced. The paper assumes familiarity with supervised fine-tuning, rejection fine-tuning, agent scaffolds, Bradley–Terry/Elo rating models, and rubric-based LLM judging.

Scope: The paper introduces GraphForge, an evidence-graph framework that synthesizes working-agent training tasks and verification rubrics from real crawled files seeded by O*NET occupations, then validates it by fine-tuning Qwen3.6-27B and Qwen3.6-35B-A3B on 2,169 trajectories and evaluating on GDPVal-AA, Workspace-Bench-Lite, and SpreadsheetBench II.

What This Paper Is About

Training agents that do real office-style work (reading diverse files, using tools, producing usable deliverables) requires large numbers of tasks built on real files with verifiable results, and few pipelines exist to build such data. Existing pipelines either generate files with a model, which the paper argues lack realism and diversity and are checked only by state-check scripts, or build tasks on real files with no task-specific verifiers, leaving result quality unchecked. GraphForge addresses both gaps by grounding the task statement and the verification rubrics in an evidence graph built over real crawled files.

Key Contributions

  1. An evidence-graph framework for grounding both task and verification in real files. Starting from occupation-grounded seeds derived from O*NET and AI4Work DIGITAL annotations, task statements and rubrics are compiled from an evidence graph over real files so that each rubric criterion is anchored to the specific files needed to verify it, and each task is validated through an initial teacher rollout before trajectory collection.

  2. Empirical validation across two base models and three benchmarks. Training Qwen3.6-27B and Qwen3.6-35B-A3B on GraphForge data yields gains on GDPVal, Workspace-Bench-Lite, and SpreadsheetBench II under multiple scaffolds, with analysis suggesting the gains reflect transferable working skills rather than memorization of benchmark content.

  3. A controlled rubric-guided rejection fine-tuning study. Three RFT arms sharing the same 462 queries, candidate pools, and optimization budget differ only in selection rule (anchored, unanchored, random-of-4), isolating within-query trajectory selection and suggesting that the evidence-anchored rubrics provide a useful selection signal.

  4. Release of data and models. The synthesized tasks, rubrics, trajectories, and trained checkpoints are released publicly (with source links provided for workspace files rather than redistributed file contents).

Main Findings

  • SFT on GraphForge data produces large benchmark gains. Qwen3.6-27B reaches GDPVal 1445.7 (+65.7) on OpenHands and 1427.4 (+63.4) on Codex, Workspace-Bench-Lite 63.7 (+7.7) on Claude Code and 65.2 (+3.8) on Codex, and SpreadsheetBench II 24.0 (+13.7) on Claude Code and 24.6 (+9.0) on Codex. Qwen3.6-35B-A3B improves by 101.7 Elo on OpenHands and 101.4 Elo on Codex on GDPVal, up to 6.6 and 7.7 points on Workspace-Bench-Lite, and up to 16.5 and 13.7 points on SpreadsheetBench II.

  • Gains transfer across scaffolds. All GraphForge trajectories were rolled out with the Codex scaffold, while evaluation covers OpenHands and Codex on GDPVal and Claude Code and Codex on Workspace-Bench-Lite and SpreadsheetBench II; the SFT model improves over its base model under every scaffold.

  • Corpus diversity. The SFT corpus contains 2,169 admitted trajectories (Q > 0.90) drawn from 3,638 materialized tasks, covering 466 distinct O*NET task types, 15 of 16 occupational sectors, and all 16 execution patterns. 164 of 256 sector–pattern combinations are realized. Data analysis and reporting accounts for 14.2% of trajectories and research and source synthesis for 13.2%; no other pattern exceeds 9%. PDF appears in 96.7% of trajectories.

  • Trajectory length. A trajectory contains 50.0 assistant steps on average (median 48, 95th percentile 82) and 162.0k tokens on average (median 158.0k, 95th percentile 224.8k). 28 sequences (1.3%) reach the 262,144-token training ceiling.

  • Rubric-guided RFT helps on Workspace-Bench-Lite and SpreadsheetBench II. Anchored selection gives the largest gains over SFT on both; unanchored selection gives smaller or negative gains; random-of-4 is consistently weakest. On GDPVal, anchored RFT improves over SFT (+7.2 and +11.0 Elo) and random selection degrades it (-9.0 and -9.5), while the unanchored arm shows higher point estimates (+34.1 and +25.3). The paper states none of the GDPVal differences is statistically resolved at this scale, and all RFT difference intervals in Appendix D include zero.

  • Contamination audit. None of the 39,201 training files coincides with any of the 260 GDPVal files (0 shared hashes at the file level). At the text level, the top-20 most similar 13-gram pairs contain no substantive shared content. Only 13 of the 44 GDPVal occupations are covered by the training taxonomy.

  • Transfer is not concentrated on covered occupations. SFT win rate against the base model is 0.739 (95% CI [0.671, 0.803]) on the 155 occupation-uncovered GDPVal tasks, versus 0.692 (95% CI [0.585, 0.800]) on the 65 covered tasks; overall 0.725 (95% CI [0.668, 0.782]).

  • Judge sensitivity is localized but coarse. Deleting the worksheet a criterion cites drops the corresponding score by 0.377 on average, while non-target criteria remain essentially unchanged (mean |ΔQ| = 0.016). Fine-grained corruptions of rows, numbers, and citations cause only small changes (below 0.03 in magnitude: -0.018, -0.022, -0.013), which the authors attribute to the capability limit of GLM-5.2 as an agentic judge.

  • Comparison to a frontier model. Claude Opus 5 leads the reported GDPVal Elo table at 1774.1 (OpenHands) and 1753.1 (Codex), above the GraphForge SFT models. The open-weight Nex-N2-Mini-35B baseline scores 1288.8 and 1342.3 on GDPVal, 33.1 and 31.6 on Workspace-Bench-Lite, and 6.5 and 10.3 on SpreadsheetBench II.

Methodology in Plain English

GraphForge turns occupational task seeds into verified training trajectories in five stages.

Seed construction. Seeds come from the O*NET database. Only tasks annotated DIGITAL by AI4Work are kept. After filtering, 246 occupations across 16 sectors and 43 sub-sectors remain, covering 891 DWA task types and 3,419 valid occupation–task-type pairs. Each seed is a tuple of occupation, Detailed Work Activity (task type), work demand, dominant execution pattern, expected input file family, and retrieved occupational evidence. Work demands must be directly supported by cited evidence; unsupported demands are dropped rather than filled to a quota. Execution patterns come from a vocabulary of 16. Seeds are chosen greedily by marginal coverage gain over occupation, task type, pattern, and input family, with a further rebalancing pass after files are downloaded.

Workspace construction. A seed names no company, event, dataset, or result. A search agent finds a coherent public case and retrieves the files needed to do the work, downloading them in native formats. Exact duplicates and invalid files are removed. Each retained file receives a hidden role: core, supporting, confuser, or ambient. These roles guide assembly and are never shown to the working agent.

Evidence graphs and rubrics. An evidence graph is built over the workspace: nodes record a source file, the fact it provides, and its role; edges record cross-file dependencies needed to interpret, compare, reconcile, or derive information. The graph is described as an intermediate representation, not ground truth, whose claims must remain recoverable from the original files. The graph compiles into a task specification containing a natural-language task statement, deliverable requirements, positive criteria, and negative penalty criteria. Each positive criterion specifies a target deliverable, requirement, expected value or computation, weight, evidence anchors, and a verification procedure. The working agent sees only the task statement and workspace; the judge additionally sees the anchors and verification instructions.

One-step revision. Each initial specification is run once with a strong teacher model. A revision agent then receives the original files, graph, task, rubrics, and the execution, and returns only the components needing change. The teacher is rerun only when the task statement changes or the initial trajectory is missing; otherwise the initial execution is reused and re-judged.

Admission and cleaning. A trajectory is admitted only after its promised deliverables are materialized. Deterministic checks verify filenames, formats, required sheets and formulas, and structural constraints. An agent judge then scores every criterion, computing a weighted score in which negative criteria subtract penalties proportional to the evidenced degree of violation (binary at this stage). Scores are used raw and may fall below zero. Trajectories with degenerate tool-use behavior (repeated non-polling calls, excessive tool use, high tool-failure rates, repeated truncation) are discarded.

Training and evaluation. GLM-5.2 powers all pipeline components. SFT uses 3 epochs of Muon, peak learning rate 2×10⁻⁵, cosine schedule, warmup ratio 0.10, weight decay 0.05, global batch size 8, and packed sequences up to 262,144 tokens. For RFT, K = 4 rollouts per query are sampled from the SFT model on 2,000 queries (125 per execution pattern); the judge scores the four candidates jointly, the highest-scoring valid trajectory is kept when its score exceeds 0.95, and 462 trajectories result. Evaluation is pass@1 with a fixed workspace interface per scaffold, with missing or invalid deliverables counted as losses. GDPVal Elo is fitted with a Bradley–Terry model over a comparison graph, anchoring GLM-5.3 (OpenHands) at 1667.

Why This Matters

Impact on research. The paper targets a structural gap in agent training data: realism versus verifiability. Prior working-agent pipelines either generate files with a model (sacrificing realism and diversity) or use real files without task-specific verifiers (leaving results unchecked). GraphForge's evidence-anchored rubrics make the same artifact that defines the task also define its verification, and the paper reports a controlled ablation showing that selection quality matters for turning self-generated rollouts into training signal.

Real-world applications:

  • Business and financial analysis: assembling deliverables from corporate filings and public reports, and reconciling figures across multiple source documents.
  • Spreadsheet and reporting workflows: producing multi-sheet workbooks with required sheets and formulas that can be checked structurally.
  • Professional services work of the kind APEX-Agents models with investment banking analysts, management consultants, and corporate lawyers operating in data-rich simulated environments.
  • Cross-file document reasoning at the scale Workspace-Bench targets, where agents must trace dependencies across tens of thousands of files.

Industry relevance. The reported gains are measured on benchmarks the paper describes as realistic work tasks (GDPVal covering 1,320 tasks across 44 occupations with expert-written rubrics; Workspace-Bench with tens of thousands of files; SpreadsheetBench II with expert-annotated multi-sheet business workbooks). The results suggest that open-weight models in the 27B–35B range can be substantially improved for this style of work using synthesized data, and the paper releases both data and checkpoints. The gains transferring across OpenHands, Codex, and Claude Code scaffolds is relevant to deployment, where the scaffolding around a model often differs from the one used in training.

Future Directions

  • Scaling the data budget. The paper states its corpus contains 2,169 trajectories and that it has not studied how GraphForge's benefits scale with larger data budgets.
  • Stronger models for synthesis and judging. Both the judge and the synthesis pipeline use GLM-5.2; the authors note that stronger frontier models could improve synthesized data quality and judging reliability, and that exploring the framework's ceiling with such models remains future work. The judge's inability to catch fine-grained row, number, and citation corruptions is a specific instance of this.
  • Transfer across model families. Experiments cover two base models from the same family, and the paper does not study how GraphForge transfers to other model families.
  • Broader task families, file types, and horizons. The authors plan to scale GraphForge to more task families and file types and to study how evidence-anchored verification interacts with longer-horizon agent scaffolds.

Target Audience

Researchers and engineers who build or train LLM agents for file-based knowledge work, especially those working on synthetic data generation, agent benchmarks, or rubric-based reward and verification. It is also relevant to practitioners who fine-tune open-weight models (in the tens-of-billions parameter range) for document, spreadsheet, and multi-tool workflows, and to benchmark designers interested in how task admission and trajectory selection interact with evaluation noise. Readers without background in agent training pipelines, SFT/RFT, and Elo-style pairwise evaluation will find the method and results sections demanding.

Authors’ abstract

Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal. The data and models are available.

Read the original paper