Skip to content
AI.info

Research

Studying Without a Syllabus: Task-Agnostic Environment Preprocessing

Overview Research area: LLM agent systems, specifically automated adaptation of agent harnesses and environments before deployment. Technical level: Advanced. The paper assumes familiarity with agent

Studying Without a Syllabus: Task-Agnostic Environment Preprocessing
arXiv
2609.10824
Published
2026-09-09
Authors
Vinay Samuel, Varun Ursekar, Vijay S. Kalmath, Apaar Shanker, Veronica Chatrath, Yuan Xue

AI summary

Overview

Research area: LLM agent systems, specifically automated adaptation of agent harnesses and environments before deployment.

Technical level: Advanced. The paper assumes familiarity with agent harnesses, retrieval pipelines, benchmark evaluation protocols (Avg@K / Best@K), and multi-model orchestration.

Scope: The paper formalizes and empirically tests "task-agnostic environment preprocessing," where an agent studies an unfamiliar environment without any downstream task examples and builds reusable artifacts for a frozen solver.

What This Paper Is About

Before an LLM agent works in a new environment, someone usually prepares it — building indices, writing scripts, or supplying guidance. Almost all automated methods for doing this depend on task examples, agent trajectories, or evaluation feedback, which are often unavailable at cold start. This paper asks whether a study agent can inspect an unfamiliar environment with no knowledge of the downstream task distribution, decide for itself how to prepare that environment, and produce artifacts that measurably help a frozen solver at test time.

Key Contributions

  1. A formalization of task-agnostic environment preprocessing. A studying system S takes a frozen solver and an environment E and returns a modified environment E_studied, incurring a study cost C_study(S,E) that must stay within a budget B_study. Crucially, S never receives real downstream task instances or any signal derived from their distribution — no traces, labels, verifier outputs, or evaluation feedback. Artifacts must be materialized as files the solver can access from inside its sandbox.

  2. Open-ended studying Meta-Agents in two variants. Meta-Agent w/o Archive must decide on its own how to explore and what to build. Meta-Agent w/ Archive is additionally given a mounted seed archive of executable skills and descriptions, including skills implementing PREPING, Corpus2Skill, and a general exploratory-study workflow resembling the unaided variant, which it may use, combine, or ignore.

  3. A comparative evaluation against two fixed, task-agnostic baselines. PREPING (environment-grounded synthetic practice distilled into a procedural playbook) and Corpus2Skill (a non-agentic pipeline that compiles a document corpus into a navigable skill hierarchy) are run under a common protocol across six heterogeneous benchmarks.

  4. Evidence that studying shifts compute from test time to pre-task time. The paper shows that studied artifacts reduce the number of test-time rollouts needed to reach a given score, while showing that larger study budgets do not reliably improve downstream reward.

Main Findings

  • Meta-Agent variants lead on five of six benchmarks. A Meta-Agent variant achieves the highest Avg@3 and Best@3 reward on 5 of 6 benchmarks, and Meta-Agent w/ Archive ranks first or second on every benchmark under both metrics.

  • Corpus2Skill is specialized, not general. It is strongest on BCP-G (Avg@3 = 0.471 ± 0.015; Best@3 = 0.680 ± 0.017) but falls below No Study on OfficeQA, Harvey LAB, and DABStep. On AppWorld it scores an Avg@3 of 0.603 ± 0.028 versus No Study's 0.576 ± 0.013.

  • PREPING is consistently beneficial but never best. It exceeds No Study on all six benchmarks under both metrics (e.g., Avg@3 on DABStep 0.464 ± 0.060 vs. No Study's 0.353 ± 0.008; Best@3 on DABStep 0.565 ± 0.042 vs. 0.487 ± 0.015), yet never ranks first.

  • Archive access produces occasional large gains, not uniform improvement. Under Avg@3, the mean improvement on the four benchmarks where the archive helps is 0.046, against a mean decline of 0.021 on the two where it hurts. The largest single gain is 0.129 on DABStep (0.400 ± 0.006 to 0.529 ± 0.097); the largest decline is 0.023. Best@3 shows the same asymmetry: a mean gain of 0.041 where it helps and a mean loss of 0.017 where it hurts.

  • More study budget does not reliably help. Budget scaling was tested on OfficeQA, Harvey LAB, and Apex Agents for budgets of $1, $5, $10, $25, and $50. OfficeQA and Apex Agents show flat or non-monotonic curves with mostly overlapping intervals. Harvey LAB is the exception: Meta-Agent w/o Archive rises from 0.271 to 0.348 and Meta-Agent w/ Archive from 0.249 to 0.332, while PREPING stays roughly flat after $5.

  • Studying reduces required test-time sampling. To exceed the strongest studied Best@1 score, No Study requires 8 rollouts on BCP-G, 2 on OfficeQA, 3 on Harvey LAB, 5 on DABStep, 2 on Apex Agents, and 4 on AppWorld. Using average per-rollout inference costs, reaching those points costs No Study 1.6–5.5 times as much task-time inference as the corresponding studied rollout. This excludes upfront study cost, so studying yields a total saving only when artifacts amortize across enough downstream tasks.

  • At a higher target the gap widens. At the Best@3 target for studied agents, No Study requires 6 rollouts on OfficeQA, 7 on Harvey LAB, and 5 on Apex Agents, and does not reach the strongest studied score within 10 rollouts on BCP-G, DABStep, or AppWorld.

  • Deeper exploration does not predict better artifacts. On OfficeQA, Corpus2Skill processes all 697 files yet trails methods that inspect fewer. Harvey LAB is the clearest split: both meta-agents discover its single benchmark-specific tool, labread, and score 0.35–0.38, while methods that never invoke it score 0.22–0.31.

  • Study agents share a diagnostic core but specialize. Inventory-taking, unavailable-resource tracking, and budget-conscious extraction recur across environments. Test anticipation concentrates in Apex Agents and AppWorld; parallel delegation is observed only in Harvey LAB and Apex Agents.

  • Archive composition is environment-sensitive but imperfect. Open-ended-study and PREPING outputs appear in every BCP-G, OfficeQA, and DABStep artifact set; Corpus2Skill appears in every AppWorld artifact set and in 86 of 93 Apex Agents artifact sets. On BCP-G the Meta-Agent omits Corpus2Skill even though it is the strongest standalone method there.

  • Gains concentrate in specific task strata. On DABStep, Meta-Agent w/ Archive improves reward by 0.185 on hard tasks but only 0.015 on easy tasks, where No Study already scores 0.792. On BCP-G, Corpus2Skill improves all ten topics, with lift rising from 0.224 with one gold document to 0.293 with four or more. Every method gets its largest AppWorld gain on two-application tasks and its largest Apex Agents domain gain in Law.

  • Artifacts can misdirect the solver. One DABStep artifact correctly supplies missing information about a fraud boundary. In Apex Agents, an incomplete artifact steers the solver toward general pricing documents and away from two task-specific files containing the required inputs.

  • Overall, 21 of 24 method–benchmark comparisons improve over No Study.

Methodology in Plain English

The researchers set up each benchmark as a Docker sandbox containing a corpus and/or tools, and treat each sandbox as an "environment" to be studied. A frozen solver — the same Claude Code harness with Claude Haiku 4.5 at temperature 1.0 in every condition — then attempts held-out tasks inside that environment. The solver's weights and harness never change; only the files it finds in its sandbox change.

Four studying methods produce those files. PREPING runs synthetic practice cycles (a proposer generates tasks from available tools, a solver attempts them, a validator filters, and a reflector/curator distills lessons), then prepends retrieved playbook entries at test time. Corpus2Skill runs a fixed, non-agentic pipeline that summarizes, embeds, clusters, and indexes documents into a navigable hierarchy. The Meta-Agent w/o Archive variant is simply a study agent told to explore the environment and build whatever file-based artifacts it judges useful. The Meta-Agent w/ Archive variant gets a mounted set of executable skills, including implementations of the two baselines, and decides which to invoke and how to combine their outputs. Claude Opus 4.8 drives Meta-Agent study and PREPING's proposal, validation, reflection, and curation; Claude Haiku 4.5 runs PREPING's synthetic tasks and all downstream tasks; Claude Sonnet 4.6 handles Corpus2Skill's clustering and labeling. Corpus2Skill keeps Qwen3-Embedding-8B for embeddings; PREPING keeps text-embedding-3-small for playbook retrieval.

For measurement, each studying method is run 3 times per environment, producing 3 artifact sets, and each artifact set is evaluated with 3 independent downstream repetitions. Avg@3 averages all three rollout rewards per task; Best@3 takes the best of three. No Study has no artifact dimension and is evaluated with 10 repetitions, with Avg@3 and Best@3 computed over all 120 three-repetition subsets; its standard deviations are population SDs across those subsets, while studying methods report sample SDs across the 3 study runs. Harvey LAB is scored using a dense rubric criterion pass rate rather than its original strict all-criteria pass indicator. The primary comparisons impose no shared study budget — each method runs its full configuration — while the budget-scaling experiment uses budgets of $1 through $50.

Why This Matters

This work reframes agent adaptation as a pre-task study phase rather than a test-time optimization loop. Instead of repeatedly sampling rollouts and using feedback to improve a harness, a system can pay once to build reusable structure — indices, scripts, guidance — that a frozen solver then exploits across many tasks. The paper also contributes a clean experimental protocol for a question that had been addressed piecemeal: when no task supervision exists, which form of preparation actually transfers, and does the answer depend on the environment?

Real-world applications:

  • Enterprise document systems. The Harvey LAB and OfficeQA settings resemble legal, financial, and compliance environments where large document stores must be made navigable before anyone knows which questions will be asked.
  • Cold-start deployments. When a new customer environment is onboarded, representative tasks may not exist yet; task-agnostic preprocessing lets a system prepare the environment anyway.
  • Long-horizon agents in heterogeneous worlds. The Apex Agents results — 31 independently studied worlds across investment banking, law, and management consulting — map onto deployments where each team or project has its own filesystem and tools.
  • Cost planning for agent products. The test-time scaling result gives teams a way to trade pre-task study spend against per-query inference spend, with amortization depending on how many queries follow.

Industry relevance: Study cost and downstream inference cost are both reported in API dollars, making the trade-off directly actionable for teams that pay per token. The finding that bigger study budgets do not reliably buy better artifacts cautions against simply throwing compute at preparation, and the archive result — occasional large gains without comparably large regressions — suggests a comparatively safe way to widen a study agent's repertoire. The Responsible Use section also flags concrete deployment risks: exploration side effects, sensitive artifacts, misleading guidance, and archive misuse.

Future Directions

  • Better strategy selection. A Meta-Agent variant leads on five of six benchmarks, but the specialized Corpus2Skill pipeline still wins on BCP-G, and the archive-equipped agent sometimes omits the strongest available workflow. A broader archive, or demonstrations of successful environment–strategy matches, could improve in-context selection.
  • Learning to select across deployments. The paper suggests that repeated deployments could let downstream evaluation feedback update the Meta-Agent's strategy-selection policy through weights or harness, turning cross-environment experience into a form of meta-learning.
  • Calibrated budget tracking and planning. Since budget scaling appeared only on Harvey LAB and only for Meta-Agents, the authors argue current models may not track their own consumption or plan differently as the allowance changes, so a larger budget may extend the same strategy rather than induce a new one. The experiment measures how the current policy responds, not what a calibrated studier could achieve.
  • Generalization beyond one solver and one archive. All studiers and the frozen solver share a model family and the Claude Code harness, so transfer across solvers is unestablished. The authors also did not compare the Meta-Agent against a non-agentic classifier that picks one fixed workflow from the same archive from an environment description, leaving the value of agentic decision-making and of individual archive entries unresolved.

Target Audience

Researchers and engineers working on agent harnesses, memory, and retrieval or workflow optimization will find the formalization and benchmark protocol most useful, as will practitioners who need to prepare specialized environments — legal, financial, consulting, document-heavy — before downstream tasks are known. The paper is written for readers comfortable with agentic evaluation design; the budget-scaling and test-time-scaling

Authors’ abstract

Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural guidance. Most automated adaptation methods, however, rely on task examples, trajectories, or evaluation feedback to decide what to build. Existing task-agnostic approaches avoid this supervision but commit in advance to a preparation strategy for a particular type of environment. We study a more open-ended setting: can an agent study an unfamiliar environment without a syllabus, i.e. before test time and without knowledge of the downstream task distribution, and choose how to prepare it? We formalize task-agnostic environment preprocessing, in which a studying system explores an environment under a budget and produces artifacts for a frozen solver. We compare unaided and archive-equipped meta-agents with fixed synthetic-practice and corpus-processing methods across six heterogeneous benchmarks. A meta-agent variant achieves the highest Avg@3 reward on five benchmarks, while fixed corpus processing remains best on the largest corpus benchmark. Larger study budgets do not reliably improve downstream reward. Nevertheless, studied artifacts reduce the test-time sampling needed to reach a given score, demonstrating how reusable preparation can shift computation from repeated test-time attempts to a pre-task study phase.

Read the original paper