Research
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
Overview Research area: Natural Language Processing / LLM evaluation for automated scientific research (research agents, scientific ideation, and experimental execution). Technical level: Intermediate

- arXiv
- 2609.10539
- Published
- 2026-09-09
- Authors
- Yiling Ma, Yilun Zhao, Sihong Wu, Manasi Patwardhan, Arman Cohan
AI summary
Overview
Research area: Natural Language Processing / LLM evaluation for automated scientific research (research agents, scientific ideation, and experimental execution).
Technical level: Intermediate. The paper assumes familiarity with LLM benchmarking, classification metrics (Macro-F1, recovery rates), and the pipeline of turning a research idea into runnable code, but it does not require deep technical background in any specific subfield.
Scope: The paper introduces and validates IdeaAmbig, a 660-instance benchmark that measures whether LLMs can detect, locate, and clarify implementation-blocking gaps in research-method specifications before those specifications are turned into code.
What This Paper Is About
A research idea can be novel, coherent, and scientifically sound, yet still be too vague for anyone (human or AI) to implement faithfully. When a method-defining detail is missing or ambiguous, an implementer silently picks something, and the resulting code may implement a different method than the authors intended. This paper builds a benchmark to test whether LLMs can notice that a specification is not ready to code, pinpoint the exact missing decision, and ask the right clarifying question to resolve it.
Key Contributions
-
A formal definition of "codification readiness." The authors define an idea specification as codification-ready when a competent implementer can build a faithful initial implementation without unsupported assumptions about the core method. They then decompose this into three evaluable tasks: readiness assessment, defect localization, and clarification action generation.
-
The IdeaAmbig benchmark. 660 single-defect instances grounded in real evidence: 163 real-world gaps (121 from GitHub issue threads, 42 from reproducibility reports) and 497 controlled synthetic gaps injected into verified codification-ready references. Each instance pairs an underspecified specification with a resolved counterpart, a defect label (3 Level-1 types, 10 Level-2 categories), and a gold clarification action.
-
A multi-model diagnostic evaluation. Thirteen LLMs were tested, revealing a consistent pattern: models are reasonably good at clarifying once told what the defect is, but poor at finding the defect on their own. The authors separate these two capabilities experimentally rather than conflating them.
-
Oracle and executable utility studies. Supplying the gold resolution to the model raised downstream codification-ready rates from 14% to 98%, and in an executable study raised test-passing implementations from 45% to 85% and faithful implementation of the target method detail from 30% to 90%.
Main Findings
-
Localization is the dominant bottleneck. The strongest model (GPT-5.6-Sol) achieved only 9.6% Macro Defect Recovery Rate on real-world instances, versus 80.6% Macro Clarification Action Success Rate when the annotated defect was handed to it. This gap held across all 13 models.
-
Readiness judgments are unreliable and poorly grounded. GPT-5.6-Sol scored 67.5 Macro-F1 on real-world readiness assessment, accepting 31% of underspecified specifications and rejecting 34% of ready ones. The best Reason Grounding Score was only 0.36 on real-world instances: even correct labels rarely came with rationales that identified the relevant blocker.
-
Models predict plausible categories without finding the actual blocker. On real-world instances, the top model reached 60.1 Level-1 accuracy and 25.2 Level-2 accuracy, but only 16.0% localization accuracy. It recovered the annotated blocker in just 16.0% of cases, identified a neighboring decision in 17.8%, and a different blocker in 66.3%.
-
Synthetic instances are easier than real ones. The same model scored 12.2 Macro DRR on synthetic instances versus 9.6 on real-world ones, and 96.2 versus 80.6 Macro-CAS. Synthetic targets are more explicitly stated, which inflates performance relative to naturally occurring gaps.
-
Failures are about missing information, not guessing. When clarification failed, the dominant error was insufficient clarification (19.6% of cases), not unsupported assumptions (4.3%). The model's No-Assumption score was 95.7 versus 80.4 Sufficiency, indicating it avoids inventing details but often fails to ask for enough.
-
Blocker availability drives the gap. In a controlled comparison on identical instances with the same model, providing the defect raised Macro-CAS from 13.6 to 80.6 and Sufficiency from 8.6 to 80.4, while No-Assumption stayed roughly flat (Δ = 67.0, 95% CI [59.0, 75.2], p < 0.001).
-
Taxonomy prediction adds difficulty but is not the whole story. A taxonomy-free ablation on 50 real-world instances showed GPT-5.6-Sol improving from 10.0% to 40.0% when scored only on identifying the same blocker without labels, meaning localization stays hard even when category prediction is removed.
-
Coarse defects are easier to catch. When an entire method component is missing or undefined, models recover it more readily than non-coarse defects, which are specific operations or local ambiguities inside an otherwise well-specified component.
-
Oracle resolution nearly eliminates the gap. With the gold defect, clarification action, and resolution supplied, the codification-ready rate rose from 14% to 98%, completeness from 30% to 98%, and missing-detail recovery from 14% to 98%, all statistically significant under paired testing. The unsupported-assumption rate fell from 6% to 0%, but with only three discordant instances this drop was not statistically significant (p_adj = 0.250) and should be read descriptively.
-
Passing tests does not mean implementing the intended method. In the executable study, oracle-resolved specifications raised test-passing implementations from 45% to 85% and faithful implementation of the target method detail from 30% to 90%, showing code can pass executable tests while encoding a different methodological choice.
-
The benchmark measures what it claims to. A construct validity study classified 94% of instances as specification-level clarification rather than idea-level refinement, and inter-annotator agreement was high (readiness κ = 0.84/0.92, Level-2 κ = 0.92/0.85 on real and synthetic subsets).
Methodology in Plain English
The authors assembled the benchmark from three sources. First, they mined GitHub issue threads from paper-linked repositories and reproducibility reports from venues like the ML Reproducibility Challenge, ECIR, and TMLR, keeping only cases where an implementer hit a genuine method-defining gap that was later resolved with evidence. Second, they took papers with verified successful reproductions and built clean "reference" specifications from the paper alone, then deliberately altered exactly one implementation-critical detail to create a precise counterfactual. Third, they used 43 executed research projects (19 human-authored ideas, 24 LLM-generated) whose final paper-code pairs documented the method as it was actually built, and removed one critical detail from each.
Every candidate went through two-stage human review, with a second annotator independently checking all retained and uncertain cases. Items that were not atomic, not evidence-supported, or not implementation-critical were discarded. Evaluated models never see the resolution evidence or downstream artifacts; they get only the reconstructed specification.
The three tasks mirror the decision sequence a real implementer faces: (1) is this specification ready to code? (2) if not, where exactly is the problem? (3) what specific action, either asking a question or inspecting an artifact, would resolve it? Task 3 is deliberately run in two conditions, one where the model receives the annotated defect ("Defect-Guided") and one where it must find it itself ("End-to-End"), to isolate how much difficulty comes from discovering the blocker versus formulating the fix.
Scoring used an LLM judge (Claude Opus 4.8) with rubric-based and semantic evaluation, validated against adjudicated human annotations. Because instances can share a source repository or paper, the authors also ran source-clustered bootstrap analysis, source-balanced estimates, and paired source-level comparisons to confirm findings were not artifacts of clustering.
Why This Matters
Impact on research. Automated research pipelines increasingly chain ideation, planning, coding, and experimentation. If the specification at the hinge between idea and code is underspecified, every downstream stage inherits the error. This benchmark makes that hinge measurable and shows that current models fail there even when they succeed at adjacent tasks. It also reframes reproducibility failures: some are not sloppy engineering but a specification that never pinned down a method-defining decision.
Real-world applications:
- Research agents and autonomous labs could use a codification-readiness gate that halts execution and requests clarification instead of silently implementing an assumed method.
- Reproducibility auditing for venues, journals, and artifact-review committees could flag submissions whose method descriptions underdetermine key choices before reviewers spend time on failed reproduction attempts.
- Paper-to-code assistants and scientific copilots could route ambiguous method details into targeted questions to authors rather than generating plausible but incorrect implementations.
- Peer review and editorial workflows could use readiness checks to identify manuscripts that need specification before publication.
Industry relevance. The same failure mode appears wherever requirements are handed off, including software specification, ML experiment configuration, data pipeline design, and API integration. The paper's distinction between "capable of acting on a known blocker" and "capable of discovering the blocker" is directly useful for teams building LLM-driven engineering tools: it identifies where to invest, namely blocker detection and clarification elicitation rather than response generation.
Future Directions
-
Multi-turn, multi-gap clarification. The benchmark deliberately isolates one defect per instance. The authors' own exploratory multi-defect ablation found no evidence that two combined defects are harder to localize, but it used only 50 pairs with unequal response budgets, so a properly matched study of interacting defects remains open.
-
Choosing between asking a person and inspecting an artifact. The clarification action type distinguishes questions from evidence-seeking, but current evaluation scores a single action rather than a strategy that decides which route is more reliable for a given gap.
-
From diagnostic to deployed gate. The work evaluates readiness before execution rather than running full implementations of every specification. Embedding readiness checks into research agents would allow direct measurement of whether clarification reduces implementation divergence, unsupported assumptions, and human correction cost.
-
Domain and scope expansion. Instances are concentrated in AI, NLP, and machine learning, with most source papers from 2020–2025. Extending collection to other computational sciences and to naturally occurring early-stage ideas would test whether the taxonomy and the readiness construct transfer.
-
Reconciling conflicting evidence. Real resolutions are distributed across papers, codebases, issue threads, and reports; the authors note that future systems need to handle conflicting evidence and determine when enough information has been gathered.
Target Audience
This paper is most valuable to researchers and engineers building LLM-based research or coding agents, and to anyone working on specification quality, requirement elicitation, or agent reliability. It is also relevant to reproducibility researchers and venue organizers who design artifact evaluation processes, to ML practitioners who have struggled to implement a paper from its description alone, and to NLP benchmark designers interested in how to construct controlled counterfactuals alongside naturally occurring real-world instances. Readers need only a working understanding of LLM evaluation, not deep expertise in any single scientific domain.
Authors’ abstract
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.