Research
What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation
What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation Overview Research area: Natural Language Processing — evaluation

- arXiv
- 2609.03254
- Published
- 2026-09-03
- Authors
- Daisuke Kikuta
AI summary
What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through ConversationOverview
Research area: Natural Language Processing — evaluation of Large Language Models on multi-turn, conversationally generated structured artifacts, combined with test-time compute methods.
Technical level: Intermediate. The reader benefits from familiarity with LLM inference, sampling-based decoding strategies, and JSON patch formats, though the paper explains its setup clearly.
Scope: This paper introduces RevPropBench, a human-annotated benchmark for measuring whether LLMs can propagate a local revision request to all dependent elements of a JSON artifact that was built through conversation, and it evaluates nine revision methods across six LLMs to find the most cost-effective way to apply test-time compute.
What This Paper Is About
When a user asks an LLM to revise an artifact (such as a plan, invoice, or configuration) in conversation, the user typically names only one element to change, but other elements depend on it and must be updated too. The paper studies this revision propagation problem in a setting where dependencies are not explicit or statically analyzable — they may be established implicitly by the conversation that produced the artifact. The goal is to build a benchmark for this setting and identify which test-time compute strategies improve accuracy most cost-effectively.
Key Contributions
- A new benchmark, RevPropBench, for evaluating the ability of LLMs to propagate revisions across dependent elements in conversationally generated JSON artifacts. It contains 150 samples: 30 development samples and 120 test samples, spanning nine domains, six propagation patterns, and three artifact sizes of 10, 50, and 100 JSON elements.
- A comprehensive evaluation of nine revision methods — three single-inference baselines, sequential reflection, four rule-based parallel-sampling merging variants, and LLM-based selection — across six LLMs: gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b-a10b.
- Practical guidance for cost-effective method selection, including cost and latency analyses and the recommendation of specific methods at specific numbers of LLM calls.
- Release of the benchmark instances, the data sampling tool, and the annotation tool for reproducibility and future extension (available at https://github.com/ntt-dkiku/llm-revision-propagation).
Main Findings
- Conversation history matters more than the artifact alone for context, but both together are best. Across every model, the single-inference baselines follow a consistent ordering, j < h < j+h (for example, 90.7% < 92.7% < 93.0% for gpt-5.4-mini and 81.7% < 86.7% < 90.3% for qwen3.5-122b). The paper attributes the gain from j to h to dependency information contained in the conversation that cannot be recovered from the final artifact alone, and the further gain from h to j+h to helping models avoid path errors and missed revisions when emitting JSON patches.
- Baseline accuracy spans 68.3–93.0%. Using the strongest baseline (j+h), completion rates span 68.3–93.0% across models. Performance generally increases with model scale: qwen3.5-9b < 27b < 122b, and gpt-oss-20b < 120b < gpt-5.4-mini.
- LLM-based selection from parallel samples is the most consistent improvement. Select improves over j+h by +3.3–12.5%, achieving the highest completion rate on four of the six models and the second-highest on the remaining two.
- Medoid selection is the second-most consistent. med improves by +1.8–7.7%, achieving the highest completion rate on one model and the second-highest on three models.
- Unanimous merging is harmful. The "and" rule falls below j+h on every model (−13.2–21.3%) because requiring unanimous agreement discards correct edits found by only some samples. The "or" and "maj" rules are competitive with Select and med on a few models but are not consistent across models (−0.8–+4.8%).
- Sequential reflection gives limited gains. Reflect improves by +0.5–8.2% but rarely matches Select or med.
- Errors are dominated by under-propagation. Grouping failures into miss, over edit, and wrong value (pooled over the six models, i.e., 3600 samples), miss accounts for the majority of errors for most methods. The "and" rule fails almost only by miss, while "or" introduces many over edits. med and Select produce the fewest failures: med reduces outlier over edits but recovers fewer misses, whereas Select's reasoning tends to target the most complete candidate, reducing misses the most but not over edits.
- The most cost-effective setting is three parallel samples. Using LLM-based or medoid selection over three parallel samples is the most cost-effective approach, improving accuracy by 2.2–9.7% compared to a single inference.
- Cost differs by model family. At five LLM calls, GPT models consume similar API cost across all methods, increasing roughly in proportion to the number of calls. For Qwen models, Select increases cost by 5.7–7.5× over j+h, mainly due to more reasoning tokens.
- Select's advantage is not merely from higher cost. When other methods are extended until their costs match or exceed Select's, the relative performance ordering among methods remains largely unchanged.
- Performance saturates around four to five LLM calls. In post-hoc analysis, the highest cost-effectiveness on average across models is achieved by either Select with three or four LLM calls, or med with three LLM calls.
- Latency favors medoid selection. med executes its samples in parallel, so its latency stays below 1.5× that of j+h even as the number of samples increases, while Select remains around 2–3× for GPT models but rises to 3.3–6× for Qwen models. On performance gain relative to latency increase, med with k=3 is the most latency-efficient method on average across all models.
- Recommended configuration. The paper recommends Select with four LLM calls as a practical choice, and med with three LLM calls as an alternative when latency is also considered.
Methodology in Plain English
The researchers built a benchmark rather than only running an existing one. They first created 50 scenarios covering nine practical domains and six propagation patterns (arithmetic, substitution, add_remove, threshold, temporal, and status_flip). A strong LLM (GPT-5.5) then generated synthetic user–LLM conversation histories from those scenarios, in which an artifact is grown element by element through JSON patches, ending with a final JSON artifact. Each conversation was paired with a local revision request whose correct answer requires propagating the change to dependent elements.
Gold patches were produced in two steps: another strong LLM (Claude-Opus-4.8) generated tentative patches, and a human annotator reviewed and corrected them through a GUI annotation tool, also consulting the LLM about its rationale. The gold patch is an ordered sequence of RFC 6902 operations, each a triple of (op, path, value), where op is replace, add, or remove. Free-form values are checked with keyword matchers rather than exact string matching, and elements that are valid to either edit or leave unchanged can be flagged optional.
Evaluation mirrors this: a model receives the conversation, the final artifact, and the revision request, and must output JSON patches. The primary metric is the completion rate — the proportion of samples where the patched artifact exactly matches the gold-patched artifact. Three finer metrics support failure analysis: miss, over edit, and wrong value.
The 150 samples were split at the scenario level so all three size variants of a scenario stay together. Ten of the 50 scenarios (30 samples) were assigned to development using random sampling with seed 42, balanced by domain, propagation size, and propagation pattern; the remaining 120 samples form the test split.
Nine methods were compared: three baselines that generate a patch in one inference with the final artifact (j), the conversation history (h), or both (j+h); sequential reflection (Reflect) starting from the j+h patch; four rule-based merges of parallel samples (or, and, maj, med — where med selects the candidate with the smallest mean leaf-level disagreement with the others); and Select, which samples in parallel and then uses the same LLM to choose among candidates. For the merge rules, candidates are compared after being applied to the artifact rather than as patches, because the same edit can be written as many different patches.
Reasoning was enabled for all six models, with effort level set to medium for the GPT models. Temperature was 0.6 for all models that support it, except gpt-5.4-mini, and the maximum output context was 32K tokens. Local models were deployed with vLLM on four NVIDIA A100 (80GB) GPUs. Each sample was evaluated five times with seeds s = 0, 42, 84, 126, 168; Reflect performs four reflection iterations (five LLM calls total), and parallel sampling draws five outputs with seeds s through s+4. Select samples four outputs and uses seed s for the additional selection call.
Why This Matters
Impact on research. Existing revision-propagation work assumes dependencies are explicit or statically analyzable — call graphs and imports in repository-level code editing, pre-existing knowledge graphs in knowledge editing, and section, figure, citation, or claim references in document editing. This paper defines and measures a distinct setting where such dependency sources are unavailable and dependencies may be established implicitly by the conversation outside the artifact. It also extends test-time compute research into that setting and shows which strategies transfer and which do not.
Real-world applications:
- Travel itinerary planning, where moving an anchor date shifts every activity scheduled relative to it.
- Invoice processing, where changing a line's quantity or unit price recomputes line amounts, the subtotal, and every charge built on the subtotal.
- Project scheduling, where removing a task relinks successors and shifts downstream dates and project totals.
- Data pipeline and deployment configuration, where crossing a stated threshold changes partition strategy or retention tiers.
Industry relevance. LLM systems routinely produce JSON as structured output for tool calls, agent actions, and system-side data processing, so a benchmark built on JSON artifacts and RFC 6902 patches is directly relevant to deployed systems. The cost and latency analyses — including the finding that performance saturates around four to five LLM calls — give practitioners concrete guidance on how much inference budget to spend. The paper also notes that in domains such as financial data processing and invoice processing, a single over-edit or miss can cause significant errors, and suggests human review safeguards such as presenting all propagated changes for approval before applying them.
Future Directions
- Continuously raise the benchmark's difficulty. gpt-5.4-mini already reaches a 93% completion rate, raising the concern that comparisons will saturate as models improve; the authors provide the full construction tool so the benchmark can be updated over time.
- Move closer to real-world conversations. The conversations are synthetically generated from human-created scenarios rather than drawn from actual human–LLM interaction data, and dependencies were deliberately kept deterministic for reliable automatic evaluation. Ambiguous cases where people would disagree about whether an element should change are not covered.
- Validate safeguards for critical revisions. The proposed approaches — presenting all propagated changes for user approval, using an LLM-as-a-judge to prioritize edits by criticality, or letting users pre-specify critical items — are described by the authors as conceptual ideas not yet experimentally validated.
- Broaden coverage beyond the current patterns. The benchmark's 150 samples cover nine domains, six propagation patterns, and three artifact sizes, but the authors state it does not exhaustively cover all possible patterns. There is also a noted bias concern: conversations were generated using GPT-5.5, so samples may be easier for GPT models, although the authors argue scenarios were created by humans with Claude-Opus-4.8 and the conversations were built from meta-level instructions to limit task overlap between generation and evaluation.
Target Audience
Researchers and engineers working on LLM-based code and document editing, agentic systems that emit structured JSON output, and test-time compute or inference-budget optimization. It is also relevant to practitioners who must decide how much inference cost and latency to spend on reliability in conversational artifact editing, and to benchmark builders interested in the synthetic sampling plus human annotation workflow described here.
Authors’ abstract
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies and propagate the revision to all affected parts of the artifact. This paper studies this ability of LLMs on conversationally generated artifacts, where the artifact context and its dependencies may be embedded in the conversation history. Toward practical use, we also explore cost-effective test-time compute for this new setting. Specifically, we introduce a new benchmark for this setting, and evaluate nine revision methods, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark. The results show that baselines achieve accuracies of 68.3--93%, and the most cost-effective method is selecting from three parallel samples using either LLM-based or medoid selection, which improves accuracy by 2.2--9.7%. Our code and dataset are available at https://github.com/ntt-dkiku/llm-revision-propagation.