Skip to content
AI.info

Research

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

Overview Research area: Natural Language Processing, specifically application-level evaluation methodology for Large Language Model (LLM) systems, prompt engineering, and regression testing. Technical

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications
arXiv
2601.22025
Published
2026-01-29
Authors
Daniel Commey

AI summary

Overview

Research area: Natural Language Processing, specifically application-level evaluation methodology for Large Language Model (LLM) systems, prompt engineering, and regression testing.

Technical level: Intermediate. The paper is a technical report aimed at practitioners who already understand basic LLM application architectures (prompting, retrieval-augmented generation, tool use) but want a disciplined approach to evaluation.

Scope in one sentence: The paper proposes the Minimum Viable Evaluation Suite (MVES), an audit-oriented structure for application-level LLM evaluation, and reports a reproducible local prompt-ablation study showing that generic prompt "improvements" can improve one task contract while degrading another.

What This Paper Is About

Evaluating LLM applications is not like traditional software testing, because outputs are probabilistic, semantically variable, and sensitive to both prompt and model changes. The paper's core problem is that teams often treat prompt edits as assumed improvements, when they are in fact empirical interventions that can introduce regressions. The goal is to provide an audit-oriented evaluation structure plus a small, reproducible harness that surfaces these trade-offs before deployment.

Key Contributions

  1. MVES framework. An audit-oriented structure that links LLM application categories to failure modes, metrics, required artifacts, and validation evidence, across general LLM applications, retrieval-augmented systems, and agentic workflows. MVES has three tiers: MVES-Core (all apps), MVES-RAG (retrieval-based), and MVES-Agentic (tool-use).
  2. A reproducible local evaluation harness. Covering structured extraction, RAG citation/content-compliance, and instruction-following checks, with hand-authored seed suites and deterministic expansion to 30 cases per task in the reported ablation.
  3. A prompt-ablation study. Five prompt conditions tested across two local models, showing that generic prompt additions can improve one task contract while degrading another, with effects varying by model and by prompt placement.
  4. Practical evaluation guidance. Consolidated advice for test-set design, metric selection, LLM-as-judge calibration, and production monitoring, intended to support reproducible application-level evaluation.

Main Findings

  • Prompt effects are not monotonic. In the expanded 30-case-per-suite runs, stronger output-contract prompts improved strict extraction for both models, but RAG citation/content-compliance declined under some generic-rule conditions. The result is not that generic prompts are bad, but that prompt changes are empirical interventions requiring regression testing.

  • The largest observed decline: Qwen 2.5 on RAG. When generic rules were appended to the user prompt (condition C, "+Rules"), Qwen 2.5 RAG all-pass fell from 26/30 (86.7%, Wilson 95% CI [70.3, 94.7]) at baseline to 9/30 (30.0% [16.7, 47.9]). Under the same condition, its check-pass rate fell from 90.7% to 51.2%.

  • Extraction gains were large, especially for Qwen 2.5. Qwen 2.5 extraction went from 4/30 (13.3%) at baseline to 30/30 (100.0% [88.6, 100.0]) under the full improved system prompt (D), and to 25/30 (83.3%) under the non-conflicting condition (E). Llama 3 extraction went from 0/30 (0.0%) at baseline to 28/30 (93.3% [78.7, 98.2]) under D and 29/30 (96.7% [83.3, 99.4]) under E. The paper reads this as evidence that minimal baselines are inadequate for strict raw-JSON contracts, not as evidence that any generic prompt is intrinsically superior.

  • RAG regressed for Llama 3 too, but less sharply. Llama 3 RAG moved from 19/30 (63.3%) at baseline to 16/30 (53.3%) under both the full improved (D) and non-conflicting (E) system prompts. Its +Rules condition (C) actually improved RAG to 22/30 (73.3%) — the opposite direction from Qwen 2.5.

  • The non-conflicting condition is informative but not uniformly protective. Condition E explicitly preserves JSON-only, exact-format, and source-only contracts. It improved or preserved some instruction and extraction checks (Qwen 2.5 instruction rose to 23/30, 76.7%; Llama 3 instruction rose to 24/30, 80.0%), but it still reduced Llama 3 RAG all-pass relative to baseline.

  • Instruction-following was the most stable task. For Llama 3, conditions A, B, and C all produced 20/30 (66.7%) on instruction following, with only D and E lifting it. Qwen 2.5 instruction actually dipped under +Rules to 17/30 (56.7%) before recovering to 23/30 under E.

  • Check-pass rates reveal partial failures. Because all-pass counts a case as failed if any single check fails, the paper also reports check-pass rates. For example, Llama 3 extraction check-pass rates were 0.0, 0.0, 11.8, 94.1, and 97.1 across conditions A through E, while Qwen 2.5 extraction check-pass rates were 23.5, 23.5, 76.5, 100.0, and 85.3.

  • Failure modes are task-specific. In extraction, the dominant baseline failure was raw JSON contract violation: outputs contained markdown fences, explanations, or malformed structures, caught by json_valid and required_keys. In RAG, failures were missing citation markers, missing expected source-derived terms, or failure to follow the unknown-answer policy, caught by must_cite and gold_contains. In instruction following, failures included wrong output counts, regex mismatches, and insufficient refusals.

  • Evaluation method trade-offs (approximate regimes from literature, not measurements from this experiment). The paper reports human evaluation as the 1.0 baseline correlation at $1000+ per 1k and Days of latency; LLM-as-judge at 0.70–0.85 correlation, $10–50, Hours; BERTScore at 0.40–0.60, $0.10, Mins; and exact match as N/A correlation at $0.01, Secs, with high regression sensitivity on specific tasks.

Methodology in Plain English

The paper is built around a four-phase evaluation loop: Define (state quality requirements in testable terms), Test (run a curated suite), Diagnose (categorize failures to find systematic patterns), and Fix (adjust prompts, retrieval, or model choice), then repeat. This differs from one-time benchmarking because it is continuous and prioritizes failure analysis over a single aggregate accuracy number.

MVES operationalizes this by requiring that every evaluation component be traceable along a chain: application class to failure mode, failure mode to tests and metrics, tests to versioned artifacts, and metrics to validation evidence. MVES-Core calls for a stratified golden set, explicit acceptance checks, versioned prompts and model identifiers, and a human-reviewed calibration sample, with 100 examples as a practical lower bound for early-stage systems. MVES-RAG adds gold source-document annotations, retrieval metrics such as Recall@k or MRR, citation presence and citation-support checks, and tests for both answerable and unanswerable questions, with a Recall@5 target near 0.8 offered as a heuristic starting point rather than a universal threshold. MVES-Agentic adds trajectory-level evaluation, per-tool success rates, sandboxed execution, and review of irreversible or high-stakes actions.

For the empirical demonstration, all runs were local: Ollama API version 0.18.2 on a Mac mini M4 with 16GB unified memory, macOS 26.5, arm64. Models were llama3:8b-instruct-q4_K_M and qwen2.5:7b-instruct-q4_K_M, with temperature=0 and num_predict=1024. Each expanded condition was run once per case, and failed model calls were recorded as failed cases with the error message preserved in the JSON output.

The expanded suites contain 30 cases per task. Each begins with hand-authored seed cases (Extraction: 20 seed; RAG: 15 seed; Instruction: 15 seed) and adds deterministic synthetic variants from version-controlled templates. Extraction covers contact information, invoice parsing, calendar events, support tickets, classification, numeric extraction, policies, meeting notes, and edge cases. RAG contains closed-context question-answering cases checked for citation markers and expected source-derived answer terms. Instruction covers format constraints, exact matching, refusal behavior, structured output, and pattern matching.

Five prompt conditions were compared: A baseline (task-specific prompt plus minimal system prompt), B baseline plus a short instruction-following wrapper, C baseline plus generic rules appended to the user prompt, D baseline plus the full improved system prompt, and E baseline plus general guidance that explicitly preserves JSON-only, exact-format, and source-only contracts. Two metrics are reported: all-pass rate (percentage of cases where every configured check passes) and check-pass rate (percentage of individual checks that pass), with Wilson 95% confidence intervals on the all-pass rates.

A note on scope: the supplied paper text is truncated at Section 6.5. Several later sections are referenced but their content is not included here — Section 7 (metrics bridging offline and online evaluation), Section 9 (LLM-as-judge), Section 11 (the Diagnose phase), and the continuation of the test-set design and monitoring guidance.

Why This Matters

Impact on research. The paper reframes prompt engineering as an empirical intervention subject to regression testing rather than an intuition-driven craft. It also argues that MVES differs from benchmark suites such as HELM and metric libraries or RAG evaluators such as RAGAS in scope: MVES specifies the minimum artifacts needed to audit an application-level evaluation decision, and can incorporate tools such as RAGAS or the lm-evaluation-harness rather than replacing them. It also differs from checklist-style guidance by requiring each evaluation component to be traceable to a failure mode, an artifact, and a validation signal.

Real-world applications:

  • Customer support chatbots and knowledge retrieval systems, where prompt drift and format drift matter and where the quality taxonomy weights helpfulness and harmlessness as high.
  • RAG knowledge bases, where the paper's strongest result applies: adding generic prompt rules collapsed Qwen 2.5 RAG citation/content-compliance from 86.7% to 30.0% in this local stack.
  • Structured extraction pipelines for API consumption, where strict raw-JSON contracts fail catastrophically under minimal baselines (Llama 3 at 0/30) but reach 93.3–96.7% with explicit output-contract prompts.
  • Medical Q&A and code generation, where the paper's quality dimension table rates correctness, helpfulness, harmlessness, and groundedness (medical) or correctness, helpfulness, and format adherence (code) as high importance.

Industry relevance. The failure pattern described — a prompt change that looks like a generic improvement silently degrading a different task contract — is a concrete argument for maintaining small, task-specific regression suites that run on every prompt or model change. The harness is deliberately small and local, which makes it suitable for teams that need reproducibility and privacy without hosted evaluation infrastructure.

Future Directions

  • Larger and more representative test suites. The expanded suites contain 30 cases per task, which the paper explicitly says is useful as a mechanism-oriented case study but insufficient for estimating production reliability. Deterministic template-derived augmentation expands schema and constraint coverage but does not substitute for production-distribution sampling.
  • Stronger baselines and broader prompt design space. The paper flags "prompt strawman risk": the baseline and generic prompts are simple by construction, and a stronger hand-tuned baseline might reduce or reverse some observed differences. The non-conflicting condition is offered as a partial mitigation, not an exhaustive exploration.
  • Stronger RAG faithfulness measurement. The current RAG checks are citation/content-compliance proxies for citation markers and expected answer terms. The paper states that claim-level faithfulness would require human annotation, entailment modeling, or a validated LLM-as-judge protocol, and that citation precision and citation recall are not established by the current suite.
  • Broader model, hardware, and stack coverage. Only two local 7–8B quantized models were tested. Larger open-weight models and hosted API models may show different sensitivity to the same prompt changes, and results may vary across quantization schemes, Ollama versions, operating systems, and hardware. The paper notes the expanded run uses one live generation per case and condition, so it should not be read as a variance estimate; repeated-run modes exist in the harness but broader determinism claims need separate validation.
  • Independent authorship and held-out production traces. Because the same author designed the tasks, prompts, and checks, the paper calls for independently authored prompts and held-out production traces as a stronger validation design.

Target Audience

The primary audience is practitioners who build and ship LLM applications: ML engineers, AI product engineers, and technical leads responsible for prompt iteration and deployment decisions. It is also useful for evaluation and QA specialists who need an auditable artifact structure rather than an ad-hoc checklist, and for researchers studying LLM evaluation methodology, prompt sensitivity, and regression testing. Readers seeking benchmark leaderboards or universal prompt-engineering rules will not find them here — the paper explicitly states that it is not intended to rank LLMs, establish universal prompt-engineering rules, or provide a comprehensive survey of evaluation tools.

Authors’ abstract

Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantically variable, and sensitive to prompt and model changes. This technical report proposes the Minimum Viable Evaluation Suite (MVES), an audit-oriented structure for application-level LLM evaluation. MVES links application categories to failure modes, metrics, required artifacts, and validation evidence across general LLM applications, retrieval-augmented systems, and agentic workflows. We pair the framework with a reproducible local evaluation harness covering structured extraction, RAG citation/content-compliance, and instruction-following checks. Using Ollama with Llama 3 8B Instruct and Qwen 2.5 7B Instruct, we evaluate five prompt conditions over expanded 30-case-per-suite ablations. The results show that, in the tested local conditions, generic prompt additions do not produce monotonic improvements: stronger output-contract prompts improve strict extraction for both models, while RAG citation/content-compliance declines under some generic-rule conditions. The largest observed decline occurs for Qwen 2.5 on RAG when generic rules are appended to the user prompt, from 26/30 to 9/30. These findings support evaluation-driven prompt iteration: prompt changes should be treated as potential regression risks and tested against task-specific suites before deployment. The accompanying repository contains the test suites, prompt variants, evaluation harness, raw result logs, and scripts needed to reproduce the reported local ablations.

Read the original paper