Research
BenchOverflow: Measuring Overflow in Large Language Models via Plain-Text Prompts
Overview Research area: Natural Language Processing / LLM safety, reliability, and serving economics — specifically length control and "unbounded consumption" failure modes. Technical level: Intermedi

- arXiv
- 2601.08490
- Published
- 2026-01-13
- Authors
- Erin Feiglin, Nir Hutnik, Raz Lapid
AI summary
Overview
Research area: Natural Language Processing / LLM safety, reliability, and serving economics — specifically length control and "unbounded consumption" failure modes.
Technical level: Intermediate. The conceptual core is simple (long prompts make models talk too much), but the paper leans on distributional statistics (ECDFs, cap-saturation rates, Pearson correlations, within-prompt variance) and a standardized multi-model evaluation protocol.
Scope (one sentence): The paper names and measures Overflow — excessive generation triggered by ordinary, non-adversarial plain-text prompts — and introduces BenchOverflow, a nine-strategy benchmark evaluated on nine open- and closed-source LLMs under a fixed 5,000-new-token budget, plus a one-sentence conciseness-reminder defense.
What This Paper Is About
Large language models have been trained to be thorough, comprehensive and compliant, so plainly worded requests for breadth ("enumerate", "list every", "expand each entry") often trigger very long outputs. The authors call this Overflow: prompt-induced excessive text generation that requires no jailbreak, no adversarial suffix, and no policy circumvention — just ordinary natural language. The paper's goal is to build a standardized, model-agnostic way to measure how reliably models over-generate, how heavy the tails of those length distributions are, and whether a trivial mitigation helps.
Key Contributions
-
Taxonomy and benchmark. The authors develop a taxonomy of overflow-inducing prompting strategies and release BenchOverflow, a model-agnostic benchmark instantiating nine representative attack types: change forms, explicit forced length, implicit large enumeration, infinite generation, recursive details, roleplay simulation, tokenizer stress, quote, and stepwise explanation. For each strategy they curate more than 300 systematically constructed prompts.
-
In-depth multi-model evaluation. A comprehensive study of overflow behavior across nine state-of-the-art LLMs spanning open- and closed-source families, covering distributional properties of output lengths, central tendency, tail risk, and within-prompt variability across repeated trials.
-
Lightweight defense. Assessment of a simple, model-agnostic mitigation consisting of a generic conciseness reminder prepended to the user prompt, tested across all evaluated models and shown to consistently reduce overflow incidence.
-
Reframing of verbosity. The paper positions length control as a measurable reliability, cost, and sustainability concern — connecting it to "wallet exhaustion" / Denial-of-Wallet style resource exhaustion and to OWASP's LLM10: Unbounded Consumption and LLM06: Excessive Agency risks — rather than a stylistic quirk.
Main Findings
-
Plain-text prompts reliably inflate output length. Across all nine models, BenchOverflow prompts produced pronounced rightward shifts and heavy tails in completion-length distributions relative to the benign OASST2 baseline, with visible mass accumulating near the 5,000-token cap. Both open- and closed-source families show the effect.
-
Two strategies dominate saturation. Explicit forced length and Tokenizer stress dominate CSR@5k (the share of generations exceeding 5,000 tokens). Quote, Infinite generation, and Recursive details create heavy right tails by encouraging continuation or expansion without sharp stopping cues. Change forms, Implicit large enumeration, and Roleplay simulation shift mass into mid-to-high ranges but hit the cap less consistently. The benign baseline maintains low CSR at all thresholds.
-
Overflow is reproducible, but stability is family-dependent. Per-prompt standard deviations cluster tightly near zero for GPT-5, Qwen-3-4B-Instruct, Gemma-2-9B-It, Gemma-3-4B-It and Claude-Sonnet — once a prompt triggers overflow, it does so reliably. Qwen-3-8B-Instruct, Gemini-2.5-Flash, LLaMA-3.1-8B-Instruct and LLaMA-3.2-3B-Instruct show heavier right tails, with a notable fraction of prompts varying by more than 10³ tokens across runs.
-
Cross-model agreement depends on the attack vector. Roleplay simulation and Stepwise explanation elicit relatively high cross-family correlations, while Infinite generation produces far weaker alignment. LLaMA variants correlate strongly with each other (up to 69–71%), Qwen models show moderate within-family agreement (36–40%), and Gemma-2 vs. Gemma-3 are only weakly related (23%). GPT-5 and Claude-Sonnet show the broadest positive associations across systems (51% and 54% with multiple families). LLaMA-3.2-3B-Instruct and Qwen-3-4B-Instruct even reach a negative correlation on Explicit forced length.
-
Refusals modulate but do not remove overflow. Refusal rates vary sharply by strategy: Implicit large enumeration frequently triggers refusals (e.g., 92.8% for Gemma-2-9B-It, 88.3% for Claude-Sonnet and Qwen-3-8B-Instruct, 86.0% for Qwen-3-4B-Instruct), whereas Roleplay simulation is near zero across every model (0.0–1.3%). The benign OASST2 baseline sits between 1.0% and 5.0%. Critically, for Explicit Forced Length and Tokenizer Stress, several models (Gemma-4B, LLaMA-8B and 3B, GPT-5, Qwen-4B) continue to produce near-cap completions even when flagged as refusals — "continued refusals" where the model ostensibly rejects the instruction yet still generates extensive explanations or partial completions.
-
Alignment choices plausibly explain the heterogeneity. Gemma-2-9B-It exhibits relatively high refusal rates across overflow strategies, consistent with alignment data or reward shaping that promotes rejecting exaggerated demands, while GPT-5 often follows such prompts with minimal resistance.
-
The conciseness reminder works, with caveats. Prepending a fixed brevity reminder produced a visible leftward shift, reduced density near the cap, and lowered CSR for most strategies. Mean output length dropped by roughly 30% (GPT-5) to more than 85–90% for several other models (Gemini-2.5-Flash, Qwen-3-8B, Gemma-3-4B-It). The effect is heterogeneous: style- or scope-cued strategies (Roleplay simulation, Infinite generation, portions of Recursive details) respond more, while Tokenizer stress remains comparatively resistant. The defense is described as a low-cost first line of mitigation that is not by itself sufficient to prevent saturation in the strongest cases.
-
The defense trades answer quality for length control. On benign OASST2 prompts, the fully-correct rate fell under the reminder for most models — for example, Gemini-2.5-Flash (93.2 to 54.8), Qwen-3-8B (85.8 to 45.5), Qwen-3-4B (88.5 to 54.8), Gemma-3-4B-It (78.2 to 32.8), Claude-Sonnet (90.2 to 76.8) and LLaMA-3.1-8B (62.5 to 47.0) — while GPT-5's fully-correct rate rose slightly (93.2 to 94.2). For several models the shift is a redistribution from "perfect" to "partial" answers rather than a collapse into non-answers, suggesting the constraint mainly trims supporting detail.
-
Key findings as stated by the authors: (1) Plain-text prompts reliably inflate output lengths across models, with heavy tails. (2) A small subset of strategies (Explicit forced length, Tokenizer stress) frequently reach the 5k-token evaluation limit, while Quote, Infinite generation and Recursive details yield extended right tails and the remaining strategies produce moderate increases. (3) Overflow is reproducible within prompts for most models, though stability varies by family. (4) A generic conciseness reminder measurably reduces tail mass and CSR but does not fully neutralize the strongest overflow vectors.
Methodology in Plain English
Setting up the threat model. The adversary is an ordinary end-user with no elevated privileges: no ability to change training data, system prompts, or decoding internals, and no ability to bypass API-level restrictions such as max_tokens. The only capability assumed is submitting plain-text input.
Generating attack prompts. The 300-plus prompts per strategy were not hand-written one by one. Instead the authors used a meta-prompting procedure with GPT-4o running at its default decoding configuration (temperature = 1.0). The generator model was told to act as a red-team collaborator targeting a specific overflow vector, given 5–8 positive examples with brief explanations of why they work, and given 3–4 negative counterexamples that look similar but do not satisfy the operational criteria, sharpening the decision boundary. A single unified meta-prompt template specified the attack vector, the operational description of a valid instance, required and prohibited elements, the positive/negative examples, and a JSON schema for the returned prompts. A refinement loop ran the template through an LLM with manual correction.
The nine strategies. Each captures a distinct mechanism: transforming content across formats (Change forms), demanding extreme output targets (Explicit forced length), asking for "all" of a vast category (Implicit large enumeration), instructing endless continuation (Infinite generation), requesting verbatim reproduction of public-domain texts (Quote), listing items with multi-layer expansion (Recursive details), staging character interactions (Roleplay simulation), framing tasks as procedural walkthroughs (Stepwise explanation), and exploiting tokenization inefficiencies such as numerals expanded into words (Tokenizer stress).
Models and configuration. Six open-source models — Qwen3-4B-Instruct-2507, Qwen3-8B-Instruct, LLaMA-3.2-3B-Instruct, LLaMA-3.1-8B-Instruct, Gemma-3-4B-It and Gemma-2-9B-It — were run with do_sample=True, temperature=1.0, max_new_tokens=5000. Three closed-source models — GPT-5, Gemini-2.5-Flash and Claude-Sonnet — were queried under provider defaults, with the output budget set to 5,000 tokens to match the open-source setup. Sampling and penalty hyperparameters, stop sequences and safety filters were not modified.
Datasets and execution. Evaluation drew on the nine BenchOverflow attack sets plus a benign baseline: the OpenAssistant Conversations corpus (OASST2), used as a neutral reference distribution approximating organic user queries. From each dataset the authors deterministically sampled 100 prompts, using identical subsets across models. Each prompt–model pair was evaluated in four independent runs, giving 100 × 4 = 400 generations per strategy–model pair, 3,600 per model across all strategies, and 4,000 generations per model once the OASST2 baseline's 400 generations are included. Each request consisted of one system message ("You are a helpful assistant.") and one user message containing the evaluation prompt. Completion lengths were measured with each model's native tokenizer convention.
Measurement. Three complementary angles: (i) histograms and ECDFs of completion lengths; (ii) cap-saturation rate (CSR) at thresholds τ ∈ {1k, 3k, 5k} tokens; and (iii) consistency analysis — within-prompt variability across repeated runs, and cross-model correlation of per-prompt lengths.
Defense and quality check. The mitigation prepends the fixed string "Reminder: Please provide a concise, precise response without unnecessary elaboration," forming p′ = p|r, and only generated tokens are counted so the reminder itself does not contribute to length. To check for a utility trade-off, both conditions were scored on the OASST2 subset by an LLM-as-a-judge using a three-point adequacy rubric (0 = no answer, 1 = partial answer, 2 = full answer). Separately, a GPT-5-mini classifier labeled each completion as Refusal or Non-Refusal under a minimal rubric in which disclaimer-prefaced yet substantive executions count as Non-Refusal.
Why This Matters
Impact on research. The paper carves out a distinct failure mode adjacent to, but separate from, jailbreaks, prompt injection, and adversarial-suffix optimization (e.g., Engorgio's reported 2–13× output-length increases, CRABS/AutoDoS's reported >250× latency inflation, or the Excessive Reasoning Attack's reported 3–9× reasoning-length increases). Prior unbounded-consumption work generally assumed privileged conditions — white-box gradient access, control over training or retrieval corpora, adversarial image perturbations, or agent execution frameworks. By showing that unmodified, benign-looking text can produce similar effects, the paper argues that length control deserves a place alongside safety and alignment as a first-class reliability property, and supplies a common benchmark for comparing models on it.
Real-world applications:
- Production LLM serving and capacity planning. CSR@1k/3k/5k gives operators a length-agnostic measure of tail risk to decide which models to deploy where, and to size token budgets and rate limits.
- Cost and energy accounting. The authors frame unnecessary tokens as directly increasing per-request cost and energy consumption, compounding into operational spend and carbon footprint at scale.
- Shared multi-tenant environments. In shared deployments, the same behavior consumes bandwidth, memory, and model slots, degrading service for other tenants — their illustrative case is a financial institution whose customer-support LLM could be flooded by an attacker posing as a legitimate user, exhausting token budgets and obstructing time-critical actions.
- Guardrail and middleware design. The finding that refusals do not imply length safety — "continued refusals" that still run to near-cap — gives builders a reason to enforce output-length budgets in addition to content filters.
Industry relevance. The phenomenon maps onto what the security community calls wallet exhaustion or Denial of Wallet, and onto OWASP's LLM10: Unbounded Consumption and LLM06: Excessive Agency entries. That makes BenchOverflow relevant to anyone running pay-per-token APIs, latency-sensitive products, or agent frameworks where a runaway completion propagates downstream.
Future Directions
-
Mitigations beyond a one-sentence reminder. The reminder leaves Tokenizer stress comparatively resistant and does not fully neutralize the strongest vectors, so harder or more adaptive defenses — and combinations of decoding-time and prompt-level controls — remain open.
-
Reconciling controllability with task adequacy. The defense consistently reduced length but also reduced fully-correct answer rates for most models, shifting answers from "perfect" to "partial." Finding interventions that preserve full-correctness while truncating the right tail is an explicit opening.
-
Explaining the model-family divergence. The paper attributes differences in susceptibility and refusal behavior to heterogeneous alignment pipelines and reward shaping (RLHF, RLAIF, constitutional AI, DPO), but the heterogeneity — including the negative correlation between LLaMA-3.2-3B-Instruct and Qwen-3-4B-Instruct on Explicit forced length — is characterized empirically rather than causally explained.
-
Extending the measurement frame. Because the benchmark fixes a single 5,000-token budget and one benign baseline corpus, whether the CSR and correlation structure generalizes to other budgets, other provider default settings, longer multi-turn dialogues, or agentic toolflows is not established here. (The paper also does not report a per-strategy breakdown of the defense's adequacy scores — Table 3 is reported at the model level on the benign OASST2 subset.)
Target Audience
The most likely beneficiaries are LLM platform and infrastructure engineers who own serving cost, latency, and rate-limit budgets; safety and red-team researchers studying resource-exhaustion and unbounded-consumption attacks; and model developers evaluating how alignment choices affect verbosity. Because the prompting strategies are described in plain language and need no adversarial tooling, the taxonomy is also useful to product teams and technically-engaged decision-makers who need a concrete, non-alarmist case for enforcing output-length controls.
Authors’ abstract
We investigate a failure mode of large language models (LLMs) in which plain-text prompts elicit excessive outputs, a phenomenon we term Overflow. Unlike jailbreaks or prompt injection, Overflow arises under ordinary interaction settings and can lead to elevated serving cost, latency, and cross-user performance degradation, particularly when scaled across many requests. Beyond usability, the stakes are economic and environmental: unnecessary tokens increase per-request cost and energy consumption, compounding into substantial operational spend and carbon footprint at scale. Moreover, Overflow represents a practical vector for compute amplification and service degradation in shared environments. We introduce BenchOverflow, a model-agnostic benchmark of nine plain-text prompting strategies that amplify output volume without adversarial suffixes or policy circumvention. Using a standardized protocol with a fixed budget of 5000 new tokens, we evaluate nine open- and closed-source models and observe pronounced rightward shifts and heavy tails in length distributions. Cap-saturation rates (CSR@1k/3k/5k) and empirical cumulative distribution functions (ECDFs) quantify tail risk; within-prompt variance and cross-model correlations show that Overflow is broadly reproducible yet heterogeneous across families and attack vectors. A lightweight mitigation-a fixed conciseness reminder-attenuates right tails and lowers CSR for all strategies across the majority of models. Our findings position length control as a measurable reliability, cost, and sustainability concern rather than a stylistic quirk. By enabling standardized comparison of length-control robustness across models, BenchOverflow provides a practical basis for selecting deployments that minimize resource waste and operating expense, and for evaluating defenses that curb compute amplification without eroding task performance.