Research
RIFT: Reordered Instruction Following Testbed To Evaluate Instruction Following in Singular Multistep Prompt Structures
Overview Research area: Large language model instruction following, prompt-structure robustness, and procedural control. Technical level: Intermediate. The paper is mathematically light (graph travers

- arXiv
- 2601.18924
- Published
- 2026-01-26
- Authors
- Andrew Jaffe, Noah Reicin, Jinho D. Choi
AI summary
Overview
Research area: Large language model instruction following, prompt-structure robustness, and procedural control.
Technical level: Intermediate. The paper is mathematically light (graph traversal notation) and the experimental setup is easy to grasp, but it helps to have some familiarity with LLM prompting, context windows, and instruction-tuned vs. reasoning-tuned models.
Scope (one sentence): The paper introduces RIFT, a controlled benchmark that holds question content and output format constant while varying only the traversal order of a single multistep prompt, in order to measure how much LLMs depend on positional continuity when following instructions.
What This Paper Is About
Existing instruction-following benchmarks such as IFEval, HELM, and BIG-Bench mix together task difficulty, linguistic complexity, and prompt ordering, so it is impossible to tell whether a model failed because a task was hard or because the instructions were arranged in an unusual order. The authors build a testbed where the same rephrased Jeopardy! question–answer pairs are presented either in sequential order (linear prompts) or in a non-sequential order defined by explicit "jump" arrows (jumping prompts), so that any performance difference can be attributed to structure alone. The goal is to determine whether instruction following in current LLMs is a genuine reasoning skill or a learned sequential pattern.
Key Contributions
- RIFT testbed: A controlled framework that disentangles prompt topology from prompt content by using an identical set of rephrased Jeopardy! question–answer pairs across baseline, linear, and jumping conditions, with the same system prompt, user prompt template, and output format in every configuration.
- A formal definition of jumping prompts: Traversal is modeled as a prompt graph where edges follow a fixed jump distance k with gcd(n, k) = 1, guaranteeing a Hamiltonian traversal that visits all n questions exactly once while breaking local positional continuity.
- A structural adherence metric: A binary flag that identifies whether a model's response matches the correct answer to a different question in the prompt, letting the authors separate factually correct answers produced out of order from genuine factual errors.
- An empirical scale-and-architecture sweep: 10,000 evaluations across six open-source models spanning 4B to 120B parameters, reasoning-tuned and instruction-only variants, using 10,000 linear and 10,000 jumping prompts.
Main Findings
-
Large accuracy drop under non-sequential structure: Accuracy dropped by up to 72% under jumping conditions compared to baseline, across 10,000 evaluations spanning six state-of-the-art open-source LLMs.
-
Baseline vs. linear vs. jumping (mean accuracy, %): gpt-oss-20b: 73.26 baseline, 17.75 linear, 4.94 jumping (jumping median 0); gpt-oss-120b: 84.83, 55.14, 13.9 (median 2.94); Qwen3-4B-Thinking-2507: 67.07, 34.35, 8.29 (median 2.23); Qwen3-4B-Instruct-2507: 55.71, 27.21, 1.76 (median 0.66); Qwen3-30B-A3B-Thinking-2507: 80.48, 51.91, 9.25 (median 2.59); Qwen3-30B-A3B-Instruct-2507: 75.30, 40.00, 2.43 (median 1.02).
-
Structural discontinuity dominates task difficulty and model scale: Baseline accuracy exceeds 55% for all models and surpasses 80% for the largest systems, showing the underlying task is within model capabilities. The relative drop Δ = (A_L − A_J)/A_L × 100 surpasses 12 percentage points for all models and often exceeds 30 percentage points, and parameter scaling alone does not reduce sensitivity to structural discontinuity.
-
Reasoning supervision helps only modestly: Only models trained with explicit reasoning supervision show non-trivial robustness to jumping prompts, and reasoning variants outperform their non-reasoning counterparts at every size (e.g., Qwen3-4B-Th beats Qwen3-30B-It in the jumping condition). Even the strongest model (gpt-120) reaches only 13.9% mean jumping accuracy, with a median below 3%.
-
Nominal vs. effective context capacity: gpt-oss-120B has a reported 131,000-token context window and Qwen3 variants a reported 262,000-token window, yet substantial degradation in linear instruction-following accuracy appears in the 2,000–5,000 token range. Jumping accuracy collapses toward a near-zero floor by roughly 3,000–5,000 tokens, meaning effective instruction-following capacity is an order of magnitude smaller than advertised limits.
-
Errors are structural, not knowledge-based: Between 33% and 60% of incorrect responses across models come from answering the wrong question. The abstract reports that approximately 50% of failures stem from instruction-order violations and semantic drift. Linear failures are typically local slips (skipping a question, repeating one, or an off-by-one misalignment), while jumping failures involve global state loss, retracing previously visited questions, misreading traversal instructions as self-referential, and premature termination.
-
Correct jumping answers are front-loaded: Nearly all correct answers in the jumping condition occur within the first approximately 20 executed questions, indicating that models can handle a few discontinuous transitions but cannot sustain non-sequential control as traversal depth grows.
-
Reasoning failure modes differ in form, not in fragility: Non-reasoning models most often fail to follow jump instructions at all, whereas reasoning models are more likely to enter self-referential loops, revisit earlier steps, terminate early after incorrectly inferring completion, or decline to produce meaningful output (an example shows gpt-oss-20b responding "OUTPUT: Unable to comply.").
Methodology in Plain English
The researchers took 83,453 Jeopardy! question–answer pairs (24,907 unique categories, clue values $200–$2000) and used an LLM to rephrase the clues into standard question–answer form, since Jeopardy! answers are conventionally phrased as questions. They then cleaned the data: removing questions that needed superscript/subscript, that had commas in the answer (to keep comma-separated output parseable), or that contained non-ASCII characters after diacritic normalization, and replacing ampersands in questions with "and."
From this pool they built prompts ranging from 10 to 300 questions (average 155). In the baseline condition, each question is fed to the model individually. In the linear condition, questions are numbered and each one ends with a directive pointing to the next number. In the jumping condition, the content is identical but each question ends with a directive pointing to a non-adjacent question, following a fixed jump distance k chosen to be coprime with the number of questions n, so that every question is visited exactly once. Prompts are generated deterministically from a configuration specifying the random seed, dataset start index, number of questions, prompt type, and jump distance (an example configuration uses seed 36, num 54, start_row 3704, jumping true, max_jump_distance 18).
A fixed system prompt casts the model as a program executor that must follow the arrows strictly and output only a comma-separated list of answers in the order answered. Correctness is judged by an LLM evaluator rather than exact string matching, so that semantically equivalent surface forms count as correct; the evaluator was manually verified on a random subset of 100 items and achieved 98% agreement. A separate structural adherence flag records when a model gives the right answer to the wrong question.
Experiments ran on Python 3.13.11 with vLLM 0.11.2 and PyTorch 2.9.0, on NVIDIA H200 or H100 GPUs with CUDA 12.8, bfloat16 precision, and Flash Attention 3. Maximum token length per prompt answer was 81,920, and total GPU runtime was approximately 198 hours.
Why This Matters
Impact on research. The paper reframes instruction-following failures as a prompt-topology problem rather than a knowledge or reasoning-depth problem, and provides a reproducible testbed for measuring topology-invariant execution. It argues that strong linear-prompt performance may reflect pattern continuation or memorized positional dependencies rather than genuine comprehension of instructional structure.
Real-world applications:
- Workflow automation, where a model must execute steps in an order that is not the order they appear in the prompt.
- Multi-agent systems, where control must pass across distant parts of a shared context rather than strictly forward.
- Conversation management and decision support, where conditional or non-sequential control flow is routine.
- Tutoring and automated reasoning, where instruction reliability cascades into downstream reliability in high-stakes contexts.
Industry relevance. The paper emphasizes that open-source models dominate production environments where cost, customization, and data privacy rule out proprietary API dependencies, and that the gap between nominal context limits (100K+ tokens) and effective instruction-following capacity (approximately 3–5K tokens for non-sequential tasks) is a currently unmeasured deployment risk. Practitioners designing long, multi-step prompts should not assume that a large advertised context window means reliable execution at that length.
Future Directions
- Explicit state tracking and graph-based attention: The authors motivate architectural interventions, including state-tracking mechanisms and graph-based attention, along with training objectives that emphasize structural robustness over sequential pattern matching.
- Beyond binary topology: The current framework compares linear with non-sequential jumping traversal only, and does not include conditional loops, state-dependent repetition, or other control flow resembling procedural programs. Extending beyond the linear/non-linear binary is stated as future work.
- Broader domains: The evaluation uses rephrased Jeopardy!-style fact-based question answering, so generalization to mathematical reasoning, symbolic manipulation, planning, and interactive tool use remains untested, and structural effects may differ in domains with richer state representations.
- Removing compute constraints: The study was restricted to two H100 and two H200 GPUs, bounding the number of models, prompt lengths, and repeated trials; more compute would allow broader exploration of model scales, longer traversal sequences, and more extensive ablations. The authors also note that LLM-based evaluation, while verified, may retain errors in borderline cases, and that imperfect baseline accuracy shows some questions fall outside the models' effective knowledge.
Target Audience
This paper is most useful for LLM researchers and engineers working on instruction following, prompt engineering, long-context evaluation, and agentic or multi-step workflow systems, as well as benchmark designers who need a controlled way to separate prompt structure from task content. Practitioners deploying multi-step prompts in production will also benefit from the nominal-versus-effective context capacity finding and the structural adherence metric.
Authors’ abstract
Large Language Models (LLMs) are increasingly relied upon for complex workflows, yet their ability to maintain flow of instructions remains underexplored. Existing benchmarks conflate task complexity with structural ordering, making it difficult to isolate the impact of prompt topology on performance. We introduce RIFT, Reordered Instruction Following Testbed, to assess instruction following by disentangling structure from content. Using rephrased Jeopardy! question-answer pairs, we test LLMs across two prompt structures: linear prompts, which progress sequentially, and jumping prompts, which preserve identical content but require non-sequential traversal. Across 10,000 evaluations spanning six state-of-the-art open-source LLMs, accuracy dropped by up to 72% under jumping conditions (compared to baseline), revealing a strong dependence on positional continuity. Error analysis shows that approximately 50% of failures stem from instruction-order violations and semantic drift, indicating that current architectures internalize instruction following as a sequential pattern rather than a reasoning skill. These results reveal structural sensitivity as a fundamental limitation in current architectures, with direct implications for applications requiring non-sequential control flow such as workflow automation and multi-agent systems.