Research
Replacing Training with Memory: Listwise Selection for Text-to-SQL
Overview Research area: Text-to-SQL generation, specifically the candidate-selection (reranking) stage of generate–execute–select pipelines, combined with retrieval-based memory and inference-time bia

- arXiv
- 2609.00834
- Published
- 2026-09-01
- Authors
- Yeonseok Jeong, Soyoung Yoon, Seongjun Lee, Seung-won Hwang
AI summary
Overview
- Research area: Text-to-SQL generation, specifically the candidate-selection (reranking) stage of generate–execute–select pipelines, combined with retrieval-based memory and inference-time bias mitigation. Posted under cs.SE.
- Technical level: Intermediate. The paper assumes familiarity with large language model pipelines, fine-tuning objectives, listwise reranking, and the "lost-in-the-middle" positional bias problem.
- Scope: The paper introduces MaP-SQL, a fine-tuning-free listwise selector for Text-to-SQL that replaces training with retrieved structured memories and group-based permutation aggregation, evaluated on BIRD-dev, Spider-test, and EHRSQL.
What This Paper Is About
Modern Text-to-SQL systems often generate several candidate SQL queries and then use a selector to pick the best one. Listwise selectors, which compare multiple candidates jointly, work well but are expensive to fine-tune because each training example bundles many queries plus their execution results into long contexts. This paper asks whether the two main jobs of fine-tuning — learning selection criteria and mitigating positional bias — can instead be handled at inference time, with no selector parameter updates.
Key Contributions
-
Reusable structured memories as explicit selection criteria. Instead of learning selection behavior in model weights, MaP-SQL builds memories from training data that encode how natural language maps to schema elements, SQL operations, and expected outputs, organized into Encoding, Translating, and Decoding groups. These are retrieved at test time by question similarity and prepended to the selector prompt as decision criteria.
-
Group-based permutation for positional-bias mitigation. Rather than shuffling all candidates globally or exploring up to O(n!) permutations, MaP-SQL groups candidates by execution result, fixes the ordering between groups (preserving the majority-voting prior), and permutes only within groups, reducing the space to O(g!).
-
Confidence-based tie-breaking as an optional secondary step. A Student's t-based confidence score on paired rank differences decides whether the top-1 candidate is genuinely better than top-2; only on a declared tie is a pointwise selector invoked, keeping the dominant path listwise.
-
A fine-tuning-free framework validated across benchmarks and generators. The method runs with off-the-shelf models over fixed candidate pools of n=8 and n=32, with reported accuracy, LLM calls, and token counts.
Main Findings
-
Accuracy over the prior state of the art: On BIRD-dev, MaP-SQL outperforms the previous state-of-the-art selector-based method R3-SQL by 2.02 execution accuracy points on average using the same candidate sets, with 2.92× fewer tokens. Across BIRD-dev, Spider-test, and EHRSQL, the improvements over R3-SQL are 2.02, 0.53, and 0.68 execution accuracy points, with 6.54×, 6.85×, and 7.29× fewer selector calls and 2.92×, 2.12×, and 4.16× fewer tokens.
-
Best single configuration: On BIRD-dev with Agentar-32B and n=32, MaP-SQL reaches 73.08%, outperforming R3-SQL by 1.11 points and the pairwise baseline by 2.02 points. A single MaP-SQL listwise selector already surpasses R3-SQL in most settings even though R3-SQL combines multiple selectors.
-
Efficiency gains: On BIRD-dev with n=32 using Arctic-R1-7B, pairwise selection averages 184.59 calls and 443,713 tokens per query, while the listwise selector uses 5.91 calls and 27,440 tokens. Compared to R3-SQL in the same setting, the full method uses 9.07× fewer calls and 4.16× fewer tokens. On EHRSQL with n=32, pairwise consumes over 1.3M tokens per query versus 48,060 for the listwise selector.
-
Generalization: On Spider-test with n=32, the full method achieves 87.59%, surpassing R3-SQL by 0.77 points using 8.43× fewer calls. On EHRSQL with n=32 it reaches 44.71%, improving over R3-SQL by 0.68 points with 11.09× fewer calls.
-
Ablation confirms both components matter: On BIRD-dev with n=8, the full system scores 72.62 on Agentar and 72.16 on Arctic. Removing permutation lowers the average from 72.39 to 72.07 (a 0.32-point drop); removing memory lowers it to 71.84 (a 0.55-point drop); removing both gives 71.64 and 71.32.
-
Memories still help a strong code selector: With a gpt-5.1-codex-mini selector, adding memory raises accuracy from 71.90% to 72.23% for Agentar-32B and from 70.01% to 70.40% for Arctic-R1-7B — a modest but consistent gain.
-
Positional bias reduction: With n=8 and shuffled initial orders, random selection yields 12.58% consistency. Baseline listwise reaches 18.98%, global permutation 23.04%, group-based permutation 29.04% (a 6.0-point gain over global), and adding tie-break 33.17% consistency with 22.63% consistency-and-correct. At four permutations, applying the tie-break reaches 33.17%, comparable to running group-based permutation eight times without it.
-
Competitive with a fine-tuned selector: Against R3-SQL's reported 71.84% on BIRD-dev with OmniSQL-7B and n=32, MaP-SQL achieves 71.90% without selector fine-tuning (majority voting under the same setting is 68.45%).
-
Tie-breaking is optional: Without tie-breaking, MaP-SQL achieves 72.23% on BIRD-dev with n=8 (versus 71.51 for R3-SQL). Reusing Qwen3-Coder for tie-breaking gives 72.42%; Contextual-RM-32B gives 72.62%.
-
Statistical reliability is partial: Paired bootstrap 95% confidence intervals exclude zero for BIRD-dev with Agentar ([0.20, 2.02], McNemar p=0.021), BIRD-dev with Arctic ([1.76, 4.37], p<0.001), and Spider-test ([0.14, 1.39], p=0.023). The EHRSQL interval includes zero ([-0.06, 1.43], p=0.125), so the authors treat that single-pool improvement as inconclusive.
-
Consistency across candidate pools: Across three pools seeded 42, 43, and 44 with n=32, MaP-SQL has the higher mean in every setting: BIRD-dev Agentar 73.01±0.36 versus 71.82±0.16; BIRD-dev Arctic 72.34±0.25 versus 69.73±0.20; Spider-test 87.89±0.26 versus 86.92±0.10; EHRSQL 45.28±0.52 versus 43.86±0.45.
-
Beats a reproduced MCS-SQL sorted-list strategy: With n=8 and the same selector and candidate pools, MaP-SQL scores 72.62, 72.16, 87.49, and 39.93 on BIRD-dev (Agentar), BIRD-dev (Arctic), Spider-test, and EHRSQL, versus MCS-SQL's 71.51, 70.47, 86.97, and 39.25.
-
Robustness to question rewording: Across the nine natural-language perturbations of Dr.Spider with n=8, listwise selection with memory achieves the highest post-perturbation accuracy at 77.80% and the smallest drop at 12.72 points; listwise without memory reaches 77.06% with a 13.13-point drop.
Methodology in Plain English
The researchers keep the standard generate–execute–select pipeline but replace selector fine-tuning with two inference-time procedures.
First, memory. For every question in the training set, they generate a compact record using the same LLM that later acts as the selector. Each record is organized into three parts borrowed from prior work: Encoding (how phrases in the question ground to tables and columns), Translating (how that meaning becomes SQL operations such as joins, aggregations, and ordering), and Decoding (what the result should look like). Memory keys include schema_grounding, join_path, filter_semantics, aggregation, ordering_and_scope, conditional_and_null, output_form, query_constraints, and extra_keywords. At test time, the top-k memories most similar to the question are retrieved with the bge-m3 dense retriever and included in the prompt, with k chosen by filling the selector's context limit rather than fixed in advance.
Second, permutation. Candidates are grouped by their execution results — the same result table, empty result, or error — with larger groups placed first because majority voting is a reliable prior. A sliding window of size w=8 and stride s=4 runs from back to front, so each selector call sees at most 8 candidates even when there are 32. Only the order within each group is shuffled, which reduces the permutation space from O(n!) to O(g!); in experiments with n=8, the average group count g was about 2. Each candidate's average rank across K runs determines the ordering, and a confidence score based on paired rank differences, evaluated against a Student's t-distribution with K−1 degrees of freedom, decides whether a tie exists (threshold τ=0.95). Ties are optionally broken by scoring only the tied candidates with a pointwise reward model.
Evaluation uses two generators, Agentar-Scale-SQL-Generation-32B and Arctic-Text2SQL-R1-7B, with Qwen3-Coder-30B-A3B-Instruct as the selector for all listwise and pairwise components, and Contextual-RM-32B only for optional tie-breaking. Experiments ran on a single node with one NVIDIA RTX PRO 6000 GPU. "Fine-tuning-free" here means selector parameters are not updated; pretrained models and labeled question–SQL pairs are still used to construct the retrieval memories.
Why This Matters
-
Impact on research: The paper reframes two things normally handled by training — selection criteria and positional-bias mitigation — as retrieval and permutation problems at inference time. It also quantifies the efficiency gap between selection paradigms, with the conclusion stating that listwise selection reduces call complexity from O(N²) to O(N)–O(N log N) and uses up to 27.92× fewer input tokens per query than pairwise selection.
-
Real-world applications:
- Enterprise analytics and business-intelligence tools where a natural-language question must be turned into a reliable SQL query over a production database.
- Clinical and electronic health record querying, as represented by the EHRSQL benchmark of 1,008 questions, where wrong queries have real consequences.
- Cross-domain database assistants, since Spider-test's 2,147 queries span many different schemas.
- Cost-sensitive deployments that cannot afford pairwise selection, given pairwise's 443,713 tokens per query on BIRD-dev and over 1.3M on EHRSQL at n=32.
-
Industry relevance: The method is designed to be compatible with off-the-shelf language models and requires no selector fine-tuning, which matters for teams that lack the data scale or compute to train a reranker. The reported reductions in LLM calls and input tokens translate directly into lower serving cost, and the reported robustness to question rewording suggests the memory retrieval does not collapse when user phrasing changes.
Future Directions
- Stronger coordination for listwise selection: The authors note the limits of sliding windows and point to advanced coordination techniques such as TourRank and Set-based ranking as unexplored alternatives.
- Closing the end-to-end accuracy gap: Reported overall execution accuracy of 70–72% lags behind SOTA systems using proprietary models at 75–76%; the authors attribute part of this to generator quality and state that their results characterize selector performance on fixed candidate pools rather than end-to-end SOTA on BIRD.
- Enterprise-scale schema handling: For databases with hundreds of tables, the authors assume upstream schema linking prunes the schema before generation and selection, leaving schema-scale context management outside the scope of the selector.
- Generalization beyond Text-to-SQL: The paper focuses on Text-to-SQL because it is a difficult listwise selection setting with long schemas, execution results, and multiple candidates; whether memory-guided selection and group-based permutation transfer to other domains is left open.
Target Audience
Researchers and engineers working on Text-to-SQL, LLM-based reranking, or retrieval-augmented inference, particularly those who care about selection-stage accuracy and serving cost rather than training a new model. It is also relevant to practitioners deploying database question-answering systems who need a selector that works with existing off-the-shelf models and limited compute, and to readers interested in positional bias in listwise ranking more broadly.
Authors’ abstract
Modern Text-to-SQL systems often follow generate-execute-select pipelines, generating multiple candidate queries then selecting the best one. Listwise selection, by jointly comparing multiple candidates, has been widely adopted, but fine-tuning listwise selectors is costly. We thus propose a fine-tuning-free listwise selector. We replace two major fine-tuning objectives with inference-time strategies: (1) learning selection criteria as ordering and (2) mitigating positional bias. First, we build reusable structured memories instead of learning selection behavior as model parameters. Given a question, MaP-SQL retrieves memories distilled from training data that encode how natural language maps to schema elements, SQL operations, and expected outputs. These memories serve as explicit decision criteria for evaluating candidates in a listwise manner. Second, to mitigate ordering bias of listwise selectors, we aggregate rankings across multiple input permutations, with inference cost optimized by execution results and pointwise scoring. Our approach improves selection accuracy while maintaining efficiency and compatibility with existing large language models. Across Text-to-SQL benchmarks, it produces more stable selection without fine-tuning and fewer unnecessary comparisons than existing methods. On BIRD-dev, it outperforms the previous state-of-the-art selector-based method R^3-SQL by 2.02 execution accuracy points on average using the same candidate sets, with 2.92x fewer tokens.