Research
From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
Overview Research area: Natural Language Processing / large language model inference (test-time scaling, agentic program search, personalization). Technical level: Intermediate. The framing is intuiti
- arXiv
- 2610.09684
- Published
- 2026-10-07
- Authors
- Xinglin Wang, Zishen Liu, Tong Zheng, Shaoxiong Feng, Peiwen Yuan, Yiwei Li, Jiayi Shi, Yueqi Zhang, Chuyi Tan, Ji Zhang, Boyuan Pan, Kan Li
AI summary
Overview
Research area: Natural Language Processing / large language model inference (test-time scaling, agentic program search, personalization).
Technical level: Intermediate. The framing is intuitive (matching user requirements), but the formal objective, controller action space, and transfer analysis require comfort with LLM inference control and replay-based evaluation.
Scope: The paper formulates Personalized Test-Time Scaling as discovery of executable TTS controllers that maximize joint satisfaction of user-specific accuracy, latency, and cost requirements, and proposes PersonTTS, an agentic framework that amortizes that discovery across users via experience reuse.
What This Paper Is About
Existing test-time scaling methods make LLMs reason better by spending more inference compute, and the efficiency literature mostly pushes either the accuracy–cost frontier or the accuracy–latency frontier — one resource dimension at a time. Real users, however, state accuracy floors and latency/cost ceilings together, and different requirement combinations can favor different controllers even on the same problem set. The paper reframes this as a preference problem — discovering executable controllers that jointly satisfy a user's accuracy, latency, and cost requirements — and attacks the practical obstacle that running policy discovery from scratch for every new user profile is expensive, cited at $39.9 and 160 minutes for a complete policy-discovery run.
Key Contributions
-
A new problem formulation. Personalized Test-Time Scaling is defined as discovering executable controllers that maximize the joint satisfaction rate (JSR) of user-specific accuracy, latency, and inference-cost requirements, rather than optimizing a single accuracy–resource Pareto frontier.
-
PersonTTS, an amortized agentic policy-discovery framework. It combines feedback-driven program search with two forms of cross-user reuse: requirement-matched controller initialization and a frozen Guide distilled from source discovery histories, while every candidate is still evaluated under the target profile.
-
Empirical validation on AIME and HMMT using six Qwen3 models, showing substantial gains over AutoTTS, ASC, ESC, and Parallel-Probe baselines on unseen user profiles and held-out problems, with reuse additionally improving policy quality and reducing discovery-agent time and cost under the same candidate-evaluation budget.
-
A transfer analysis. The paper decomposes the change in a policy's JSR between source and target profiles into a threshold-change term and a requirement-conditioned execution term, and gives bounds on target satisfaction in terms of observed execution deviation — explaining why source performance can inform but not determine target performance.
Main Findings
-
Personalized discovery dominates scalar-objective and fixed-strategy baselines. On AIME24-25 target profiles, PersonTTS reaches 96.57 JSR versus 35.58 for AutoTTS (β=1.0), 16.41 for AutoTTS (β=0.5), 10.19 for ParallelProbeSR, and 0.14 for both ASC and ESC.
-
Gains persist on held-out problems and other benchmarks. PersonTTS scores 83.85 on held-out AIME26 target profiles, 90.86 on HMMT24 discovery target profiles, and 77.28 on held-out HMMT25 target profiles. AutoTTS (β=1.0) gets 35.47 on AIME26 and 0.15 on HMMT25.
-
The main gain comes from the joint (personalized) objective, not reuse alone. The no-reuse variant, PersonTTS (w/o reuse), already scores 85.57 on AIME24-25 target, 79.22 on AIME26 target, 87.04 on HMMT24 target, and 69.16 on HMMT25 target — above all baselines — while reuse adds further improvement on top.
-
The two reuse mechanisms play different roles. Warm-start mainly improves the initial search point, whereas the Guide informs subsequent revisions and yields more consistent held-out gains across benchmarks. The paper reports their gains are not uniformly additive: removing warm-start gives 96.45 (AIME24-25 target), 83.50 (AIME26 target), 88.27 (HMMT24 target), 79.09 (HMMT25 target); removing the Guide gives 95.13, 79.60, 91.19, and 74.04 respectively.
-
The Guide drives most of the efficiency gain. Relative to no reuse, Guide-only discovery is reported to reduce agent-call time by approximately 46% and cost by approximately 36% across the two benchmarks, whereas warm-start alone mainly shortens elapsed time and can slightly increase cost.
-
Total per-profile discovery cost over five rounds (AIME). PersonTTS (w/o reuse) 74.63 minutes | 10.30 USD; PersonTTS (w/o Guide) 66.33 | 10.64; PersonTTS (w/o warm-start) 39.69 | 6.64; PersonTTS 40.41 | 7.40.
-
Total per-profile discovery cost over five rounds (HMMT). PersonTTS (w/o reuse) 85.49 minutes | 12.34 USD; PersonTTS (w/o Guide) 76.30 | 12.76; PersonTTS (w/o warm-start) 46.29 | 7.87; PersonTTS 48.51 | 8.84.
-
Discovery JSR is an informative but imperfect search signal. Round-wise improvements in discovery JSR generally accompany stronger held-out policies, and PersonTTS finishes above independent discovery on both benchmarks, but a persistent discovery–held-out gap remains.
-
Bigger experience banks help discovery but not uniformly. Enlarging the source bank consistently improves discovery-problem JSR (AIME24-25: 92.24 at 20 pairs, 94.51 at 60, 96.57 at 100; HMMT24: 86.22, 87.40, 90.86). On held-out problems the trend holds for AIME26 (80.30, 81.38, 83.85) but reverses on HMMT25 (86.09, 85.91, 77.28), which the authors attribute to a limitation of retrieval-side scaling.
Methodology in Plain English
A user profile is a triple: an accuracy floor, a latency ceiling on per-question replay latency, and a cost ceiling on mean per-question inference cost. Success is measured by the joint satisfaction rate, the fraction of evaluation seeds in which a controller meets all three at once.
The output of discovery is not a number but code — an executable controller that decides which model to call, how many branches to open, how deep to reason, when to self-refine, when to prune, and when to stop, based on public observations such as branch progress, intermediate answers, and accumulated latency and cost. Correctness labels and reference answers are never exposed to the controller.
A policy is evaluated by replaying pre-recorded, checkpointed reasoning trajectories, so no new model calls are needed. An LLM discovery agent then iteratively proposes new controller code, receiving not just the scalar JSR but also constraint pass rates and margins plus sanitized execution traces, so it can tell an accuracy shortfall apart from a resource violation. Only strict JSR improvements replace the incumbent.
To avoid paying full discovery cost for every new user, the system keeps a Policy Experience Bank of prior source-profile searches. For a new profile, it retrieves the source controller whose requirement profile is closest after standardizing accuracy and log-transformed resource ceilings, replays that controller under the target profile to get honest target-side feedback, and then uses it as the starting point. Separately, an agent compares source candidates against their requirement profiles and distills the comparisons into a frozen Guide — rules about when to revise, what to prioritize, and how to evaluate — which is injected into every subsequent proposal. The Guide is never updated from target feedback; target evaluation remains the sole basis for selecting the final controller.
Why This Matters
Impact on research. The paper moves TTS efficiency work from single-axis Pareto optimization to a multi-constraint, user-conditioned objective, and it treats LLM-driven policy search as something that should be amortized across users rather than repeated. The transfer analysis also formalizes why a controller tuned for one profile cannot be assumed to work for another, which is a caution for anyone reusing searched artifacts across settings.
Real-world applications.
- Serving LLM reasoning in products where different customers have different latency and budget constraints (for example, interactive assistants versus batch analysis).
- Routing between small and large models within a model family to hit an accuracy floor without exceeding a per-question cost budget.
- Deploying reasoning agents in latency-sensitive settings where a worst-case per-question latency ceiling matters more than average throughput.
- Multi-tenant or tiered API offerings where the same underlying problem set must satisfy different contractual accuracy, latency, and cost tiers.
Industry relevance. The cost figures matter operationally: the paper frames repeated policy discovery as a real expense and reports discovery-agent time and dollar cost separately from replay inference cost, with cross-user reuse reducing both under a fixed candidate-evaluation budget. That is the kind of accounting a serving team can act on.
Future Directions
-
Calibrated coverage and confidence. The bank-scaling results show that more source profiles help discovery but can hurt held-out generalization, so retrieval should be governed by coverage or confidence estimates rather than bank size alone.
-
Request-frequency-aware deployment. The authors propose reusing matched controllers when the experience bank is reliable, while triggering background discovery for frequent or poorly covered profiles whose search cost can be amortized over future requests.
-
Closing the discovery–held-out gap. Since discovery JSR is an informative search signal rather than a direct estimate of out-of-sample satisfaction, better selection criteria or validation protocols are an open problem.
-
Validation beyond offline replay. The ethics statement notes that deployment would require validation on the intended workload and resource accounting; the reported metric is defined by accuracy and resource constraints in the replay setting and does not measure subjective satisfaction.
Target Audience
Researchers and engineers working on LLM inference efficiency, test-time scaling, and agentic system design who need to satisfy multiple resource constraints simultaneously; practitioners building model-routing or multi-model serving systems with per-request latency and cost budgets; and readers interested in how LLM agents can be used to synthesize executable policies and how such search can be amortized across users.
Authors’ abstract
Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.