Research
Persona Dosing: Calibrated Activation Steering for Graded Trait Control
Overview Research area: Interpretability and activation-level control of large language models, specifically steering persona traits through residual-stream interventions. Technical level: Advanced. T

- arXiv
- 2609.36388
- Published
- 2026-09-28
- Authors
- Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen, Jingyuan Zhang, Yuxuan Zhang, Xinjie Shen
AI summary
Overview
Research area: Interpretability and activation-level control of large language models, specifically steering persona traits through residual-stream interventions.
Technical level: Advanced. The paper assumes familiarity with activation steering, flow matching, isotonic regression, and behavioral rubric scoring.
Scope: The paper introduces PersonaDose, a calibrated activation-steering method that maps a requested behavioral intensity score to an intervention strength for a shared, description-conditioned controller across three model families and seven traits.
What This Paper Is About
Activation steering lets researchers scale up a trait by increasing a coefficient, but the coefficient itself has no behavioral meaning: the same value produces different expression levels across traits and models. This makes it impossible to specify an experiment as "generate responses at sycophancy 60." The authors study persona dosing, an interface where a trait description selects what to express and a requested score selects how strongly, using a shared controller specialized on persona responses and a post-training calibration curve that converts scores into flow times.
Key Contributions
- Persona control in behavioral units. A trait description plus a requested mean score selects an intervention through persona-response specialization and post-training calibration, without pairing training responses with requested target intensities.
- Higher expression under a coherence floor. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, persona specialization raises aggregate core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition (CAA), with held-out evaluations supporting the advantage.
- Graded access to learned behaviors. Calibrated strengths realize intermediate mean intensities on new questions across seven trained traits, with mean targeting errors of 4.7–6.2 points, and recalibration supports graded requests after a controller update.
- Separation of reach from precision. The design and evaluation distinguish the behavioral range a controller supplies from the accuracy of requests within that range, reporting reachability, held-out coherence, and error together.
Main Findings
- Expression at the coherence floor: PersonaDose reaches core-trait expression scores of 75.1 (Llama), 82.3 (Qwen), and 86.0 (Gemma) at mean coherence of at least 75, compared with CAA's 41.9, 64.0, and 68.2. Generic-flow scores are 13.3, 46.2, and 42.7; RepE scores are 18.5, 31.4, and 40.3; the linear separator scores are 31.7, 58.6, and 50.2.
- Per-trait gains: On Llama, specialization raises sycophantic expression from 9.7 to 90.5 and impolite expression from 1.6 to 72.3. Llama improvements over CAA are 38.3, 26.9, and 34.6 points across the three core traits. On Qwen and Gemma, the largest improvement is on evil, rising from 3.3 to 71.5 and from 21.2 to 68.9 respectively.
- Held-out expression advantage: After calibration-only strength selection, PersonaDose achieves held-out mean expression of 75.6, 80.3, and 84.8 versus CAA's 43.5, 63.3, and 68.0, with paired gains of 32.1 [26.4, 37.8], 17.0 [10.4, 23.7], and 16.8 [12.4, 21.5]. Mean coherence is 75.5, 83.2, and 79.4, but only 1/3, 2/3, and 3/3 traits retain the floor on test questions, so the calibration coherence constraint does not consistently transfer at the trait level.
- Intermediate dosing: Across four targets per trait on seven traits (28 trait–target cells per model), PersonaDose achieves mean targeting MAE of 6.1 (Llama), 6.2 (Qwen), and 4.7 (Gemma) over 22/28, 20/28, and 14/28 reachable requests. Of all 28 requests per model, 19, 20, and 13 are both reachable and above the test mean-coherence floor.
- Trait-specific mappings: In the Qwen seven-trait study, sycophancy targets 20, 40, 60, and 80 yield held-out means of 12.9, 37.1, 56.3, and 80.5, with mean coherence of 88.9–92.9. Target 60 uses flow time T = 1.4864, and all settings share the same weights.
- Controller updates shift the map: Continuing to train the Qwen controller raises held-out expression from 78.7 to 93.7, a paired gain of 15.0 [7.3, 23.3], at mean coherence 81.9, with all three traits retaining the floor. For sycophancy 60, the fitted map selects T = 1.4864 for the initial checkpoint and T = 0.8745 for the continued checkpoint.
- Calibrated dosing versus CAA: On the continued Qwen controller over the three core traits and twelve requests, PersonaDose supports 11/12 coherent requests versus 8/12 for calibrated CAA, with mean coherence 86.3 versus 76.5 and MAE 6.2 versus 8.2. The paper notes these MAEs cover different requests and do not establish a paired targeting-accuracy gain.
- Uneven reachability: Reachable target counts vary by trait: optimistic has only one reachable target on each model, and Gemma hallucinating has none.
- Reach does not equal precision: In the continued Qwen core-trait study, mean targeting error is 6.2 with coherence 86.3 across 11/12 reachable requests, and all eleven cell means exceed coherence 75, yet the largest errors include evil/80 at 10.6 and hallucinating/80 at 15.5 points. A separate Llama continuation supports 11/12 requests with error 4.2 and coherence 82.2, with nine cell means above 75.
Methodology in Plain English
The authors freeze a language model and attach a small trainable module that transforms the model's internal activations at decoder layer 20. A trait description conditions this module, so a single controller per base model can express any of the seven traits; the "flow time" parameter controls how far the transformation is applied, and thus how strong the trait is.
Training uses responses the base model already produced at trait score of at least 50 under the Persona Vectors rubric. Crucially, responses are never labeled with a requested intensity; instead the same response is supervised at flow times sampled uniformly from 0.5 to 2.0, so the controller learns the behavior rather than a specific dial position. Each model's corpus has 1,250 rows: 150 examples for each of seven persona concepts plus 200 generic replay examples. The objective combines response cross-entropy with a concept-diversity regularizer weighted at 0.1, which discourages identical transformation directions across concepts.
After training, the authors measure trait expression on a grid of strengths using ten calibration questions per trait, fit an increasing isotonic regression to the calibration means, and interpolate to build a response curve. Inverting the curve gives the strength that realizes a requested score, with plateaus resolved to the smallest strength and out-of-range targets marked unreachable before test evaluation. The scheme is validated on a disjoint set of ten test questions per trait, with targets of 20, 40, 60, and 80. Comparisons use CAA, RepE, a logistic linear separator, and a generic flow controller without persona specialization, all intervening at the same layer with the same questions, response counts, decoding settings (temperature 1, 256-token limit), and GPT-4.1-mini judging protocol. Uncertainty is reported via 2,000 question-bootstrap resamples within traits.
Why This Matters
Impact on research. Graded persona control is a prerequisite for studying behaviors like sycophancy at controlled intensities. The work reframes the steering coefficient as a measurement problem: the coefficient has no intrinsic behavioral meaning, so the same requested score selects different flow times across models and traits, and a controller update changes the calibration map. Reporting reachability, held-out coherence, and targeting error as separate quantities gives a template for evaluating other steering methods.
Real-world applications (as suggested by the paper's framing):
- Studying sycophancy and other graded traits by generating responses with measurable degrees of agreement to the same user premise.
- Behavioral red-teaming and controllability characterization for undesirable traits such as evil, hallucinating, and impolite expression.
- Building benchmark suites that require several measurable intensity levels of a behavior rather than a single steering strength.
- Rapid re-calibration after a model or controller update so that existing behavioral study protocols remain valid.
Industry relevance. Because one shared controller per base model handles seven traits and switching traits does not require loading a different adapter, the approach is practical for teams that need per-trait intensity control without maintaining separate steering artifacts. The paper's caution that trait and coherence scores are not safety guarantees, and that such interventions can amplify harmful behavior, is directly relevant to deployment decisions.
Future Directions
- Generalization beyond trained traits. The evaluation covers seven trained traits and short responses; whether the interface works for unseen traits or long conversations is not established.
- Stale calibration maps after updates. The effect on targeting error from applying an old calibration map to an updated checkpoint has not been measured, even though the paper shows the map shifts (sycophancy 60 moving from T = 1.4864 to T = 0.8745).
- Isolating continuation effects. Controller continuation changes response sampling, chat-prefix construction, and sequence length together, so the individual contribution of each change is unattributed.
- Paired targeting comparisons. Since MAEs average each method's own reachable set, a paired evaluation on common requests would establish whether calibration accuracy itself improves; secondary-judge MAE intervals include zero.
Target Audience
Researchers in mechanistic interpretability and activation steering who need graded behavioral control for experiments; safety and alignment teams characterizing controllable traits; and engineers building persona-controlled generation systems who want a behavioral interface (trait description plus requested score) rather than an opaque steering coefficient. The paper is most useful to readers already comfortable with flow-based interventions and rubric-based LLM judging.
Authors’ abstract
An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.