Research
Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30
Overview Research area: Natural language processing and AI safety, specifically the psychometric evaluation of large language models (moral profiling) and mechanistic/prompt-level steering of model be
- arXiv
- 2609.21636
- Published
- 2026-09-18
- Authors
- Hans Andersen, David Dichas
AI summary
Overview
Research area: Natural language processing and AI safety, specifically the psychometric evaluation of large language models (moral profiling) and mechanistic/prompt-level steering of model behaviour.
Technical level: Intermediate. Readers need some familiarity with transformer internals (residual-stream activations, layers, logits) and with psychometrics (Likert scales, Cronbach's alpha, Mahalanobis distance), but the paper's argument is stated plainly.
Scope: The paper administers the Norwegian MFQ-30 moral foundations questionnaire to six open-weight LLMs, compares their profiles against 1,282 Norwegian human respondents, and tests whether prompt-level persona steering and activation-level ActAdd can move model profiles toward the human distribution.
What This Paper Is About
LLMs are increasingly used in human-facing settings, so it matters whether their moral profiles resemble those of the people they serve. Prior work applies human psychometric questionnaires to LLMs, but it is unclear whether the questionnaire measures anything stable in a model, or whether the resulting profile can be moved toward a target human population. This paper asks both questions in a Norwegian setting: do six open-weight LLMs answer the Norwegian MFQ-30 coherently, and can those answers be steered toward the mean of 1,282 Norwegian respondents using a persona prompt or an activation-level intervention.
Key Contributions
-
A Norwegian MFQ-30 elicitation pipeline with an attention gate. The authors build a first-token probability decoding method over the digits 1–6, add a "Svar:" completion suffix, and define an attention-check gate whose threshold (a 92.2% human pass rate) is derived only from the human reference sample, never from model output.
-
A two-tier empirical split across six models. Half the tested models engage with the questionnaire under the attention check; the other half produce flat or central-tendency outputs that look near-human on average without tracking item content.
-
A head-to-head comparison of prompt-level and activation-level steering. Persona prompting (nordic_b) moves engaging models 44–77% closer to the Norwegian mean in Mahalanobis d², while one-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations.
-
A documented instance of "cognitive phantoms." For Qwen3-8B under reversed scale, the same persona that shifts the foundation profile also induces questionnaire engagement that was absent at baseline — raised from roughly 30% to roughly 98% reversed-scale attention — which the authors connect directly to the caution raised by Peereboom et al. (2025).
Main Findings
-
The attention split is stark and stable. Gemma 4-E4B-it, Qwen3-14B and Qwen3-8B pass the forward attention check; NorMistral-7B, NorMistral-11B-Thinking and Qwen2.5-1.5B-Instruct fail. Across every per-item condition run (both scale orders, baseline and persona), no model lands between 37.2% and 93.8% joint pass probability, so any gate placed inside that interval gives the same classification.
-
Qwen3-8B passes forward but not reversed. It passes forward at 99% but drops to 30% reversed, the signature of a model reading digit position rather than label semantics. Gemma 4 and Qwen3-14B pass under both scale orders and are treated as the robust passers.
-
Baseline distances from the human centroid. Gemma 4 goes from d² = 5.42 forward to 2.80 reversed; Qwen3-14B from 8.98 to 4.42. Their canonical forward/reversed midpoint d² values are 3.65 and 6.26. Qwen3-8B is reported forward-only at d² = 11.59 because its reversed attention failed.
-
Baseline profiles overshoot on the individualizing foundations. All three passers sit above the human pentagon on fairness. Gemma 4 lifts care and fairness the most, by about 6 points each (33.0 vs human 26.7 on care; 33.3 vs 27.0 on fairness) but stays near human means on loyalty, authority and purity. Qwen3-14B has a smaller individualizing lift plus a separate purity gap (26.7 vs human 20.0). Qwen3-8B shows a similar purity gap and is the only passer below the human mean on care.
-
Only Gemma 4 reproduces the human foundation ordering. The human sample puts care and fairness on top (26.7, 27.0) with loyalty, authority and purity below (21.6, 22.0, 20.0). Gemma 4 reproduces this ordering with a wider two-cluster gap (about 11 points vs the human 5). In Qwen3-14B and Qwen3-8B, purity climbs into or past the individualizing range, breaking the binding cluster.
-
Theory-grounded personas move passing models, not failing ones. The individualizing-vs-binding persona contrast is 32.6 for Gemma 4-E4B-it, 21.9 for Qwen3-14B and 20.0 for Qwen3-8B. Two of the three failers stay below 2 (NorMistral-7B at 1.8, Qwen2.5-1.5B-Instruct at 0.4). NorMistral-11B-Thinking responds at first-token resolution (approximately 11.5) but still fails the batched attention check.
-
The neutral Nordic persona closes most of the gap. The nordic_b persona drops canonical-midpoint d² from 3.65 to 2.05 for Gemma 4-E4B-it (44%) and from 6.26 to 2.37 for Qwen3-14B (62%). For Qwen3-8B, compared forward-only, d² drops from 11.59 to 2.69 (77%); its midpoint under nordic_b is 1.81. A residual gap of about 2 remains across all three models.
-
The persona writes a profile the model was never shown. nordic_b contains no foundation scores, no answer frequencies and no example responses — only a short demographic persona (an adult Norwegian web-panel respondent with a moderate moral worldview in a Nordic welfare state). It nevertheless pulls three model profiles substantially closer to a human answer distribution the model has no access to.
-
Persona steering also changes engagement. Under nordic_b, Qwen3-8B becomes a robust two-direction passer, with reversed-scale attention climbing from roughly 30% to roughly 98%. Gemma 4 and Qwen3-14B stay close to 100% attention under both scale orders. The persona does not rescue the hard failers: at the canonical midpoint, nordic_b leaves NorMistral-7B at 17.1%, NorMistral-11B-Thinking at 19.3% and Qwen2.5-1.5B-Instruct at 33.1%.
-
One-pair ActAdd does not steer selectively. At layer 15 with α = 5, neither contrastive construction produced isolated single-foundation steering. On the loyalty/betrayal pair, Qwen3-8B and Qwen3-14B collapse foundation differentiation: per-foundation expected scores fall within 0.13 of one another across all five foundations at α = +5, and within 0.04 at α = −5. Under the individualizing and binding presets, Qwen3-14B's per-foundation scores fall to about 8.8 and 7.0, close to the floor of the 6–36 scale, and its d² rises from 9.0 at baseline to 31.5 and 38.4.
-
Gemma 4's ActAdd shifts are not foundation-specific. At α = +5, loyalty moves +5.7 relative to the no-steering baseline, but authority moves +8.4 and purity +5.9. At α = −5, loyalty shifts by +1.0 rather than the symmetric −5.7 a clean linear axis would predict, indicating the extracted vector is not antisymmetric in α.
-
The ActAdd collapse is monotone in coefficient. In the NorMistral-7B α-sweep on the binding preset, the per-item parsed-score range goes from four distinct values at α = 5, to two at α = 15, to a single value at α = 20 — a tuning artefact rather than a poorly chosen constant.
-
Batched presentation failed. Under batched presentation, only Gemma 4 produced parseable output across both scale orders (27/32 forward, 19/32 reversed); the NorMistrals worked partially forward only, and all three Qwen models produced zero parseable items in either order, so batched presentation was excluded.
-
Suffix wording affects ranks only weakly. In the four-suffix robustness sweep (Svar:, Mitt svar er:, Tall:, Svar (1-6):), Kendall's W = 0.963 for Qwen2.5-1.5B and W = 0.762 for NorMistral-7B, both p < 0.01. Absolute foundation sums and absolute d² values are not invariant to suffix wording. The no-suffix baseline breaks rank concordance entirely (W drops to 0.35 and 0.62).
Methodology in Plain English
The researchers took an existing, validated Norwegian version of the Moral Foundations Questionnaire (MFQ-30), created and validated by Enstad and Finseraas (2024), and put its 30 items to six open-weight models: Gemma 4-E4B-it (4B), Qwen3-14B, Qwen3-8B, Qwen2.5-1.5B-Instruct, NorMistral-11B-Thinking and NorMistral-7B-Warm-Instruct. Each foundation is scored as the sum of six items on a 6–36 range, and the human comparison comes from the 1,282 respondents released alongside the questionnaire (Kantar web panel, September 2021, with a slight over-representation of high-education adults).
Rather than asking a model to write a digit and parsing it with a regex — which broke down, with NorMistral producing Norwegian prose and Qwen2.5-1.5B looping on a single digit — they read the model's next-token probabilities directly. At each item they keep only the six logits for digits 1–6, softmax them, and take the weighted average sum of k × p_k. This means the model never commits to a single digit, no sampling is needed, and the values are deterministic given the prompt. Before reading those logits, they append the string "Svar:" ("Answer:") to give the model an answer-completion context.
They added an attention check on the two control items, MFQ1_6 and MFQ2_6, passing a model only if MFQ1_6 ≤ 3 and MFQ2_6 ≥ 4 as a joint analytic probability. The threshold is the human pass rate of 92.2%, which comes only from the human sample and not from any model output, so a model clears the gate only by being at least as attentive as a screened human respondent. They note a stricter version of the rule would fail humans 33% of the time with this dataset, which is why they did not tighten it.
Similarity to humans is summarised by squared Mahalanobis distance d² between the model's five-foundation mean vector and the human centroid, using the 5×5 human covariance matrix. This weights drift along directions humans actually vary on (such as the care–fairness axis, r = 0.61) more cheaply than drift along directions they do not.
The researchers also varied two presentation conventions on a 2×2 grid — per-item vs batched presentation, and forward (1 = low) vs reversed (1 = high) Likert labelling — all at T = 0.7, with batched corners using max_new_tokens = 512. Because forward and reversed per-item runs disagreed, they report the midpoint of the two as the canonical score for models that pass attention under both orders, while stating they cannot tell which corner is closer to the model's actual profile.
For Phase 2 they prepended Norwegian persona descriptions to the system prompt, keeping the user message identical: theory-grounded individualizing and binding personas as a steerability sanity check, a demographic nordic_a, the main treatment nordic_b (nordic_a plus a welfare-state anchor), and nordic_c, which names the foundations and was excluded to avoid data leakage.
For Phase 3 they followed one-pair ActAdd: run positive and negative contrastive prompts, take the mean hidden state at the output of layer 15 over all token positions, subtract to get a steering vector, and add α times that vector at layer 15 during every MFQ forward pass, with no persona in the system prompt. Layer 15 sits between roughly 36% and 54% of network depth given layer counts from 28 (Qwen2.5-1.5B) to 42 (Gemma 4-E4B-it). They fixed α = 5 after a sweep on NorMistral-7B over {5, 10, 15, 20} on both two-cluster presets, choosing the smallest coefficient that still injects non-trivially. They did not search over layer indices.
All runs used a single Apple M1 Max with 64 GB unified memory, the PyTorch MPS backend, and fp16 precision. Code is available at https://github.uio.no/haan/IN5550-llm-moral-foundations. The paper also discloses the use of Anthropic's Claude (Opus 4.6 and 4.7, 2026) for first-pass rephrasing and most of the code implementation, with all quantitative claims checked against the authors' own pipeline outputs.
Why This Matters
Impact on research. The paper pushes back on the assumption that a psychometric questionnaire measures the same construct in a model as in a person. Its attention gate and its finding that failers can look human-like on average purely by regression to the midpoint are methodological warnings for anyone running survey instruments on LLMs. It also supplies a clean negative result: a lightweight one-pair ActAdd setup at a single layer and coefficient does not deliver selective foundation steering here, which is useful information for the interpretability community.
Real-world applications
- Cross-cultural model evaluation. The work shows why evaluating in the target language on target-language-specialised models matters, rather than translating an English run — a point reinforced by Aksoy (2025), who found MFQ-2 profiles vary by prompt language.
- AI alignment auditing. Persona-based steering offers a cheap way to test whether a model's value profile can be shifted toward a target population, relevant to deploying assistants in public-sector or civic contexts.
- Survey-methodology safeguards. The attention-check gate provides a reusable pattern for detecting LLMs that produce plausible survey answers without engaging with item content.
- Localised public-sector deployment. With generative AI reported at 53% adoption in the general population and 88% among organisations (Sajadieh, Sha et al., 2026), knowing whether a Norwegian-specialised model reflects Norwegian moral profiles matters for public services in that market.
Industry relevance. The comparison of prompt-level versus activation-level steering speaks directly to teams deciding how to shape model behaviour in production. The result here is practical: a short persona prompt achieved a 44–77% reduction in Mahalanobis distance to the target population, while the activation-level route at this configuration broke rather than steered the target behaviour. The paper's caution that cross-item correlation structure cannot be recovered from single-turn per-item queries also bears on how vendors should — and should not — claim that models "match human values."
Future Directions
-
Layer and pair search for activation steering. The paper injects at layer 15 across all models and selects α = 5 from a sweep on NorMistral-7B only. The authors name the averaged-pair extension of ActAdd (contrastive activation addition, Panickssery et al., 2024) as a plausible fix for the foundation-differentiation collapse they observe, and note they did not search over layer indices.
-
Validating personas on a held-out Norwegian sample. The nordic_b canonical d² values are computed against the same sample that motivated the persona's design. A separate Norwegian sample would
Authors’ abstract
Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis $d^2$. One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.