Skip to content
AI.info

Research

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

Overview Research area: AI safety and LLM evaluation, specifically psychometric measurement of language models (self-report instruments as predictors of downstream behavior). Technical level: Advanced

arXiv
2606.12730
Published
2026-06-10
Authors
Rafal Kocielnik, Pengrui Han, Peiyang Song, Myrl G. Marmarelis, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez

AI summary

Overview

Research area: AI safety and LLM evaluation, specifically psychometric measurement of language models (self-report instruments as predictors of downstream behavior).

Technical level: Advanced. The paper assumes familiarity with psychometrics, meta-analytic pooling, factorial experimental design, and behavioral task paradigms.

Scope: A 2×2×2 factorial study across 4 behavioral tasks and 11 frontier LLMs testing when and why LLM self-reports predict behavior, contrasting the fine-grained Theory of Planned Behavior against the coarse Big Five.

What This Paper Is About

Prior work found that LLMs produce internally coherent personality profiles that fail to predict how they actually behave in tasks, but left open whether this gap is a property of the models or an artifact of the instrument used to measure them. This paper tests that question directly by swapping the dominant Big Five inventory for the Theory of Planned Behavior, a fine-grained instrument anchored to specific behaviors, and by systematically varying whether self-report and behavior happen in the same conversation, in separate conversations, and under parameter-grid versus persona-based identity induction.

Key Contributions

  1. A 2×2×2 factorial design isolating the source of the self-report–behavior gap. The design crosses framework (Big Five vs. Theory of Planned Behavior), session context (shared message thread vs. separate sessions), and identity induction (parameter perturbation vs. psychologically grounded persona prompting), applied across 4 behavioral tasks and 11 frontier LLMs.

  2. A theoretical account of selective coherence. The paper distinguishes common-cause coupling (self-report and behavior both shaped by stable model state) from within-session context priming (coherence depends on the self-report remaining in the prompt window at behavior time), and shows which of the two each task pattern reflects.

  3. A demonstration that Big Five does not predict behavior at all under these conditions. The 88-cell Big Five matrix shows only 3 of 88 cells reach p < .05, and only 1 of those in the theoretically expected direction.

  4. A practical mapping of when self-reports are and are not diagnostic of behavior, including the safety-relevant finding that persona prompting stabilizes self-reports across sessions but does not restore behavioral coupling.

Main Findings

  • Within-session coherence exists and matches human baselines — for the fine-grained instrument. Pooled across tasks, the Fisher-z-aggregated mean within-model correlation between Theory of Planned Behavior self-report and behavior was r = +0.25, 95% CI [+0.22, +0.28]. Excluding the theoretically dissociated implicit bias (IAT) task, this rises to r = +0.40, 95% CI [+0.37, +0.43], within the range of human meta-analytic intention–behavior correlations (r ≈ 0.25–0.50). Of 77 cells, 41 (53.2%, Wilson 95% CI [42.2%, 64.0%]) were both theory-aligned and significant at p < .05 — 21.3× the 2.5% expected under a pure null (z = 28.5, p < .0001).

  • Per-task patterns follow the theory's scope. Honesty × Attitude: r = +0.67, CI [+0.63, +0.72]. Sycophancy × Intention: r = +0.47, CI [+0.39, +0.53]. Columbia Card Task × Intention: r = +0.22, CI [+0.14, +0.30]. IAT showed the theoretically expected explicit–implicit dissociation at r = −0.59, CI [−0.64, −0.53], consistent with documented human compensatory-effort inversions (human explicit–implicit r ≈ 0.15–0.25, often negative when motivation to suppress is high).

  • Strong per-model heterogeneity. Claude 4.5 Haiku showed the strongest coherence (r = +0.75, CI [+0.70, +0.79]), followed by Qwen 235B (r = +0.72), LLaMA 4 Maverick (r = +0.50), and LLaMA 3.3 70B (r = +0.38). Two models fell below zero: Phi-4 (r = −0.11) and Claude 3.7 Sonnet (r = −0.53), the latter dominated by inverted coherence on the three volitional tasks (Honesty r = −0.73, Sycophancy r = −0.40).

  • Big Five is not merely weaker — it does not predict at all here. Under identical within-session, shared-context conditions, best Big Five r_aligned across the three volitional tasks ranged only +0.06 to +0.07, with every 95% CI crossing zero. The per-task framework gap was Δ = +0.61 (Honesty), +0.40 (Sycophancy), +0.16 (CCT). Theory of Planned Behavior beat Big Five in 8 of 11 models; mean per-model r_aligned was +0.21 for Theory of Planned Behavior versus +0.01 for Big Five. Largest gaps appeared in the strongest models: Claude 4.5 Haiku (Δ = +0.74), Qwen 235B (+0.64), LLaMA 4 Maverick (+0.48).

  • Cross-session survival is task-dependent, not uniform. Honesty partially survived: r = +0.67 → +0.53, Δr = +0.14, CI [+0.04, +0.25], p < .001. Sycophancy collapsed completely: r = +0.47 → −0.07, Δr = +0.54, CI [+0.39, +0.69], p < .001. CCT showed marginal, non-significant reduction: r = +0.22 → +0.12, Δr = +0.10, CI [−0.06, +0.26]. IAT was stable and slightly strengthened: r = −0.59 → −0.66, Δr = +0.07, ns.

  • Only 2 of 11 models retained cross-session coherence. Claude 4.5 Haiku (r = +0.75 → +0.65, Δr = +0.10) and LLaMA 3.3 70B (r = +0.38 → +0.66, Δr = −0.27). The largest collapse was Qwen 235B (+0.72 → −0.14, Δr = +0.86).

  • The collapse is driven by behavior, not by self-report drift. Cross-session self-report consistency was high across all four tasks (Honesty +0.81, Sycophancy +0.59, CCT +0.23, IAT +0.52). Behavior consistency tracked survival instead: IAT +0.98, Honesty +0.45, CCT +0.41, Sycophancy −0.02, with several models showing strongly negative cross-session behavioral correlations on sycophancy (Qwen 235B −0.96, Gemini 2.5 Flash −0.90, DeepSeek V3.1 −0.89).

  • Persona induction does not rescue coherence. No model met the rescue criterion. Sycophancy showed only partial rescue: −0.07, CI [−0.18, +0.02] under grid to +0.09, CI [−0.01, +0.18] under personas, Δr = +0.16, CI [+0.00, +0.32], p < .01, with the bootstrap CI [−0.05, +0.35] less conclusive. Honesty attenuated: +0.53 → +0.38, Δr = −0.15, CI [−0.28, −0.03], p < .001. CCT and IAT were induction-invariant. The two previously coherent models attenuated rather than recovered (Δr = −0.25 and −0.15).

  • Personas change what models say, not what they do. Self-report diversity was positive in 50% of (model × task) cells and above 0.3 SD in 11 of 44, with the largest gains on Gemini 2.5 Flash CCT (+0.70), Claude 3.7 Sonnet IAT (+0.55), and Claude 4.5 Haiku Sycophancy (+0.53). Self-report stability was positive in 75% of cells (mean Δr = +0.14), strongest for LLaMA 3.3 70B IAT (+0.50), Gemini 2.5 Flash CCT (+0.61), and Mistral Large IAT (+0.42). None of this translated into behavioral coupling.

  • Same-session probes are not neutral. Asked to evaluate a policy, models tend to adopt it, shifting self-report or behavior toward it despite not being asked to do so — which is why same-session coherence for context-loaded tasks cannot distinguish priming from disposition.

Methodology in Plain English

The researchers treated the question as a measurement design problem. They first built the most favorable possible test: place a fine-grained questionnaire and a behavioral task inside a single conversation so the model can see what it just said about itself, then check whether its stated intentions correlate with its choices. The questionnaire was the Theory of Planned Behavior, adapted for each task by anchoring items to a specific Target, Action, Context, and Time — for example, "When making risky decisions in this card game, I intend to flip cards carefully" rather than a generic trait item like "I see myself as someone who is cautious."

They then relaxed one condition at a time. First, they swapped the fine-grained instrument for Big Five, holding everything else constant, to see whether the granularity mattered. Second, they moved self-report and behavior into separate API calls that shared only initialization settings — matched temperature, seed, and system prompt — but no message history, mirroring how a model's stated dispositions and its actual behavior rarely share a conversation in deployment. Third, they replaced random sampling variation with named persona descriptions (30 PersonaHub characters selected for diversity across demographic, occupational, and personality dimensions at fixed temperature 0.2) to test whether providing a stable identity could restore cross-session coherence.

Across all conditions, the unit of analysis was a (model × task × construct) cell, with within-model Pearson correlations estimated from roughly 54 observations under grid induction and 60 under persona induction. Cell-level correlations were combined using inverse-variance-weighted Fisher-z meta-analysis to produce pooled r values with 95% confidence intervals, directly comparable to human meta-analytic benchmarks. Proportion-based results used Wilson confidence intervals against a null baseline of 2.5%. Primary findings were validated with pooled OLS using Mundlak within/between decomposition and cluster-robust standard errors, a policy-contrast difference-score specification that removes response-style variance, and model-resampling bootstrap confidence intervals. Critically, neither phase instructed the model to behave consistently: the questionnaire gave no indication a task would follow, and the task never referenced the questionnaire.

Why This Matters

Impact on research. The paper reframes a documented dissociation as a measurement-design problem rather than a model deficiency. It argues that the field's reliance on Big Five — a deliberately cross-situational taxonomy that predicts specific behaviors weakly even in humans — has been generating null results for structural reasons. It also offers a methodological rule: because same-session probes conflate priming with disposition for context-loaded tasks, behavioral safety probes meant to predict deployment should elicit self-report and target behavior in separate sessions. Validation studies that link self-report to generation through matched persona descriptors alone can recover context-independent properties but cannot generalize to context-loaded ones.

Real-world applications:

  • Auditing risk-taking behavior in financial advising deployments, anticipating how a model will behave before it is trusted with consequential decisions.
  • Testing whether models communicate confidence reliably in clinical decision support and medical advice settings.
  • Evaluating models used for educational tutoring, where stated pedagogical intentions need to match actual interactions with students.
  • Building a benchmarking framework that lets stakeholders test LLM behavior in specific agentic decision-making deployment settings.

Industry relevance. The persona finding is directly safety-relevant: persona-customized deployments may produce confidently distinct self-reports without correspondingly distinct behavior, meaning a model that says it is a cautious advisor may not act like one. As LLMs increasingly shape user behavior, characterizing when cheap self-report probes can substitute for expensive behavioral batteries is a prerequisite for scalable behavioral auditing.

Future Directions

  • Self-correction under exposure to decoherence. Pairing self-report items with feedback about prior self-report–behavior mismatches, to test whether reasoning-trained models can close their own gap.

  • LLM-native self-report frameworks. Existing instruments are imported wholesale from human psychometrics; building instruments designed around LLM response patterns from the ground up could shift the field from adapting human scales to creating native ones.

  • Mechanistic interpretability of the dissociation. Testing whether self-report and behavior generations share early-layer activations that diverge in later layers — restricted for now to small open models, since the proprietary frontier-scale models dominating this sample are inaccessible to internal probing.

  • Cross-task generalization. The paper notes that even where coherence is measured in one task, it may not translate to another, implying a need for broader testing before any instrument is trusted as a general behavioral predictor.

Target Audience

AI safety and evaluation researchers who need to know when cheap psychometric probes can stand in for expensive behavioral testing; psychometricians and social scientists interested in whether human measurement instruments transfer to LLMs; model deployment and red-teaming teams who build persona-customized or agentic systems; and policy stakeholders designing behavioral auditing standards for AI systems in high-stakes domains such as finance, medicine, and education.

Authors’ abstract

Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior. Recent work documented substantial SR-behavior dissociation in LLMs, but relied on broad personality traits (Big 5) that predict specific behaviors weakly, even in humans. Furthermore, the isolation of conversational sessions combined with weak context matching left open whether LLMs truly lack coherence or whether the conditions needed to detect such coherence were not met. We contrast Big 5 with the Theory of Planned Behavior (TPB), which measures intention targeted to a specific behavior and predicts human behavior substantially better than broad traits. We run experiments across four behavioral tasks and 11 frontier LLMs, while also varying session context and identity induction. We find that SR-behavior coherence exists but is selective. 1) Within a shared conversation, the Theory of Planned Behavior reaches human-level coherence; Big 5 does not. 2) Across separate conversations, coherence survives only for behaviors anchored outside the immediate prompt, such as implicit bias shaped by training, and collapses when behavior is strongly primed by context, as with sycophancy. 3) Persona prompting makes self-reports more consistent across conversations, but does not bring behavior into alignment. These findings suggest that coarse personality frameworks, such as Big 5 may not be the best tools for testing deployment behavior. More task- and behavior-specific instruments are needed, and even these must be evaluated across tasks and contexts.

Read the original paper