Research
Beyond Fixed Psychological Personas: State Beats Trait, but Language Models are State-Blind
Overview Research area: Natural Language Processing, with ties to affective computing, persona/personalization research, and RLHF alignment. Technical level: Intermediate. Scope: The paper introduces
- arXiv
- 2601.15395
- Published
- 2026-01-21
- Authors
- Tamunotonye Harry, Ivoline Ngong, Chima Nweke, Yuanyuan Feng, Joseph Near
AI summary
Overview
Research area: Natural Language Processing, with ties to affective computing, persona/personalization research, and RLHF alignment.
Technical level: Intermediate.
Scope: The paper introduces the Chameleon dataset of 5,001 contextual psychological profiles from 1,667 Reddit users and uses it to show that psychological expression in text is mostly context-driven (state), yet LLMs ignore that context while reward models respond to it inconsistently.
What This Paper Is About
Existing persona datasets such as PersonaChat and PANDORA treat a user's psychology as a fixed, stable attribute, capturing only how one person differs from another. Drawing on Latent State-Trait theory, the authors argue this misses the larger part of the picture: the same person expresses different psychological characteristics across different contexts. They build a dataset that measures the same users in multiple contexts and then test whether language models and reward models handle that context correctly.
Key Contributions
- Chameleon dataset: 5,001 contextual psychological profiles from 1,667 Reddit users across 645 subreddits, spanning 26 dimensions across four validated psychological frameworks, with each user measured across multiple contexts. The paper states this enables state-trait decomposition in NLP for the first time.
- Empirical variance decomposition: Evidence that 72–74% of psychological variance in text is within-person (state), while only 26–28% is between-person (trait), replicated across two methodologically distinct extraction pipelines.
- Two applied evaluations: Showing that LLMs are "state-blind" in generation (they detect persona framing but fail to differentiate profiles) and that reward models are context-aware but inconsistent in evaluation (they disagree on direction about the same users).
- Public release: The dataset is released at
https://huggingface.co/datasets/tonyeh/chameleon-datasetto support work on psychology-aware AI systems, affective computing, personalized dialogue, and RLHF alignment.
Main Findings
- State beats trait: SEANCE-derived profiles show a mean ICC of 0.26 (range 0.25–0.27) with all 26 dimensions below the 0.30 threshold and occasion specificity of 74%. LangExtract-derived profiles show mean ICC 0.28 (range 0.25–0.31) with 25 of 26 dimensions below threshold and occasion specificity of 72%. Fused profiles show mean ICC 0.27 (range 0.25–0.30) and 73%. Context shapes expressed psychology 2–3 times more than stable individual differences.
- Robustness of the state finding: ICCs are consistent across 26 diverse constructs (SD = 0.02) and across two independent extraction methods. After residualizing out each subreddit's mean score, mean ICC decreased only slightly from .273 to .266 and all 26 scales stayed below .30.
- Cross-method agreement: Scale-level MTMM agreement is modest (mean r = .06, range .01–.12), but within-profile agreement is high (mean r = .71, median = .76); 69.9% of posts show r > .70 and 91.8% show r > .50. The authors read this as calibration differences rather than disagreement about content.
- Hypothesis tests confirmed: All five literature-derived hypotheses were confirmed by both extraction methods: r/SuicideWatch → elevated neuroticism (SEANCE β = .15, LangExtract β = .30, fused β = .23); r/SuicideWatch → reduced competence (SEANCE β = -.43, LangExtract β = -.38, fused β = -.41); r/depression → elevated neuroticism (SEANCE β = .14, LangExtract β = .39, fused β = .27); r/personalfinance → elevated security (SEANCE β = .16, LangExtract β = .38, fused β = .26); r/personalfinance → elevated achievement (SEANCE β = .15, LangExtract β = .65, fused β = .40).
- Users span multiple archetypes: Clustering the 5,001 posts into k = 6 psychological state archetypes (k-means on z-normalized fused profiles), 94.7% of users express posts in at least two different archetypes and 50.7% appear in three distinct archetypes; only 5.3% show the same archetype across all posts.
- LLMs are state-blind: Prompting GPT-4o, Llama-3.1-8B, and Qwen2.5-14B with 127 questions under 7 conditions (2,667 responses total), models differed significantly in psychological sensitivity (F = 48.31, p < .0001), but not in the expected direction. Mean pairwise semantic similarity: Llama-3.1-8B .768 (SD .068, most sensitive), GPT-4o .819 (SD .074, moderate), Qwen2.5-14B .846 (SD .048, least sensitive). The smallest model was most sensitive.
- Shallow persona detection: Models deviated from the no-profile baseline when any persona was present (mean 20.6%), but failed to differentiate between archetypes (F = 2.18, p = .054). Distressed-Vulnerable and Driven-Assertive users receive essentially identical responses.
- Reward models violate state-invariance in opposite directions: All three reward models (DeBERTa-RM, Skywork-RM-8B, ArmoRM-8B) systematically violate state-invariance with large effect sizes (d > 1.0), explaining 7–30% of score variance. ArmoRM-8B rewards profiles (Cohen's d = +0.76 for Distressed, +0.31 for Driven), while DeBERTa-RM (d = -1.08 and -1.11) and Skywork-8B (d = -1.12 and -1.02) penalize them.
- The vulnerable user paradox: The same Distressed-Vulnerable user is maximally favored by ArmoRM (+0.76) and maximally penalized by Skywork (-1.12), meaning RLHF training would either prioritize or deprioritize vulnerable users depending on arbitrary reward model selection.
Methodology in Plain English
The authors build on Latent State-Trait theory, which says observed behavior is a mix of a stable trait component, a context-specific state component, and measurement error. They operationalize "context" as the subreddit a post appears in, reasoning that subreddits are self-organized communities with distinct norms and topics. From the Webis-TLDR-17 corpus (about 3.8 million Reddit posts from 27,406 subreddits), they select users who posted in at least three distinct subreddits, sample three posts per user, and require posts to be at least 50 words. The final dataset has 5,001 posts from 1,667 users across 645 subreddits (41 subreddits have n ≥ 10 posts; median post length is 186 words, range 50–3,053). The most represented subreddits include AskReddit (n = 1,558), relationships (n = 923), relationship_advice (n = 268), offmychest (n = 198), depression (n = 129), SuicideWatch (n = 43), and personalfinance (n = 49).
Each post is processed through two parallel extraction methods: SEANCE, an open-source rule-based lexicon tool computing 250+ indices (254 lexicon features used here), and LangExtract, an LLM-based structured extraction tool using GPT-4o as the underlying model. Both feature sets are then assessed against 26 psychological scales using GPT-4o, which responds to validated scale items as if it were the post's author. The four frameworks are the Big Five Inventory (5 dimensions), the Schwartz Value Survey (10), Self-Determination Theory scales (5, using items adapted from the Work Preference Inventory and Intrinsic Motivation Inventory), and DOSPERT risk attitudes (6), totaling 171 items. Scores from the two methods are z-normalized per dimension and fused by averaging.
Variance decomposition uses the intraclass correlation coefficient (ICC) from a one-way random effects model treating posts as nested within users; ICC below 0.30 is treated as state-dominant. Validation uses MTMM-style cross-method agreement, five pre-registered-style literature-driven hypotheses tested with linear mixed-effects regression (subreddit as fixed effect, baseline r/AskReddit, random intercepts for users), and k-means clustering into six archetypes. The applications then prompt three LLMs across 7 conditions (six archetypes plus a no-profile baseline) on 127 questions (77 from GlobalOpinionQA plus 50 psychological dilemma scenarios) and score fixed reference responses with three reward models.
Why This Matters
Impact on research: The paper challenges an assumption baked into persona and personalization datasets — that a psychological profile is a fixed user attribute. If roughly three-quarters of psychological variance is contextual, then aggregating a user's posts into one static profile discards most of the signal. The Chameleon dataset makes within-person variation a measurable object of study rather than noise to be averaged away.
Real-world applications:
- Personalized dialogue and advising: Systems that adjust framing and tone based on whether a user is anxious or confident, rather than producing near-identical responses.
- Mental health and emotional support contexts: The r/SuicideWatch and r/depression results show that expressed distress is context-specific, which matters for how support systems interpret and respond to users.
- RLHF and reward model auditing: Because reward models disagree on direction for the same vulnerable user, teams can audit which users their reward signal systematically favors or penalizes.
- Fairness and vulnerable-user protection: The finding that the same Distressed-Vulnerable profile is maximally rewarded by one model and maximally penalized by another raises the risk of inadvertent, undetected bias against vulnerable users.
Industry relevance: Companies using RLHF to tune assistants inherit whatever bias their reward model embeds, without deliberate design choice. The paper argues this arbitrariness propagates into training undetected, and that it may explain why aligned models are state-blind: the reward signal itself is inconsistent about context.
Future Directions
- Establish criterion validity: The authors note that human annotation of profiles was not conducted, and that reliable psychological construct rating requires specialist training; they call for ecological momentary assessment with consenting users.
- Increase observations per user: The current design uses k = 3 observations per user, which is described as producing a likely upper bound on the 26% between-person estimate. The authors suggest k = 10+ to enable more precise variance decomposition.
- Disentangle context components: Subreddits conflate topic, audience, and community norms; future work should separate these factors and control for specific linguistic features rather than subreddit means.
- Improve reward models: The paper states that resolving the problem requires reward models that respond to psychological context in principled, consistent ways, and notes generalization beyond Reddit — plus the fact that the corpus spans 2006–2016 — remains untested.
Target Audience
Researchers and practitioners in NLP personalization, persona-based dialogue, affective computing, and RLHF alignment; psychology researchers interested in applying Latent State-Trait theory to text; and teams building or auditing reward models who need to understand how user context enters — or fails to enter — model behavior. Readers should be comfortable with statistical concepts such as intraclass correlation, mixed-effects models, and effect sizes.
Authors’ abstract
User interactions with language models vary due to static properties of the user (trait) and the specific context of the interaction (state). However, existing persona datasets (like PersonaChat, PANDORA etc.) capture only trait, and ignore the impact of state. We introduce Chameleon, a dataset of 5,001 contextual psychological profiles from 1,667 Reddit users, each measured across multiple contexts. Using the Chameleon dataset, we present three key findings. First, inspired by Latent State-Trait theory, we decompose variance and find that 74% is within-person(state) while only 26% is between-person (trait). Second, we find that LLMs are state-blind: they focus on trait only, and produce similar responses regardless of state. Third, we find that reward models react to user state, but inconsistently: different models favor or penalize the same users in opposite directions. We release Chameleon to support research on affective computing, personalized dialogue, and RLHF alignment.