Research
Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict
Overview Research area: Natural language processing and LLM evaluation, specifically the measurement of value–action alignment in large language models, drawing on psychometrics (IUIPC privacy scales,
- arXiv
- 2601.03546
- Published
- 2026-01-07
- Authors
- Guanyu Chen, Chenxiao Yu, Xiyang Hu
AI summary
Overview
- Research area: Natural language processing and LLM evaluation, specifically the measurement of value–action alignment in large language models, drawing on psychometrics (IUIPC privacy scales, the Prosocialness Scale for Adults, and a data-sharing acceptance instrument) and multi-group structural equation modeling (MGSEM).
- Technical level: Intermediate. The paper explains SEM and MGSEM concepts from the ground up, but readers benefit from some familiarity with LLM prompting protocols, Likert-scale psychometrics, and latent-variable modeling.
- One-sentence scope: The paper proposes a context-based questionnaire protocol and a new human-referenced metric, Value–Action Alignment Rate (VAAR), to test whether ten LLMs reproduce the human-like directional relations by which privacy concern negatively and prosocial motivation positively predict acceptance of data sharing.
What This Paper Is About
When an LLM decides whether to share personal data, two motives can pull in opposite directions: privacy concern (which should reduce sharing) and prosocial motivation (which should increase sharing, for example when sharing supports public health). Existing evaluations tend to measure these attitudes or sharing intentions in isolation, or collapse them into a single "value–action gap" score, which becomes ambiguous when two attitudes point in different directions. The paper's goal is to evaluate alignment relationally: instead of asking whether a model's stated values and actions match on average, it asks whether the estimated relationships between privacy concern, prosocialness, and data-sharing acceptance carry the signs established in human behavioral research.
Key Contributions
- A context-based assessment protocol that sequentially administers standardized questionnaires for privacy attitudes (IUIPC-derived dimensions of Control, Awareness, and Collection), prosocialness (the 16-item PSA), and acceptance of data sharing (AoDS, with subscales for Sacrifice Privacy, Past Acceptance, and Future Willingness) within one bounded, history-carrying session. Up to three prior questionnaire steps are summarized into the next prompt using dimension means, without changing item content or scale.
- Value–Action Alignment Rate (VAAR), a human-referenced directional agreement metric. VAAR uses multi-group structural equation modeling to estimate six focal cross-domain paths (three PSA→AoDS and three Privacy→AoDS) per model, converts each standardized coefficient into a directional probability for the human-expected sign via a normal approximation of the z-score, and averages the resulting path-level log-losses. Smaller VAAR means more probability mass on human-consistent directions.
- An empirical evaluation across ten LLMs — GPT-4o, GPT-4-turbo, GPT-4, GPT-3.5-turbo, GPT-4o-mini (OpenAI API); DeepSeek-R1, Llama3-70B-Instruct, Mistral-7B-Instruct, Titan-Text-Express (AWS Bedrock); and Qwen3-14B (HuggingFace) — with 100 independent runs per model at temperature t = 0.7 under a fixed Privacy → PSA → AoDS order.
- A diagnostic argument about measurement validity: under multi-attitude conflict, gap-style value–action evaluators implicitly select a single reference attitude, so some apparent misalignment may be an artifact of the evaluator rather than a property of the model. The paper also reports cases where the fixed SEM specification cannot be estimated at all, and treats those as
NArather than as alignment or misalignment evidence.
Main Findings
- Cross-model heterogeneity with within-model stability: Models show clearly different Privacy–PSA–AoDS profiles (Kruskal–Wallis tests across models, all p < .001), yet each model maintains a stable mean profile across 100 independent runs. Dispersion concentrated in a subset of models: GPT-4o had the lowest average within-scale SD at 0.73, versus Qwen3-14B at 1.84 and Amazon Titan at 1.52 (GPT-4 1.02, GPT-4-turbo 0.84, GPT-3.5-turbo 0.84, Llama3-70B 0.91, Mistral-7B 0.94, DeepSeek-R1 1.16).
- Human-similar marginals, not uniform means: Human reference levels were Privacy M = 5.84, PSA M = 5.82, and AoDS M = 2.86 on a 1–7 scale. Model Privacy means fell in the same high range (from GPT-4 at 6.43 down to Amazon Titan at 5.13), PSA was systematically lower than the human reference across models (GPT-4 5.76 down to Titan 4.65), and AoDS means spanned both sides of the human baseline (GPT-4 3.69 and Qwen3-14B 3.81 above; Mistral-7B 2.49 and DeepSeek-R1 2.71 below).
- Limited and highly model-dependent value–action alignment: VAAR ranged from strong alignment to severe misalignment, spanning more than an order of magnitude. GPT-4o (0.111), GPT-4-turbo (0.225), and Llama3 (0.234) showed strong alignment; Amazon Titan (0.474) moderate; GPT-3.5-turbo (0.858) and GPT-4o-mini (0.864) weak; Mistral (2.266) and Qwen3-14B (4.914) misaligned. The descriptive tiers used were Strong [0, 0.3), Moderate [0.3, 0.7), Weak [0.7, 1.0], and Misaligned > 1.0.
- Two models were unestimable: GPT-4 and DeepSeek-R1 returned
NAVAAR because the fixed MGSEM specification could not be reliably estimated, attributed primarily to variance-structure collapse (an ill-conditioned covariance/information matrix) rather than to numerical noise or merely low marginal standard deviations. - Dispersion is descriptively linked to divergence but not sufficient: Noisier profiles often coincided with larger VAAR, but Mistral showed strong divergence despite moderate dispersion, and GPT-3.5-turbo and GPT-4o-mini remained only weakly aligned despite relatively concentrated profiles.
- Robustness to sampling and temperature: In a stateless condition with strictly independent, single-item prompts (n = 50 runs per model), cross-run SD and drift were low, with all drift tests non-significant (p > .05) — for example Llama SD 0.13/0.19 (median/max) and drift 0.43/0.57; Titan SD 0.32/0.48 and drift 0.83/1.47. Re-running the full pipeline at t ∈ {0.1, 0.7, 1.0} preserved the qualitative contrast, with Llama at 0.00 (Strong), 0.23 (Strong), 0.58 (Moderate) and Mistral staying above 1 (Misaligned), above 1 (Misaligned), and 0.85 (Weak).
- Order sensitivity: Conclusions were stable when AoDS was elicited last, and swapping the two value scales caused only minor VAAR changes. Stress-test orders that elicited AoDS first substantially increased VAAR for otherwise aligned models such as GPT-4o, GPT-4-turbo, and Llama3, consistent with classic priming effects. Models unestimable under the main order remained
NAor misaligned across orders.
Methodology in Plain English
Each model was put through a single, continuous questionnaire session. First it answered privacy items, then prosocialness items, then data-sharing items, all on a 7-point Likert scale, in a strict "QUESTION_ID: SCORE" format. After each questionnaire, the responses were averaged into dimension means and a short summary of the previous conversation was carried into the next prompt, so the model had a bounded memory of what it had just said. This was repeated 100 times per model at a fixed temperature of 0.7.
The analytical core is a structural equation model. SEM distinguishes a measurement model, which links individual questionnaire items to latent constructs such as privacy concern or prosocialness, from a structural model, which specifies directed relationships among those constructs — essentially a system of regressions that lets several attitudes act on one outcome at once. Multi-group SEM fits the same structure to several groups while allowing group-specific coefficients; here each group is one LLM. The specification was fixed in advance and estimated in lavaan with robust maximum likelihood, mean structures, and the Theta parameterisation.
Six "focal paths" were examined: three from prosocialness to the three AoDS subscales and three from privacy concern to those subscales. The human reference template assigns +1 to every PSA→AoDS path and −1 to every Privacy→AoDS path, based on prior privacy and prosociality research. For each path, the standardized coefficient is divided by its robust standard error to give a z-score, which is converted into a probability that the path has the human-expected sign using the normal CDF. The negative log of that probability is the path loss, and VAAR is the average of these losses across the estimable focal paths. The authors also report that the fixed specification satisfied at least configural invariance between convergent model groups, supporting comparability across models under a shared structural template. They stress that SEM is used here as a controlled structure extractor, not as evidence of causal or psychological mechanisms inside models.
Why This Matters
Impact on research. The paper argues that one-dimensional value–action gap scores are structurally ill-defined when multiple values with opposing effects shape a single behavior: the same sharing choice can be consistent with prosocial motivation and inconsistent with privacy concern simultaneously, so an apparent gap may reflect the evaluator's choice of reference attitude. It reframes alignment as a question about relations among constructs rather than about marginal scores, and provides a reusable human-referenced template and metric for that purpose. It also documents a failure mode — variance-structure collapse — that leaves some models unmeasurable under a fixed specification.
Real-world applications:
- Privacy-preserving deployment decisions, where teams need to know whether a model's stated privacy stance actually predicts its data-sharing recommendations in public-health or surveillance-adjacent scenarios.
- Regulatory and governance assessment, giving auditors a structured instrument for probing how a model resolves trade-offs between individual privacy and collective benefit.
- Model selection for data-handling agents, where the choice of model may hinge on whether privacy concern reliably suppresses sharing behavior.
- Alignment training and evaluation pipelines, offering path-level diagnostics that localize which value–action link deviates, rather than reporting a single aggregate number.
Industry relevance. The finding that human-like marginal attitudes do not guarantee human-aligned value–action relations matters for organizations that screen models using attitude surveys alone. The reported robustness checks (stateless prompting, temperature variation, questionnaire ordering) give practitioners a sense of which results are stable model properties and which are protocol artifacts. The ethics statement cautions against reading the results as a normative endorsement — for example, that more privacy-sacrificing models are preferable because they support public goods — and notes that structurally inferred associations are not a substitute for formal privacy or safety guarantees.
Future Directions
- Extend the MGSEM-based alignment framework to other value conflicts beyond privacy versus prosociality.
- Explore causal interventions on value–action relations, moving from the paper's explicitly non-causal structural extraction toward tests of what changes those relations.
- Study how training and alignment methods influence relational coherence, given the authors' suggestion that model-specific training and alignment pipelines plausibly shape the observed heterogeneity.
- Broaden the operationalisations and settings: other instruments for privacy and prosociality, other languages, other prompts, and deployment conditions. The authors also note that a finite, evolving set of models was evaluated at specific points in time, and that independent assessments revealed limited test–retest reliability for some constructs — an open issue for large-scale stateless prompting.
Target Audience
Researchers working on LLM alignment, evaluation, and value-consistency measurement will find the metric and the critique of gap-based scores most directly useful. Behavioral and privacy researchers interested in applying psychometric instruments and SEM to machine respondents are a second core audience, as are practitioners who need to assess data-sharing behavior in deployed models. Readers without a background in latent-variable modeling can follow the results tables but will need the paper's SEM primer to engage with the method.
Authors’ abstract
Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concerns and prosocial motivations can push choices in opposite directions. Existing evaluations often measure privacy-related attitudes or sharing intentions in isolation, which makes it difficult to determine whether a model's expressed values jointly predict its downstream data-sharing actions as in real human behaviors. We introduce a context-based assessment protocol that sequentially administers standardized questionnaires for privacy attitudes, prosocialness, and acceptance of data sharing within a bounded, history-carrying session. To evaluate value-action alignments under competing attitudes, we use multi-group structural equation modeling (MGSEM) to identify relations from privacy concerns and prosocialness to data sharing. We propose Value-Action Alignment Rate (VAAR), a human-referenced directional agreement metric that aggregates path-level evidence for expected signs. Across multiple LLMs, we observe stable but model-specific Privacy-PSA-AoDS profiles, and substantial heterogeneity in value-action alignment.