Research
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors
Overview Research area: Natural Language Processing, specifically LLM persona simulation applied to legal reasoning (common-law jury simulation), with cross-disciplinary ties to legal psychology and A
- arXiv
- 2609.09887
- Published
- 2026-09-09
- Authors
- Cho-Ying Wu
AI summary
Overview
- Research area: Natural Language Processing, specifically LLM persona simulation applied to legal reasoning (common-law jury simulation), with cross-disciplinary ties to legal psychology and AI bias evaluation.
- Technical level: Intermediate. Readers need basic familiarity with LLM prompting, regression-based effect analysis, and U.S. criminal-law concepts such as justification defenses and mens rea, but the paper is written accessibly.
- Scope: A systematic benchmark study of how 20 frontier LLMs behave as simulated jurors in 500 controversial U.S. criminal cases, measuring how defendant courtroom statements, background affinity, and juror ideology shift verdict severity.
What This Paper Is About
Courts in the U.S. common-law system rely on lay juries to decide facts and verdicts, and lawyers increasingly use mock juries to test trial strategies. This paper asks whether LLMs can credibly stand in for those mock jurors, and specifically how a defendant's courtroom statement changes an LLM juror's decision. The goal is to determine which factors — emotional appeal, expressed remorse, or shared background between juror and defendant — most strongly move LLM verdicts, and whether those patterns resemble findings from human legal psychology.
Key Contributions
- The first systematic study of LLM-simulated juries under the common-law system, focused on emotional persuasion as the central courtroom dynamic, in contrast to prior civil-law courtroom simulations built from post-verdict judgment documents.
- JuryBench, a benchmark of 500 controversial criminal cases spanning roughly 400 potential charges under the U.S. Criminal Code, with controlled defendant backgrounds (stereotype-conforming "Group-A" and stereotype-subverting "Group-B"), long-form defendant statements, and 12 diverse juror profiles per case.
- Large-scale evaluation of 20 frontier LLMs, producing 432,000 juror decisions with written rationales, enabling per-model behavioral profiling of sensitivity to emotion, remorse, and background affinity.
- A human-evaluation anchor on 25 selected cases with 12 U.S. lay participants, used to check whether the LLM reasoning patterns resemble human juror reasoning.
Main Findings
-
Statements rarely change verdicts, but when they do, they usually backfire. Roughly 70–80% of juror decisions were unchanged after hearing the defendant's statement. This "decision inertia" mirrors human-jury research suggesting most jurors form impressions during opening remarks. Among models that did shift, most became harsher rather than more lenient. Individual model results ranged widely: Gemini 3 Flash showed the largest severity-reducing effect (SE = +0.55), while Qwen3 showed the largest severity-increasing effect (SE = −0.59).
-
Emotional appeal is a double-edged sword. Higher emotional contagion reduced severity for a modest majority of models, but increased it for others. Jurors sometimes read intense emotion as performative, inconsistent, or a signal of guilt — one Claude Haiku juror reasoned that the defendant's "charismatic speech" did not overcome the financial evidence.
-
Remorse helps some models and hurts others. DeepSeek V4, Gemini, and some GPT models treated expressed remorse as mitigating. Grok and several GPT models treated it as an implicit admission of responsibility. For about half the models, remorse had no measurable effect.
-
Background fit is the strongest and most consistent driver. A "match score" computed from the product of defendant affinity and juror ideology showed significant positive coefficients for a substantial majority of models, and no model showed a significant negative effect. Jurors were systematically more lenient toward defendants whose backgrounds aligned with their ideological leanings, and harsher toward mismatched defendants.
-
Group-A versus Group-B defendants diverge. Stereotype-conforming defendants received lower severity from ideologically aligned jurors, while stereotype-subverting defendants were treated more harshly by mismatched jurors. Each model showed a consistent internal tendency to widen or narrow this gap after statements, and the gap moved more strongly in mismatched cases.
-
Ideology shapes baseline severity. Conservative-identified jurors consistently assigned harsher verdicts than neutral jurors, who in turn were harsher than liberal jurors. Strong ideologues showed the strongest effects, and conservative jurors were roughly 70% more likely to change severity after the defendant's statement.
-
Joint effects reveal coherence matters. In an interaction regression, isolated emotion or isolated remorse often backfired, but the combination of emotion plus remorse plus background affinity produced the strongest severity reduction. The authors interpret this as jurors rewarding coherent persuasive narratives rather than one-dimensional appeals.
-
Human and LLM patterns broadly align. Human participants also became harsher on average after hearing statements (SE = −0.29), citing remorse as evidence of guilt and performative delivery as reasons — the same rationales appearing in LLM juror outputs.
Methodology in Plain English
The authors generated 500 controversial criminal scenarios using GPT-5.4, each containing a defendant background, a case background, and a list of evidence. For every case they created two contrasting defendants: one with a background matching common stereotypes (Group-A) and one deliberately subverting those stereotypes (Group-B), for example a member of a marginalized group who is actually wealthy and successful. Each defendant also produced a roughly 15-sentence courtroom statement designed to appeal to sympathy, justify conduct, or express remorse. A practicing lawyer reviewed all materials for legal plausibility and rated each statement on emotional contagion and remorse intensity.
Twelve simulated jurors were created per case, each with an ideology score from strongly conservative to strongly liberal. Every juror made a verdict choice — not guilty, or guilty of a specific charge from a provided list — both with and without the defendant's statement. The researchers then mapped each possible charge onto a numerical severity score (0 for acquittal or valid justification, up to 15 for the most serious felony) with expert labeling.
This design yielded 960 predictions per case (20 models × 12 jurors × 4 statement/background conditions). The authors quantified a "statement effect" as the severity drop after hearing the statement, and used ordinary least squares regression with three standardized predictors — emotional contagion, remorse, and background fit — to identify which factor mattered most for each LLM. A second interaction model tested how those factors combine. To validate the setup, they measured how diverse the verdicts were across models: the most common charge accounted for only about 65% of decisions on average, and only 3 of 500 cases produced unanimous verdicts, confirming the cases were genuinely contested. A human study with 25 cases and 12 lay participants served as a comparison point.
Why This Matters
This work sits at the intersection of legal practice, AI evaluation, and legal psychology. It provides the first controlled testbed for asking whether LLM "mock juries" reproduce the well-documented biases of real juries — or invent new ones. The finding that background affinity dominates emotional appeal is consequential for anyone hoping to use LLMs as neutral trial-strategy tools, because it suggests the models inherit affinity-driven skews that lawyers may accidentally exploit or be misled by.
Real-world applications:
- Mock-jury and focus-group simulation. Law firms already use LLMs informally to rehearse trial strategies; JuryBench provides a structured method for measuring how much a given model's feedback is driven by bias rather than case facts.
- Jury-selection research. The background-fit findings offer a controlled way to study how voir dire and peremptory challenges interact with juror-defendant demographic alignment, without the ethical and logistical burdens of live human mock juries.
- LLM bias auditing in high-stakes domains. The framework generalizes beyond law: any decision task involving persona-conditioned judgments (hiring, medical triage, credit) can be probed with the same controlled stress-testing approach.
- Legal AI governance and disclosure. Regulators and bar associations evaluating whether to permit LLM-assisted trial preparation now have empirical evidence about what these systems actually do and where they fail.
Industry relevance centers on legal-tech vendors, AI safety teams, and any organization deploying LLM agents in adversarial or evaluative roles. The paper explicitly warns that the same sensitivity to emotional framing that makes LLMs plausible juror simulators also makes them vulnerable to manipulation.
Future Directions
- Simulate full trial dynamics. The current pipeline compresses opening statements, examination, and evidence presentation into a single stage, and omits cross-examination and the jury's group deliberation toward a unanimous verdict. Adding these could substantially change findings, especially since real deliberation can override individual bias.
- Extend beyond text. Emotional contagion and remorse are partly conveyed through body language, tone, and facial expression. Multimodal courtroom simulation is an obvious next step that the authors flag as outside their current scope.
- Broaden jurisdictional and demographic coverage. The work is limited to U.S. common-law criminal cases with synthesized profiles; testing civil, international, or non-U.S. systems would clarify whether the affinity and ideology effects generalize.
- Track model evolution and validate against real juries. The 20-LLM snapshot will date quickly as providers retrain models. Longitudinally tracking bias signatures and pairing them against large, representative human jury studies would test whether LLM-juror simulation is genuinely predictive or merely superficially similar.
Target Audience
Legal-tech researchers and developers building courtroom simulation or mock-jury tools; NLP researchers studying persona simulation, agent bias, and persuasion; legal psychologists interested in whether LLM behavior replicates human jury findings; AI safety and governance teams auditing high-stakes decision systems; and practicing attorneys or litigation consultants evaluating whether LLM mock juries are trustworthy enough to inform trial strategy.
Authors’ abstract
LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias, and background-based affinity. To support the analysis, we introduce JuryBench, a benchmark containing controversial criminal cases in U.S. criminal law. In each case, a defendant can claim various plausible justifications to support acquittal or reduced liability. We fix the base case and design defendants of different backgrounds, who give courtroom statements with varying emotional appeal or rebuttal. Jurors with diverse ideological profiles across the spectrum are simulated. We examine 20 frontier LLMs, resulting in a total of 432K decisions and rationales, and quantify changes in verdict severity. Our findings show that LLM-jury simulation echoes many human-jury findings. First, emotional persuasion can be detrimental, since jurors may perceive it as evidence of guilt or inconsistency. Next, we show that background fit between jurors and defendants is a stronger and significant factor than other isolated factors, and that jurors are in general harsher toward opposite-background defendants and lenient toward same-background ones. Finally, we find that juror ideology also strongly shapes severity judgments. These findings highlight both the promise and risks of using LLMs to model jury reasoning and call for careful evaluation. The data and code are available at https://github.com/choyingw/JuryBench