Skip to content
AI.info

Research

DeFrame: Debiasing Large Language Models Against Framing Effects

Overview Research area: Natural language processing — fairness evaluation and bias mitigation for large language models (LLMs), specifically the study of framing effects in fairness contexts. Technica

DeFrame: Debiasing Large Language Models Against Framing Effects
arXiv
2602.04306
Published
2026-02-04
Authors
Kahee Lim, Soyeon Kim, Steven Euijong Whang

AI summary

Overview

Research area: Natural language processing — fairness evaluation and bias mitigation for large language models (LLMs), specifically the study of framing effects in fairness contexts.

Technical level: Intermediate. The core ideas (a prompt changes wording but not meaning, and the model's bias score shifts) are intuitive, but the paper operates with formal bias metrics, logit-based discrimination scores, and LLM-as-judge evaluation.

Scope: The paper (arXiv:2602.04306v2, cs.CL, by Kahee Lim, Soyeon Kim, and Steven Euijong Whang at KAIST) defines a metric called "framing disparity," measures it on 8 instruction-tuned LLMs across three augmented fairness benchmarks, and proposes DeFrame, a prompting-based debiasing framework that reduces both overall bias and framing-induced inconsistencies.

What This Paper Is About

LLMs often look fair under standard fairness benchmarks but behave inconsistently or unfairly once those benchmarks are rephrased — a form of hidden bias. The authors identify "framing" (expressing the same stereotype or decision in a positive versus negative way, e.g., "A is better than B" vs. "B is worse than A") as an underexplored cause of this gap, and their goal is to quantify that instability and then reduce it. Their answer is a three-stage prompting method, DeFrame, that forces a model to look at the opposite framing, write fairness guidelines, and revise its own initial answer.

Key Contributions

  1. The concept of framing disparity. The authors formalize framing disparity (FD) as the difference between a model's bias under a positive framing and its bias under a negative framing, using whatever bias metric a benchmark already defines, and note that FD is bounded by that underlying metric (bounds given in the paper's Appendix E.1).

  2. A framing-augmented evaluation suite and a systematic measurement. They augment three fairness benchmarks with alternative framings — BBQ (multiple-choice stereotype QA), DoNotAnswer (open-ended harmful-response generation, producing "DoNotAnswer-Framed"), and 70Decisions (yes/no decision-making, producing "70Decisions-Framed") — and report framing disparity for 8 instruction-tuned LLMs, plus five additional models in the 30–70b range in the appendix.

  3. Evidence that existing debiasing does not fix framing disparity. They show that current prompting-based debiasing methods (PR, IF-BASE, IF-CoT, TFS-PP, TFS-SR, TFS-IP, SD-EXP, SD-REP) improve frame-averaged fairness but often fail to reduce, or can worsen, the gap between framings.

  4. DeFrame, a framing-aware debiasing framework. DeFrame explicitly incorporates an alternative framing through three stages — Framing Integration, Guideline Generation, and Self-Revision — and is reported to reduce framing disparity by 92% and bias score by 93% on average on BBQ, while using 4 LLM calls per question.

Main Findings

  • Framing shifts bias substantially on BBQ. Bias under negative framings is on average 2x larger than under positive framings, reaching up to 4x in some demographic categories. The magnitude varies by category: LLaMA-3.2-3b-Instruct shows an FD of -41.388 on disability status versus -7.772 on race/ethnicity.

  • Vulnerability direction is task-dependent, not random. On BBQ, models are more vulnerable under negative framing. On DoNotAnswer-Framed, the reversal occurs: models produce more harmful responses under positive framing, which the authors attribute to positively-framed prompts making harmful stereotypes harder to recognize. On 70Decisions-Framed, positive framing tends to produce more favorable decisions for minority groups, an effect that weakens under negative framing.

  • Framing can flip which group a model favors, not just how much. In the 70Decisions-Framed non-binary category, LLaMA-3.1-8b and Gemma-3-12b keep similar bias magnitudes (≈0.25) when framing changes, but the group they favor differs: LLaMA continues to favor the minority group while Gemma flips to favor the majority group, reversing the sign of the bias.

  • Larger models show smaller and more stable bias overall. The paper reports that the overall magnitude of bias tends to decrease as model size increases, and framing disparity generally becomes smaller and more stable in larger models, though not uniformly across every setting.

  • Baselines reduce average bias but leave framing disparity. On BBQ, all compared methods reduce framing disparity, with DeFrame improving the most across categories. On DoNotAnswer-Framed, baselines lower the harmful response rate but the disparity remains positive. On 70Decisions-Framed, baselines are unstable — sometimes improving, sometimes worsening — while DeFrame is reported to improve consistently.

  • Reasoning and multi-perspective methods outperform instruction-only ones. Methods incorporating reasoning or multiple perspectives generally outperform those relying only on debiasing instructions or self-revision, and baselines that perform well on BBQ often lose effectiveness on DoNotAnswer-Framed and 70Decisions-Framed.

  • All three DeFrame components are needed for stable framing-robustness. Progressive ablations on LLaMA-3.1-8b show that adding Self-Revision cuts bias relative to the plain model, and Guideline Generation adds further gains, with the largest reductions often coming from the earlier steps; however, only the full DeFrame configuration consistently reduces framing disparity across framings.

Methodology in Plain English

The authors start by defining a way to measure inconsistency: take a fairness benchmark, rewrite its prompts in a positive framing and in a negative framing, score the model's bias separately on each set, and subtract the two scores. That difference is the framing disparity. A model that reports the same bias level regardless of wording gets a value near zero.

They then pick three benchmarks chosen to cover different fairness concepts and task types: BBQ for multiple-choice stereotype questions (they use the ambiguous setting, where "Unknown" is the correct answer, across seven demographic categories), DoNotAnswer for open-ended generation scored by an LLM judge that computes a harmful response rate, and 70Decisions for yes/no decisions scored with a logit-based discrimination score relative to a majority baseline. For DoNotAnswer-Framed, they take 95 stereotype-related prompts, find 52 that can be inverted, and paraphrase each prompt 4 times to reach 520 prompts. For 70Decisions-Framed, they use the explicit version with 9,450 questions (gender and race), generate opposite-framing counterparts, and end up with 18,900 questions. When negative framings are introduced, they treat "no" as the favorable decision rather than "yes."

To debias, DeFrame mimics the dual-process idea of a slow, deliberative "System 2" step. Given a prompt, the model first answers normally (its intuitive "System 1" answer). It then detects the prompt's framing polarity and rewrites the prompt into the opposite framing, uses both versions to write a short fairness guideline, and finally revises its initial answer using that guideline. This costs 4 LLM calls per question, compared with 1 to 3 calls for the baselines. The authors compare DeFrame and the baselines on the same three benchmarks, then run ablations that add one component at a time — plain model, then Self-Revision, then Guideline Generation, then Framing Integration — using four instruction-tuned models (LLaMA-3.1-8b, Qwen-2.5-7b, Gemma-3-4b, Mistral-7b).

Why This Matters

Fairness evaluations that use a single wording can report a model as fair while the same model behaves differently under a rephrased prompt. This means reported fairness scores may not transfer to deployment, and the failure mode is invisible to standard evaluation. The paper's framing disparity metric gives researchers a way to check for that instability without discarding existing bias metrics.

Real-world applications:

  • Screening and allocation decisions. The paper explicitly notes that real-world decisions such as screening or allocation may change solely because of prompt wording, since framing reversed which group a model favored in 70Decisions-Framed.
  • Hiring and role assignment. Prior work cited by the authors (Bai et al., 2024) shows LLMs appearing unbiased when directly asked about stereotypes but showing bias in practical tasks such as role allocation.
  • Safety filtering and content moderation. DoNotAnswer-Framed measures whether a model produces harmful responses; positive framings made harmful stereotypes harder for models to recognize, a risk for moderation systems that keyword-match on tone.
  • Automated decision-making in general. Any pipeline where a decision template can be reworded — loan, admissions, service triage — is exposed to the inconsistency the paper measures.

Industry relevance: the method is prompting-based, meaning it is model-agnostic and needs no training data or parameter updates, so it can be dropped into inference pipelines. The trade-off is cost: DeFrame requires 4 LLM calls per question, which the authors flag as a limitation.

Future Directions

  • Beyond binary framings. The paper focuses on positive versus negative polarity and only briefly discusses extending to multiple framings in Appendix D, leaving richer linguistic variation for future work.
  • Tracing bias to its roots. DeFrame reduces bias without disentangling where it comes from; the authors identify training data and training dynamics as likely sources requiring deeper analysis of corpora and internal representations.
  • Cheaper and more robust debiasing. Multiple LLM calls per question add overhead, and the authors note that adversarial prompting may expose or amplify underlying biases and potentially circumvent mitigation, motivating more efficient and robust framing-aware strategies.
  • Larger models and intersectional bias. Main analyses focus on 3b–14b models due to computational constraints, with 30–70b models examined only at a high level, and evaluation does not capture intersectional or context-specific biases arising from more complex social group interactions.

Target Audience

This paper suits LLM fairness and AI safety researchers who evaluate bias benchmarks, NLP practitioners and prompt engineers who need model-agnostic debiasing at inference time, and teams deploying LLMs in screening, allocation, or moderation settings where decision templates can be reworded. Readers wanting a single fairness number per model will find the framing-disparity framing of the problem the most transferable idea; readers looking for a training-free mitigation method will find DeFrame directly applicable, with the 4-call cost noted.

Authors’ abstract

As large language models (LLMs) are increasingly deployed in real-world applications, ensuring their fair responses across demographics has become crucial. Despite many efforts, an ongoing challenge is hidden bias: LLMs appear fair under standard evaluations, but can produce biased responses outside those evaluation settings. In this paper, we identify framing -- differences in how semantically equivalent prompts are expressed (e.g., "A is better than B" vs. "B is worse than A") -- as an underexplored contributor to this gap. We first introduce the concept of "framing disparity" to quantify the impact of framing on fairness evaluation. By augmenting fairness evaluation benchmarks with alternative framings, we find that (1) fairness scores vary significantly with framing and (2) existing debiasing methods improve overall (i.e., frame-averaged) fairness, but often fail to reduce framing-induced disparities. To address this, we propose a framing-aware debiasing method that encourages LLMs to be more consistent across framings. Experiments demonstrate that our approach reduces overall bias and improves robustness against framing disparities, enabling LLMs to produce fairer and more consistent responses.

Read the original paper