Skip to content
AI.info

Research

Collective Bias Mitigation via Model Routing and Collaboration

Collective Bias Mitigation via Model Routing and Collaboration Overview Research area: Natural Language Processing — fairness and bias mitigation in large language models, specifically multi-model col

Collective Bias Mitigation via Model Routing and Collaboration
arXiv
2610.03240
Published
2026-10-02
Authors
Mingzhe Du, Luu Anh Tuan, Xiaobao Wu, Yichong Huang, Yue Liu, Dong Huang, Huijun Liu, Bin Ji, Jie M. Zhang, See-Kiong Ng

AI summary

Collective Bias Mitigation via Model Routing and Collaboration

Overview

  • Research area: Natural Language Processing — fairness and bias mitigation in large language models, specifically multi-model collaboration and routing.
  • Technical level: Intermediate. The framework concepts (routing, voting, debate, committee) are accessible, but the bias metrics, bootstrap analysis, and topology formulations assume some familiarity with LLM evaluation.
  • Scope: The paper introduces CrowdEval, a dataset of per-query model bias behavior, and Collective Bias Mitigation (CBM), a framework that routes each query to a selected set of diverse open-source LLMs and organizes them into topologies (Sequential, Voting, Debating, Committee) to produce less biased responses.

Note: the paper opens with a content warning that it contains explicit statements of offensive or upsetting language.

What This Paper Is About

LLMs trained on large web corpora can absorb and even amplify stereotypes, which is a problem when they are used in public health, finance, and governance. Existing approaches such as self-debiasing ask a single model to find and fix its own bias, but the paper argues that a model's intrinsic knowledge is often insufficient for deeply ingrained stereotypes, and that without external supervision models may even use stereotypical knowledge to justify their answers. The goal is to instead learn fine-grained per-query behavior across a large pool of LLMs and then let several of them collaborate so that their collective output is more neutral than any single model's.

Key Contributions

  1. CrowdEval benchmark. A new dataset for evaluating fine-grained bias in LLM responses, built by collecting responses from more than 50 open-source LLMs to bias-eliciting questions drawn from the ambiguous subset of the BBQ dataset, with a per-response bias label (bias-target / non-target / neutral). It covers eight social dimensions, with sizes of 1,024 (Age), 1,024 (Gender), 778 (Disability), 1,024 (Nationality), 1,024 (Race), 600 (Religion), 1,024 (SES), and 432 (Sexual Orientation); dimensions marked with an asterisk contain fewer instances in BBQ, so all available questions were included.
  2. The CBM framework. Described by the authors as the first collective LLM debiasing framework that synergizes the knowledge of diverse LLMs to mitigate their holistic bias, combining an LLM-based model router with several coordination topologies.
  3. Extensive experimental evaluation. Comprehensive experiments over 50 leading LLMs assessing the effectiveness of CBM across multiple social dimensions.
  4. An LLM-based model router. Rather than training a dedicated classifier on a lightweight model such as BERT or T5 from scratch (as most model-selection work does), the authors fine-tune a pre-trained LLM as the router, replacing model names with unique identifiers (e.g., model_{index}) to prevent overfitting to dominant names like "Llama" or "Qwen", and ranking candidates by predicted token probability.

Main Findings

  • Routing accuracy scales with router size and then stabilizes. Micro accuracy across the eight dimensions rose from 0.424 (1B) to 0.665 (3B), 0.801 (9B), 0.831 (14B), and 0.851 (32B). The 32B router (Qwen-2.5-32B) achieved the highest accuracy of 0.851. Routing performance stabilized once router parameters exceeded 9B.
  • Routing precision beats random selection. Overall micro precision was 0.471 for random selection versus 0.582 (1B), 0.651 (3B), 0.804 (9B), 0.883 (14B), and 0.941 (32B). Precision did not increase linearly with model size, with diminishing improvements after 9B.
  • The router generalizes to unseen bias dimensions. With SES and SO excluded from router training, the 32B router reached accuracy of 0.883 (SES) and 0.809 (SO), and precision of 0.785 (SES) and 0.781 (SO), both substantially above random selection of 0.125. The authors describe this as notable zero-shot generalization once the router reaches sufficient scale (9B or above).
  • Sequential topology struggles and can compound bias. Each model's response feeds directly into the next, and the bias score increases as chain length grows, highlighting the risk of compounding bias from earlier models.
  • Voting gives a stable improvement. Despite its simplicity, Voting consistently outperformed the Single baseline across the eight social dimensions by averaging multiple responses and diluting individual biases; it can achieve better performance under the model routing setting.
  • Debating achieves lower bias at high cost. Iterative exchange drives down overall bias scores, but the Debating topology requires approximately 27 times more computational resources than the Single baseline.
  • Committee balances mitigation and cost. Although Debating often achieves the lowest absolute bias score, Committee showed more consistent results (tighter variance) with lower model inference cost. In the top-7 setting, Committee lowered the age bias score from 0.25 for the Single baseline to 0.10.
  • CBM outperforms self-debiasing on larger models. In the comparison reported in the paper's Table 7, average bias scores were: Qwen2.5-32B-Instruct 0.114, Llama-3.3-70B 0.104, DeepSeek-R1-Distill-Llama-70B 0.199, and CBM (ours) 0.095.
  • Model ordering matters in the Sequential topology. Reversing the router-recommended order (which runs from less biased to more biased) changes results noticeably; placing less biased models later in the sequence appeared to make CBM more resilient to earlier, potentially more biased decisions.
  • Individual models vary substantially. The per-model bias scores reported across the eight dimensions show wide variation, including negative bias scores indicating anti-stereotypical polarity.

Methodology in Plain English

The researchers start by building a picture of how each model behaves. They take bias-eliciting questions from the ambiguous subset of BBQ — scenarios that deliberately lack enough information to decide the answer, so any preference for the stereotyped option reveals implicit bias. Each question is presented with a context, a question, and three answer choices (bias-target, non-target, and neutral). They send every question to a pool of more than 50 open-source text-generation models from HuggingFace, spanning roughly 1 billion to 56 billion parameters and varied architectures and training corpora, generating responses with greedy decoding. Each response gets a bias label, producing the CrowdEval dataset.

Next, they fine-tune a pre-trained LLM (Qwen2.5-32B in the main setup, with routers also tested at 1B, 3B, 9B, and 14B) on CrowdEval using an Adam optimizer, one epoch, a learning rate of 5×10⁻⁵. The router is trained on two tasks: identifying the social dimension of a query (bias detection) and picking the models most likely to answer neutrally (model selection). At inference, the router reads the query and ranks candidate models by their token probabilities.

The chosen models are then arranged into one of several communication topologies and asked to produce a final answer:

  • Single — the top-ranked model answers alone, the baseline.
  • Sequential — models pass responses down a chain, each seeing prior responses; self-debiasing is a special case of this topology with the same model repeated.
  • Voting — models answer independently in parallel and a majority vote decides.
  • Debating — models answer independently, then iterate by exchanging all responses until consensus is reached.
  • Committee — a designated coordinator model collects the other models' responses, drafts a consolidated motion, and seeks approval until consensus, with the coordinator always placed at m₀ and a consensus threshold of 50%.

Evaluation uses BBQ as the benchmark and a Bias Score adapted from BBQ, which multiplies the proportion of non-neutral responses by the tendency of those responses toward the bias target; positive scores indicate stereotypical polarity and negative scores anti-stereotypical polarity. Disambiguated BBQ instances are excluded so the focus stays on inherent bias rather than bias-versus-reasoning. Router quality is measured by accuracy (correct social-dimension classification) and precision (the proportion of recommended models that give neutral responses), with bootstrap sampling over 512 iterations to estimate uncertainty.

Why This Matters

Impact on research. The paper reframes debiasing as a collective, routing-based problem rather than a single-model property, and it argues that existing benchmarks only report a holistic bias score per model, obscuring fine-grained behavior. CrowdEval's per-query, per-model structure makes it possible to study which models are neutral on which kinds of queries, and the finding that routers above 9B generalize zero-shot to held-out bias dimensions suggests a scalable path for detecting bias categories that were never labeled.

Real-world applications:

  • Public health — reducing stereotyped assumptions in triage or patient-facing advice, a sector the authors explicitly cite as a deployment context.
  • Financial services — the paper cites financial services deployments where biased defaults could affect lending or advisory interactions.
  • Governance and public administration — the paper names governance as a domain requiring both accuracy and societal value alignment (citing Aaronson, 2023 and Duan et al., 2025).
  • Workplace and hiring tools — the paper opens its motivation with prevailing workplace gender bias, implying HR and recruitment screening as affected uses.

Industry relevance. Because the model pool is entirely open-source and the framework is described as designed for scalability with seamless integration of additional LLMs, the approach is practical for teams that already run multiple models and can afford routing overhead. The Committee topology's trade-off between bias reduction and inference cost is directly relevant to anyone balancing fairness targets against serving costs; the reported roughly 27× cost of Debating relative to Single makes that balance a concrete engineering decision.

Future Directions

  • Cultural and linguistic scope. The authors note that BBQ is constructed in English and grounded in United States cultural and societal norms, so its framing of social bias may not apply universally. Extending CrowdEval to other languages and cultural contexts is an open problem.
  • Reducing the cost of the strongest topology. Debating achieves the lowest bias scores but costs approximately 27 times the Single baseline; the paper does not report figures on how to close that gap beyond pointing to Committee as a compromise.
  • Router scale and zero-shot coverage. The router generalizes to SES and SO when they are excluded from training, but the paper does not report how far this extends to other entirely unseen bias categories, nor whether many held-out dimensions can be handled simultaneously.
  • Pool composition and openness. All selected models are open-source, ranging from 1 billion to 56 billion parameters; whether the same routing and collaboration benefits transfer to closed or much larger proprietary models is not reported.
  • Topology design. The paper presents a set of topologies but leaves open the question of whether other coordination structures, or adaptive selection of a topology per query, would further improve the bias-reduction versus cost trade-off.

Target Audience

This paper is most useful for machine learning researchers and engineers working on LLM fairness, evaluation, and multi-model systems; practitioners responsible for deploying LLMs in regulated or high-stakes sectors such as health, finance, and governance; and benchmark or dataset builders interested in fine-grained, per-query bias measurement rather than single aggregate scores. Readers who want a purely theoretical treatment will not find one here — the contribution is empirical, dataset-driven, and systems-oriented.

Authors’ abstract

Large language models (LLMs) are increasingly deployed in public health, finance, and governance, requiring both accuracy and societal value alignment. Despite recent advances, LLMs often perpetuate or amplify bias embedded in their training data, posing challenges to fairness. While self-debiasing encourages an LLM to identify and correct its own biases, relying on a single model's intrinsic knowledge may be insufficient to address deeply ingrained stereotypes. To address this limitation, we introduce Collective Bias Mitigation (CBM), a framework that alleviates bias by learning fine-grained model behavior and fostering knowledge sharing among diverse LLMs. This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses. Experiments show CBM substantially outperforms standalone baselines (e.g., in the top-7 setting, Committee lowers the age bias score from 0.25 to 0.10). Our Debating and Committee topologies achieve substantial bias reduction, with the latter balancing mitigation effectiveness and inference cost, highlighting the potential of CBM for fairer LLMs.

Read the original paper