Skip to content
AI.info

Research

BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation

Overview Research area: Multi-agent systems and large language model (LLM) consultation — specifically, how a central model should decide whether to trust, weight, or ignore advice from peer models. T

BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation
arXiv
2609.35551
Published
2026-09-28
Authors
Peilin Feng, Zhengyang Huang, Soujanya Poria

AI summary

Overview

Research area: Multi-agent systems and large language model (LLM) consultation — specifically, how a central model should decide whether to trust, weight, or ignore advice from peer models.

Technical level: Intermediate to Advanced. The paper combines Bayesian linear regression, Kalman-style online posterior updates, and attention-logit modification inside a frozen LLM, so readers should be comfortable with probabilistic modelling and transformer attention.

Scope: The paper proposes and empirically evaluates BaRe-Mem, an online Bayesian reliability memory that estimates the trustworthiness of advisor responses (and the central model's own competence) from verified interaction history, and uses those estimates both to reweight advisor responses and to decide between consulting and reasoning autonomously.

What This Paper Is About

When multiple LLMs collaborate, advisor models can help or hurt: a model that is reliable on one task may fail on another, and misleading advice can cause a central model to abandon an initially correct answer. The paper's goal is to give a central model a persistent memory that estimates how reliable each advisor is for the current question, and — critically — that also estimates whether consulting anyone at all is better than simply reasoning alone. BaRe-Mem addresses both questions with a single Bayesian memory updated from verified correctness outcomes.

Key Contributions

  1. A Bayesian reliability memory conditioned on internal belief representations. BaRe-Mem encodes each candidate answer (advisor responses plus the central model's own autonomous answer) using hidden representations from the frozen central model, splitting them into a question belief ψ_q(t) shared across candidates and an answer content belief ψ_c(t,k) specific to each candidate. A linear predictor then decomposes reliability into a source-reliability term, an answer-content-reliability term, and a bias.

  2. Exact online updates via rank-one Kalman updates. Verified correctness outcomes are modelled with Bayesian linear regression, and the posterior is maintained exactly through a Kalman gain update rather than recomputing the posterior from scratch, so the memory can scale with an expanding task stream.

  3. Reliability-guided attention with no trainable parameters. Estimated advisor reliabilities are injected as an additive term β_{t,k} = γ·log(p_{t,k} / max_{k'} p_{t,k'}) into the softmax attention weights, leaving the most reliable advisor unchanged and progressively downweighting less reliable ones. The paper states this introduces no trainable parameters or additional training.

  4. An explicit consult-or-reason decision rule. The method models consultation ability as an interpolation A(T) = T·ρ + (1−T)·(κ − δ) between a "reliable evidence" regime and an "unreliable evidence" regime, estimates κ (autonomous ability) and T (trust) directly from the memory, learns ρ and δ from verified history, and consults only when A(T_t) ≥ κ_t — yielding a threshold T_t* = δ̂ / (ρ̂ + δ̂ − κ_t).

  5. Extension to worker routing in agent teams. The same mechanism is repurposed from reweighting a fixed candidate set to selecting which worker receives a sub-task, evaluated on MuSiQue.

Main Findings

  • Baseline comparisons across two capability regimes. Experiments span nine benchmarks and six central models, with Qwen3-14B and Phi-4 used as representative central models in the main text. In the capability-challenging regime, both Question + Peers and Debate (2 rounds) fall below the No consultation baseline once the misleading-information ratio exceeds 50% for both central models. Majority voting is described as substantially more vulnerable, deteriorating rapidly as misleading information becomes dominant.

  • BaRe-Mem stays above autonomous reasoning in the harder regime. With Qwen3-14B on capability-challenging tasks, No consultation is 66.5 across all misleading ratios, while BaRe-Mem scores 76.1, 72.8, 71.6, 69.9 and 68.6 at 0%, 25%, 50%, 75% and 100% misleading information. With Phi-4 the corresponding numbers are 63.4 for No consultation versus 71.2, 68.5, 66.8, 65.3 and 65.0 for BaRe-Mem.

  • In the capability-supported regime, accuracy remains high and consultation stays on. For Qwen3-14B, No consultation is 72.9, BaRe-Mem is 77.7, 77.3, 76.0, 75.1 and 74.5 across the same five ratios; the consultation ratio only falls from 87% to 85%. For Phi-4, BaRe-Mem is 76.5, 75.2, 73.9, 73.0 and 71.8 against 70.2 for No consultation, with the consultation ratio moving from 85% to 83%.

  • The system adapts when to consult. In the capability-challenging regime the consultation ratio drops from 90% to 21% for Qwen3-14B and from 85% to 18% for Phi-4 as the misleading ratio goes from 0% to 100% — matching the performance patterns.

  • Weighting advisors is not enough without the option to abstain. An ablation called Advisors + memory keeps reliability-guided attention but removes the autonomous option, forcing consultation on every question. It improves robustness substantially over Question + Peers, but in the capability-challenging regime its accuracy eventually falls below the No consultation baseline for both central models, which the authors attribute to relative advisor weighting being unable to decide whether the advisor pool as a whole is worth consulting.

  • The estimated autonomous ability κ is meaningful. Estimated κ broadly tracks changes in empirical autonomous accuracy along the question stream for both central models, and empirical accuracy increases monotonically across groups of questions with similar κ values.

  • Predicted consultation gain tracks real gain. Questions are sorted by predicted gain Δ̂ = A(T) − κ and divided into 16 equal-sized groups, where real gain Δ = y^Consult − y^Direct ∈ {−1, 0, +1}. Real gain increases with predicted gain at all misleading ratios, and the curves cross zero close to Δ̂ = 0, indicating alignment between the predicted and real decision boundary.

  • Learning works from sparse feedback. Varying the fraction of questions whose verified outcomes are written to memory from 0% to 100%, both models benefit substantially from small amounts of feedback. The magnified region below 1% (approximately 170 samples) shows clear gains over the no-feedback setting, and performance largely saturates well before full feedback.

  • Reliability memory helps agent-team routing. On MuSiQue (2,417 tasks and 6,404 sub-tasks), BaRe-Mem achieves the highest task completion across both lead agents and all verification settings, outperforming random routing and routing by historical success counts. Verification improves all routing strategies, with exact dataset verification yielding the highest completion rates. When the lead agent is allowed to try more workers per sub-task, BaRe-Mem reaches higher task completion with fewer worker calls.

  • Unverified or unspecified details. The paper does not report the numeric value of γ, the prior precision λ, or the number of attention heads modified in the retrieved content; these are stated as design choices rather than measured quantities.

Methodology in Plain English

The central model is frozen — nothing in it is trained. For each question, the model produces hidden representations for K advisor answers plus its own answer. These representations are split into two parts: one that describes the question itself (shared across all candidates) and one that describes the specific candidate answer. A simple linear score combines these with a one-hot indicator of which advisor produced the answer, yielding a predicted reliability.

The memory itself is Bayesian linear regression. Each verified outcome — correct or incorrect — is converted to +1 or −1 and treated as a noisy observation of the reliability score. Because the model is linear-Gaussian, the posterior can be updated one observation at a time with a Kalman gain, so the memory never has to reprocess its whole history. The resulting reliability estimate is the probability that a candidate is correct, computed from a normal CDF of the predicted mean divided by the square root of one plus the predicted variance. With no evidence, all advisors sit at p = 1/2.

Those reliabilities are then used twice. First, they steer attention: each advisor's relative reliability is converted into a log-scale bonus added to its attention weights before softmax, so less reliable advisors are downweighted while the most reliable one is unchanged. Second, they feed a separate decision. The highest advisor reliability T and the central model's own estimated reliability κ are plugged into an interpolated consultation-ability function A(T). The two unknown parameters of that function, ρ and δ, are learned with the same online Bayesian regression from verified questions, and the model consults only when A(T) is at least κ.

To stress-test robustness, the authors replace a controlled fraction of advisor responses (0% to 100%, in 25% steps) with answers that are fluent, on-topic and well-formed but verified to be incorrect.

Why This Matters

Impact on research. The paper reframes multi-agent consultation from an aggregation problem into an estimation problem: rather than pooling responses at inference time, an agent should maintain calibrated beliefs about who — including itself — is likely to be right. The result that relative advisor weighting fails when the whole advisor pool is unreliable is a useful negative finding for the multi-agent debate and voting literature, and the calibration analysis (predicted gain versus real gain crossing zero near Δ̂ = 0) offers a template for evaluating decision-theoretic agent designs rather than just accuracy.

Real-world applications:

  • Enterprise AI orchestration, where a router must choose among specialized models, tools or services and needs to know when the pooled evidence is collectively untrustworthy.
  • High-stakes domains such as medicine, law or finance, where an agent must recognize when to fall back on its own reasoning rather than defer to a fluent but wrong advisor.
  • Agent-team task pipelines such as the MuSiQue setup, where a lead agent decomposes tasks and must assign sub-tasks to workers and know when to retry with a different worker.
  • Cost-constrained deployments, where verified feedback is expensive — the sparse-feedback results indicate gains are available from roughly 170 verified samples and below 1% feedback coverage.

Industry relevance. Because the method adds no trainable parameters and no additional training, and because the posterior updates are rank-one, it is a lightweight add-on to existing frozen model deployments. The demonstrated benefit in worker routing suggests direct applicability to agent frameworks that already maintain some notion of worker reputation.

Future Directions

  • Moving beyond relative weighting within a fixed advisor pool. The ablation shows that weighting alone fails when the pool is collectively unreliable; extending the framework to dynamically recruit, drop or replace advisors based on the same reliability memory is a natural next step.
  • Reducing dependence on verification. The memory is updated only from verified correctness outcomes. The paper does not report how the method behaves when verification is noisy or self-reported rather than exact, despite showing that verification quality changes absolute routing performance.
  • Scaling and calibration of the reliability representation. The reliability score is linear in the belief representation; whether richer or non-linear mappings of ψ_q and ψ_c improve calibration — and how the hyperparameters γ and λ should be chosen — is left open in the retrieved content.
  • Broader deployment scenarios. The authors frame worker routing as an extension of "response level consultation," implying further settings such as hierarchical teams, long-horizon task streams and tool selection remain to be tested.

Target Audience

Researchers and practitioners working on multi-agent LLM systems, agentic orchestration and LLM routing will get the most from this paper, particularly those interested in trust, calibration and when-to-delegate decisions. It is also relevant to applied machine learning engineers building production systems where models consult external services or peers and need a robust fallback to autonomous reasoning. Readers without a background in Bayesian inference or attention mechanics will need to work through the derivations in the appendix, which cover the Bayesian regression posterior, the Kalman gain, and the reliability and consultation-parameter derivations.

Authors’ abstract

In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advisor reliability based on the central model's internal belief representations and updates these estimates from historical interactions. These estimates modulate the influence of advisor responses and guide the choice between consultation and autonomous reasoning. Across nine benchmarks and six central models, BaRe-Mem is more robust to misleading advisor information than debate and majority voting. On the more challenging tasks, it remains above autonomous reasoning across all tested misleading levels. Moreover, we extend the BaRe-Mem mechanism to worker allocation in agent teams. On the MuSiQue benchmark, BaRe-Mem improves task completion over routing by historical success counts and identifies capable workers earlier.

Read the original paper