Skip to content
AI.info

Research

ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis

Overview Research area: Multimodal medical AI, sequential/active evidence acquisition, and cost-aware decision-making around frozen vision-language models (VLMs). Technical level: Advanced. The paper

arXiv
2610.11140
Published
2026-10-08
Authors
Weiwei Ma, Xiaobing Yu, Peijie Qiu, Jin Yang, Zhaoqi An, Xuanzhao Dong, Xiaoqi Zhao, Xiaofeng Liu

AI summary

Overview

  • Research area: Multimodal medical AI, sequential/active evidence acquisition, and cost-aware decision-making around frozen vision-language models (VLMs).
  • Technical level: Advanced. The paper combines a sequential decision formulation, a reward-weighted offline imitation objective, and clinical benchmark evaluation, though the core idea is understandable without deep math.
  • Scope (one sentence): The paper proposes ActiveMedAgent, a lightweight learned controller that decides which clinical evidence channels a frozen VLM should be shown, and when to stop, balancing diagnostic quality against acquisition cost on NEJM Image Challenge, MIDAS, and OLIVES.

What This Paper Is About

Most medical multimodal AI is evaluated as a static prediction problem: the model sees all available evidence at once and produces a diagnosis. Real clinical diagnosis instead works as staged escalation, where cheap evidence is used first and expensive tests are ordered only when they are expected to resolve uncertainty. ActiveMedAgent learns this escalation behavior around a frozen, API-accessed VLM by tracking belief distributions over candidate diagnoses and scoring each potential acquisition by its diagnostic benefit minus its cost.

Key Contributions

  1. A structured, hierarchy-constrained formulation of multimodal diagnosis. Instead of free-form query generation, the agent acts over a finite, clinically grounded action space of named evidence channels at tiered costs, plus a COMMIT action.
  2. Trajectory-level supervision over belief state transitions. Rather than learning only from final predictions, the method scores each step by reciprocal-rank improvement minus normalized acquisition cost, labeling steps as beneficial, redundant, or harmful, and trains a policy from these intermediate signals.
  3. Evidence that offline policy learning beats prompt-based and self-reflective acquisition. For example, PolicyNet raises NEJM accuracy from 18.6% to 31.2% on GPT-4o-mini, where naive acquisition fails.
  4. Identification and quantification of an information overload effect. In 175 cases, selective acquisition produces a correct diagnosis while the full-modality baseline is wrong, motivating evaluation based on accuracy and cost rather than accuracy alone.

Main Findings

  • Learned acquisition beats unguided acquisition. On GPT-4o-mini, Zero-shot fell below Passive on all three datasets (NEJM 18.6% vs. Passive 18.9%; MIDAS 20.4% vs. 24.6%; OLIVES 23.5% vs. 29.4%), and its negative acquisition efficiency on MIDAS indicates that unguided requests can actively worsen the differential. PolicyNet reached 31.2% (NEJM), 37.3% (MIDAS), and 52.9% (OLIVES) on the same backbone.

  • PolicyNet matches or exceeds Full Modality in all six settings. Examples: NEJM/GPT-4o 48.3% vs. 46.7%; NEJM/GPT-4o-mini 31.2% vs. 30.9%; MIDAS/GPT-4o 47.5% vs. 47.1%; MIDAS/GPT-4o-mini 37.3% vs. 36.1%; OLIVES/GPT-4o 54.2% vs. 41.2%; OLIVES/GPT-4o-mini 52.9% vs. 47.1%.

  • Ranked differentials improve, not just top-1. PolicyNet's largest MRR gain over Full Modality is +21.0 points on NEJM/GPT-4o, with GPT-4o-mini gains of +8.2 on MIDAS and +17.7 on OLIVES. Because the VLM is frozen, these gains come from better evidence trajectories.

  • Cost savings per correct diagnosis. The paper reports savings of up to $1,768 per correct diagnosis (NEJM/GPT-4o-mini), where PolicyNet costs $706 versus $1,125 for Full Modality, with $/Cor of 1,875 versus 3,643.

  • Adaptive stopping is real, but partial. Early-stopping rates for PolicyNet range from 17.3% (MIDAS/GPT-4o-mini) to 27.3% (OLIVES/GPT-4o-mini); channels acquired range from 2.5 (NEJM/GPT-4o-mini) to 3.6 (OLIVES/GPT-4o), compared with 3.0 to 4.0 for Fixed-order and Full Modality.

  • Information overload. Across all six settings, 175 cases exist where the selective agent is correct while Full Modality fails. The illustrated OLIVES/GPT-4o-mini case: the agent commits after clinical_measurements ($20) and biomarker_hints ($100) with 2 channels, MRR=1.0, correct, $120, whereas Full Modality uses 4 channels, MRR=25%, wrong, $570.

  • PolicyNet dominates under distractor stress tests. On NEJM with +2 distractors (GPT-4o), PolicyNet reaches 0.617 top-1 at $787 mean cost, versus Reflective 0.577/$1,034, EntropyGreedy 0.535/$1,084, SelfAsk 0.506/$1,072, ReAct 0.504/$1,117, AllAtOnce 0.493/$1,125, and Passive 0.248/$0. The paper describes this as Pareto-dominance: higher accuracy at 30% lower acquisition cost. Oracle scores 0.461 at $1,125.

  • Comparison with prior sequential-diagnosis systems on a matched NEJM protocol. ActiveMedAgent reaches 0.625 top-1 with 2.0 average acquisitions, versus MAI-DxO 0.610/3.5, SDBench-Tool 0.590/4.0, VoI (Bayesian, hand-tuned) 0.585/3.0, Reflexion + tool calls 0.560/3.2, and Plain CoT prompting 0.495/3.0. The paper states this is a protocol-matched reproduction, not a leaderboard number.

  • Ablations show what carries the policy. On NEJM+2dist (GPT-4o), full PolicyNet scores 0.625 (AE 1.00); removing the case-text embedding drops to 0.585 (0.85); removing the channel-name embedding head to 0.600 (0.90); removing validation early-stopping to 0.610 (0.95); removing distractor-aware training to 0.575 (0.78); removing the cost penalty to 0.615 (0.98); swapping numerical EIG for keyword-IG to 0.605 (0.91); a vanilla MLP with all removals falls to 0.520 (0.60).

  • Cost penalty is a controllable knob. Increasing lambda smoothly reduces mean cost; lambda = 5 cuts cost by 55% with only a 1.4 point accuracy drop.

  • Training is cheap; inference is not. Policy training finishes in under two minutes on CPU, while the dominant cost is VLM trajectory collection.

Methodology in Plain English

The researchers treat a clinical case as a set of "channels" of evidence. Some channels are free and available at presentation (for example, basic measurements or clinical text), and others can be requested at a stated cost, organized into a tiered hierarchy. A frozen VLM, accessed only through an API, is asked at each step to return a ranked differential diagnosis with confidence scores in a strict JSON format. Scores that do not sum to one are repaired by renormalization, and omitted candidates get uniform probability.

To build training data, the system runs the VLM through training cases in a loop: reveal a channel, get an updated differential, and score the step. The score for each step is the change in reciprocal rank of the correct diagnosis (how much the correct answer moved up the list) minus a normalized cost penalty. Steps are then tagged beneficial, redundant, or harmful.

A small neural network controller, PolicyNet, is trained offline on these scored transitions. Its input is a compact state: which channels have already been acquired, a dataset identifier, the top-1 confidence, the gap between the top two confidences, belief entropy, the step number, and cumulative cost, plus a frozen embedding of the presentation text. Each possible channel is scored using a learned channel-name embedding, and COMMIT has its own output. The network is a 3-layer MLP (D_s to 64 to 32) with ReLU activations and dropout of 0.1, trained by reward-weighted imitation learning: transitions with high utility get exponentially more weight, and harmful transitions are exponentially suppressed. The action space is small (at most 5 actions), which is why the authors chose offline imitation over online reinforcement learning such as PPO or DQN, since online RL would require repeated expensive VLM calls.

At inference, the agent loops: the policy either chooses a channel or commits, and stopping is also value-based. For each remaining channel, the system estimates expected net value as estimated information gain minus cost. Numerical information gain uses an extra VLM call to replay the candidate channel and measure the entropy reduction plus a confidence-gap term; a cheaper keyword-based proxy exists but underperforms and is not the default. The agent commits when PolicyNet selects COMMIT or when all remaining channels have non-positive expected value.

The controller is decoupled from the backbone (the VLM is never modified), auditable (a small MLP over interpretable features), and trajectory-supervised (able to distinguish helpful from redundant or harmful acquisitions). Training settings: Adam at 1e-3, batch size 64, 200 epochs, 20% validation early stopping, reward temperature alpha = 1.0, cost penalty lambda = 0.5.

Datasets and protocol: NEJM Image Challenge (947 cases; examination, investigations, image), MIDAS (635 cases; close-view photographs and dermoscopy), and OLIVES (1,268 cases; measurements, biomarkers, OCT evidence), each split 60/40 into train/test. Backbones are GPT-4o and GPT-4o-mini. Metrics are top-1 accuracy, MRR, channels acquired, early stopping, acquisition efficiency (AE), mean cost, and cost per correct diagnosis, with bootstrap 95% confidence intervals. The paper also states that policy training finishes in under two minutes on CPU, that malformed VLM responses are retried up to 3 times with persistent failures excluded (under 2% of cases), and that reward-weighted ICL used 3 retrieved demonstrations. The authors explicitly do not claim cross-backbone transfer, since GPT-4o and GPT-4o-mini showed different uncertainty and acquisition profiles.

Why This Matters

Impact on research. The paper reframes multimodal medical diagnosis as an evidence-control problem rather than a static prediction problem. It shows that a full-modality baseline is not a reliable upper bound under this prompting regime, which challenges the common assumption that more evidence is always better. It also separates reasoning (the frozen VLM) from decision-making (the learned policy), giving a template for studying acquisition policies without retraining or accessing model internals.

Real-world applications.

  • Clinical decision support that recommends a next test only when expected to change the differential, with explicit cost accounting.
  • Triage and resource planning in settings where expensive imaging or biomarker tests are constrained.
  • Auditing and reviewing AI diagnostic behavior, since each request, stop, and error is tied to a named channel and a small, interpretable controller.
  • Benchmark design for medical VLMs, since the results argue for evaluating accuracy alongside acquisition cost and channel count.

Industry relevance. Cost-aware escalation maps directly to payer and health-system incentives, where unnecessary testing has both financial and patient-burden implications. Because the controller learns offline and requires no gradients through the backbone, it is compatible with API-served commercial VLMs. The paper reports savings of up to $1,768 per correct diagnosis on NEJM/GPT-4o-mini and a Pareto improvement of higher accuracy at 30% lower acquisition cost under distractor stress tests, which is the kind of operating-point argument that matters for deployment decisions.

Future Directions

  1. Prospective and local-cost validation. The authors state that experiments are retrospective and use protocol-level costs rather than institution-specific costs, turnaround times, or patient burden estimates.
  2. Calibrated beliefs. VLM-reported probabilities are imperfectly calibrated; the framework leans on rank changes and compact uncertainty features. Extracting better-calibrated beliefs from natural-language outputs is an open problem.
  3. Safer escalation behavior. The paper flags the risk that a policy could learn to under-request costly evidence for rare or ambiguous cases, appearing efficient while delaying necessary escalation, and that blind spots in the backbone can be inherited by the policy.
  4. Human-in-the-loop and subgroup review. The authors call for studying how clinicians use or override learned escalation recommendations, along with human override, audit logs, and subgroup monitoring, which the appendix describes as safety checks.

Target Audience

Researchers and practitioners in medical machine learning, multimodal vision-language modeling, and clinical decision support who are interested in sequential or active evidence acquisition under cost constraints. It is also relevant to health-system informatics teams evaluating test-ordering efficiency, and to benchmark designers who want to evaluate medical VLMs on the evidence path rather than only final accuracy. Readers without a machine learning background will find the motivation and case study accessible, but the full method and evaluation require comfort with probabilistic belief states, reciprocal rank, and offline imitation learning.

Notes on scope: The provided paper content is truncated, so appendix-level details (in the authors' described order: external validity, design robustness, and mechanism plus safety) are only partially visible. The paper reports no prospective clinical trial data and does not claim cross-backbone transfer.

Authors’ abstract

Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost. A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit. Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines. Notably, we identify an information overload effect. In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.

Read the original paper