Skip to content
AI.info

Research

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories

Overview Research area: Natural Language Processing applied to biomedicine — specifically, benchmarking and training large language models to simulate multidisciplinary tumor board meetings for cancer

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
arXiv
2609.32810
Published
2026-09-26
Authors
Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian, Chi-Yu Chen, MingYu Lu, Huan-Yu Hsu, Yu Gu, Yue Guo, Sheng Wang, Wei Qiu, Hanwen Xu

AI summary

Overview

  • Research area: Natural Language Processing applied to biomedicine — specifically, benchmarking and training large language models to simulate multidisciplinary tumor board meetings for cancer decision-making.
  • Technical level: Intermediate. The tasks and results are described in clinical and evaluation terms that a general ML reader can follow, though familiarity with LLM benchmarks, supervised finetuning, and reinforcement learning helps.
  • Scope: The paper introduces OpenTumorBoard, a benchmark of 611 real patient cases and 19,157 discussion turns transcribed from 12,534 minutes of publicly recorded tumor board videos, and uses it both to evaluate 14 LLMs and to train smaller models on real discussion trajectories.

What This Paper Is About

Tumor boards are meetings where specialists from different disciplines jointly review a cancer patient's multimodal evidence and agree on what to do next, and they are often a last hope when standard options run out. Existing biomedical benchmarks do not test LLMs on the kind of cases, questions, and multi-specialist back-and-forth discussions that actually occur in these meetings. The goal of this paper is to build a benchmark grounded in real recorded tumor board discussions — complete with presentation slides, specialist-turn questions, and final board consensus — and to use it to measure and improve how well LLMs can reason, respond, and simulate the meeting end-to-end.

Key Contributions

  1. A large-scale benchmark grounded in real tumor board practice. OpenTumorBoard contains 611 patient cases curated from 219 recorded tumor board meetings spanning diverse cancer types and clinical modalities, compared with MTBBench, which includes 66 patient cases.
  2. Complete naturally occurring discussion trajectories across ten specialist roles. The paper states this is, to the best of its knowledge, the first benchmark providing complete tumor board discussion trajectories paired with observed clinical consensus, enabling evaluation of both individual specialist turns and full meetings.
  3. A development testbed for post-training tumor board LLMs. Using discussion trajectories and consensus as supervision, the authors build a staged training framework combining supervised finetuning and reinforcement learning, evaluated on Board Simulation.
  4. M.D. expert-validated benchmark quality. Three M.D. experts with oncology expertise review a subset of the benchmark and validate the extracted specialist responses and board conclusions as high-quality, clinically grounded references.

Main Findings

  • Benchmark scale and structure: 611 patient cases from 219 tumor board meetings, totaling 12,534 minutes of recording and 19,157 discussion turns across ten specialist roles. The test set contains 66 recordings, 184 cases, and 4,844 questions. Recordings were split 60/10/30 into training, validation, and test sets. Across all cases, discussion trajectories have a median duration of 17 minutes, a median of 7 slides, and 28 specialist turns per case.
  • Breadth of content: The dataset spans 17 cancer sites, 10 specialist roles, 9 specialist turn types, and 11 input modalities. More than one-third of discussions focus on metastatic disease. Surgeons, medical oncologists, and radiation oncologists are the most frequently involved specialists.
  • Specialist Turn performance is weak: Evaluated across 9 open-weight models, the best result is a clinical equivalence score of only 3.43 out of 5, achieved by DeepSeek-V4-Pro in reasoning mode.
  • High error rates in specialist responses: Critical-error rates range from 5.6% to 43.5%, and unsupported-claim rates span 10.0% to 74.5% across the models evaluated on Specialist Turn.
  • Medical-specific models do not outperform general-purpose models on Specialist Turn: DeepSeek-V4-Pro and DeepSeek-V4-Flash score highest, followed by Gemma 4 31B and Ministral 3 14B, while three of the four medical models rank at the bottom. Meditron3-70B is the only medical model with a clinical equivalence score above 3, merely matching the general-purpose Ministral 3 14B. HuatuoGPT-3-32B and MedReason-8B have the highest unsupported-claim rates at 74.5% and 51.9%.
  • Difficulty varies by question type: Across 9 models, clarification questions achieve the highest average clinical equivalence at 3.30. Performance is much lower for clinical trial suggestion and evidence discussion; for the best-performing model, unsupported-claim rates rise from 11.8% overall to 31.5% and 23.5% on those two types respectively. Questions requiring supporting evidence are consistently the most difficult.
  • Board Simulation remains far from solved: Across 14 models (5 proprietary frontier models and 9 open-weight models), no model reliably recovers the ground-truth meeting conclusion. Gemini 3.7 Flash is best at only 2.78 out of 5 for conclusion alignment. Leading proprietary reasoning models (Gemini 3.7 Flash, Grok 4.6, GPT-5.6 Sol, and Claude Opus 5) tend to outperform open-weight models.
  • Longer trajectories correlate with better alignment: Across the 14 models, longer discussion trajectories are strongly associated with higher conclusion alignment (Spearman rho = 0.63).
  • Training on real trajectories improves a small model: Supervised finetuning of Qwen2.5-VL-3B with only 366 training cases raises conclusion alignment from 1.58 to 1.67 (approximately 5.7%). Reinforcement learning with Dr.GRPO further raises it to 1.86, closing half the gap to MedGemma 27B, and outperforms the base model across therapy recommendations, surgical plans, next actions, and clinical trial matching. The conclusion section reports the finetuning and RL pipeline yields an approximately 18% relative improvement in conclusion alignment for the small 3B model.
  • Expert audit supports annotation quality: Based on 447 question-answer pairs, 15 case summaries, and 15 consensus conclusions, 92.6% of extracted answers are rated "High" for response correctness and 94.4% for discussion support. Case coverage, case factuality, and consensus fidelity are rated "High" for 100% of cases. Question relevance is "High" or "Medium" for 94.9% of turns. Annotator agreement before adjudication ranged from 60.0% (question relevance) to 93.3% (case factuality).
  • Pipeline audit: Three annotators reviewing fifteen randomly selected recordings found 100.0% case count accuracy and 98.1% case-boundary accuracy (106 boundaries; 104 of 106 identified across 53 cases), slide extraction precision of 93.0% over 313 frames (99.7% with duplicates excluded) and recall of 93.6% over 311 presented slides (291 recovered), and role-inference accuracy of 96.3% over 136 speakers (131 correct). Inter-rater agreement was 85.3% for specialist role inference (Fleiss kappa 0.820) and 93.4% for slide extraction (kappa 0.526).
  • Case study of discussion value: In one example, an early proposal to re-irradiate the prostate bed is retracted after the radiation oncologist identifies the prior 71.8 Gy dose, and the board converges on surveillance keyed to PSA doubling time after 32 turns.

Methodology in Plain English

The authors treat YouTube as a source of real tumor board meetings, retrieving recordings with "tumor board" in the title and a duration above twelve minutes, then using LLMs to read transcripts and discard generic lectures. They build a nine-step automated pipeline with two decoupled tracks. The audio track establishes speaker identity and timing using WhisperX (Whisper large-v2) for transcription and the pyannote pipeline for diarization, configured for one to ten speakers per recording; GPT-5.4 then maps each voice to one of ten named specialist roles or "other," segments the transcript into cases, and extracts evidence, questions, and conclusions. The visual track samples frames every five seconds (0.2 frames per second), prefilters them with CLIP ViT-B/32 using zero-shot similarity to slide versus speaker/audience prompts, localizes the projected slide with Grounding DINO, crops it with SAM2.1, removes near-duplicates at CLIP cosine similarity above 0.85, and verifies and captions retained slides with a GPT-5.4 vision pass. Multimodal temporal alignment then links slides to the question being posed at that moment by timestamp, and a de-leaking step removes information about the decision point.

From these artifacts the authors define two tasks. In Specialist Turn, a model is given the case summary and slides, assigned a specialist role, and must answer a question that was actually posed in the meeting, with the real specialist's answer as ground truth. In Board Simulation, the model must generate the entire multi-specialist discussion and reach a consensus conclusion, scored against the real board's conclusion across therapy recommendations, surgical plans, next actions, and clinical trial matching. Evaluations use Qwen3.8-27B as an LLM judge with a clinical-equivalence rubric and a conclusion-alignment rubric, each on a 1–5 scale, plus critical error rate, unsupported claim rate, ROUGE-L, and BERTScore. For training, the recorded discussions and conclusions serve as supervision for supervised finetuning, followed by reinforcement learning with Dr.GRPO using a reward comparing predicted conclusions to board decisions.

Why This Matters

Impact on research. The paper argues that existing biomedical benchmarks fall short along four dimensions at once — real-world tumor board cases, a multidisciplinary team, ground-truth discussion, and multimodal data — and shows that general-purpose reasoning gains do not translate into clinically meaningful specialist reasoning. It also demonstrates that real discussion trajectories can serve as supervision, giving the field a training ground rather than only a test set. The authors release an automated curation pipeline so the benchmark can be scaled or extended to newer recordings.

Real-world applications.

  • Timely decision support when specialist availability delays tumor board discussions, which the paper notes can affect even patients needing urgent care.
  • Virtual specialist response generation for naturally occurring clinical questions such as findings interpretation, treatment recommendation, and clinical trial suggestion.
  • End-to-end simulation of multidisciplinary meetings to produce therapy recommendations, surgical plans, next actions, and clinical trial matching.
  • Local extension of the curation framework to institutional datasets, subject to privacy and governance requirements.

Industry relevance. The results matter for teams building clinical LLM products, because medical-specific models did not beat general-purpose models here, and because weaker models frequently produced clinically consequential errors or unsupported claims (critical-error rates up to 43.5%, unsupported-claim rates up to 74.5%). The evaluation and training code, prompts, judge protocols, and gated dataset access are released, and a leaderboard is publicly available, which supports standardized comparison across commercial and open models.

Future Directions

  • Establishing agreement between rubric-based LLM judging and clinician assessment, which the authors state remains important future work; they plan to incorporate clinician evaluation to validate and calibrate scores.
  • Addressing representativeness, since the benchmark is derived exclusively from publicly available YouTube recordings; the released curation framework is intended to enable local extension to institutional datasets under applicable privacy and governance requirements.
  • Scaling and continuously growing the benchmark from newly posted recordings, toward what the authors describe as a self-evolving benchmark that grows as model capabilities advance.
  • Pushing post-training further, since supervised finetuning plus reinforcement learning closed only half the gap between a 3B model and MedGemma 27B on conclusion alignment.

Target Audience

Researchers and engineers working on clinical NLP, biomedical LLM evaluation, or multi-agent and long-context reasoning; clinicians and informatics teams interested in how well automated systems approximate tumor board decisions; and model developers who want a realistic, expert-audited benchmark plus a released pipeline for post-training and for extending the dataset to their own recordings.

Authors’ abstract

Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube. The benchmark evaluates two settings: SPECIALIST TURN, in which an LLM responds to a clinically significant question posed during a real discussion, and BOARD SIMULATION, in which it generates an entire back-and-forth discussion and reaches a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Evaluation of 14 general-purpose frontier and medical LLMs reveals substantial limitations: the best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions. Supervised finetuning and reinforcement learning improve performance on a held-out test set, suggesting that real-world discussion trajectories can support model adaptation. Three M.D. experts review a subset of the benchmark, finding high information coverage and factuality of patient cases and strong fidelity of extracted consensus conclusions. We will release OpenTumorBoard and its automated curation pipeline to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.

Read the original paper