Skip to content
AI.info

Research

A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care

Overview Research area: Computer Vision / Vision-Language Models, applied to medical and tactical emergency care (multimodal dataset construction and benchmarking). Technical level: Intermediate. The

A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care
arXiv
2610.07339
Published
2026-10-05
Authors
Junseob Kim, Jade Chng, Ayman Ali, Victor Moas, Yichun Lee, Po-Chun Chin, Sunil Hwang, Rishikesan Kamaleswaran

AI summary

Overview

Research area: Computer Vision / Vision-Language Models, applied to medical and tactical emergency care (multimodal dataset construction and benchmarking).

Technical level: Intermediate. The paper assumes familiarity with visual question answering (VQA), vision-language models (VLMs), retrieval pipelines, and natural language inference (NLI) verification, but its construction logic is described step by step.

Scope: The paper introduces TC3-VQA, a doctrine-grounded visual question answering dataset of 1,860 questions over 581 items built from public Tactical Combat Casualty Care (TC3) videos and authoritative TC3 documents, together with construction methodology, quality validation, and baseline evaluations of five open VLMs.

What This Paper Is About

Supporting responders with AI in Tactical Combat Casualty Care requires training data that ties what is visible in a scene—equipment, interventions, anatomy—to clinical guidance that can be traced back to a source document. Existing medical VQA datasets focus on radiology, pathology, and related imagery, and public video material is prone to visual hallucination, temporal misalignment, and unsupported citations. The paper's goal is to build a dataset in which every question is grounded in verified visual evidence and every doctrinal answer reproduces a verbatim passage from an authoritative source, and then to measure how well current open VLMs perform on it.

Key Contributions

  1. TC3-VQA, a doctrine-grounded dataset: 581 items (431 answerable, 150 refusal) spanning 11 concepts and 1,860 questions, with answers to doctrine, reasoning, and procedural "how" questions anchored to verbatim corpus spans recorded by chunk and character offsets, drawn from 190 videos.
  2. A verification-first construction pipeline: visual annotation, passage retrieval, deterministic provenance checking, NLI entailment filtering, cross-family model consensus, and human ratings, designed to independently verify both what a scene supports and what a clinical source substantiates.
  3. Rich visual and provenance annotations: a 17-class equipment ontology with 1,565 boxes, answer-region boxes, eight anatomical body-region labels, temporal segments, source metadata, license status, safety-critical flags, and visibility tiers released per item.
  4. Baseline evaluation and ablation evidence: benchmarks of five open VLMs (7–38 billion parameters) on recognition, doctrine multiple choice and free-form answer, refusal, and a Risk-Weighted Hallucination Rate (RWHR), plus frame-substitution, temporal-truncation, and direct-generation comparisons.

Main Findings

  • Dataset scale and composition: 1,860 questions across 581 items (431 answerable and 150 refusal) from 190 videos and 11 concepts. The answerable set comprises 431 recognition, 427 doctrine, 425 reasoning, and 427 "how" questions, with all four types available for 418 items. Of the 581 items, 309 use a single frame and 272 use two to four frames; each video contributes a median of three items and a maximum of eight.
  • Short visual context: Median segment duration is 10 s (interquartile range 4–30 s). Questions are one or two sentences, with median lengths of 9 words for recognition up to 27 for refusal; reference answers range from a 3-word concept name to a 30-word procedural passage.
  • Answer provenance: ATP 4-02.11 supplies approximately half of the cited answers and 73% of procedural answers, with the CoTCCC and Joint Trauma System guidelines supplying most of the remainder. Refusal questions most often request blood-loss volume, elapsed time, or vital signs, with no category exceeding 26% of refusal items.
  • Concept-label reliability varies with visibility: Of 431 released answerable items, 245 were classified clearly visible and 186 partly visible. Model consensus was unanimous for 147 items, five of six for 117, four for 91, and three for 76 (mean 4.8 of 6). Individual-model agreement with the released label ranged from 57% for MedGemma-27B to 94% for Qwen2.5-VL-72B. Both physicians accepted 82% of clearly visible labels versus 44% of partly visible labels.
  • Automated audit results: The majority verdict judged 98.7% of refusal questions unanswerable versus 10.8% of answerable controls on the same frames; 94.6% of "how" questions were anchored to their paired frames versus 11.0% on substituted frames. Doctrine questions were less scene-specific (58.8%), reflecting that 37% asked about indications or guidance applicable across scenes of the same intervention. The reasoning audit accepted 91.5% of released answers versus 33.0% of within-concept controls.
  • Reasoning audit sensitivity gap: The reasoning audit accepted 48 of the 49 answers accepted by both physicians, but also 35 of the 38 flagged by at least one physician, indicating lower sensitivity to deficiencies identified by expert review.
  • Equipment annotations: 1,565 boxes across 743 of 903 frames, covering 16 of 17 ontology classes. Among the 431 answerable items, 383 contain equipment boxes, 393 contain answer-region boxes, and 420 contain at least one of the two. A YOLOv8s detector trained on 595 frames and evaluated on 148 held-out frames reached mAP50 = 0.49 and mAP50-95 = 0.35, with class-level mAP50 of approximately 0.6 for hands, decompression needles, hypothermia wraps, and surgical-airway devices, versus 0.28 for hemostatic dressings.
  • Human ratings: Two physicians (one TC3-trained) and two medical students evaluated 120 candidate items, 88 of which were retained. Both physicians rated 69% of recognition labels correct (AC1 0.80); medical students accepted more labels (93–94% correct), including 23 of the 30 questioned by at least one physician. Across the three quoted-answer types, outright incorrect ratings ranged from 0 to 11%. Physician severity agreement was lower (AC1 0.27), though mean physician severity increased from 1.0 to 2.05 and 2.4 across concepts assigned RWHR weights of 1, 2, and 3.
  • Baseline performance: Recognition accuracy ranged from 0.67 to 0.85. Qwen2-VL-7B achieved the highest recognition accuracy (0.85) but only 26% refusal accuracy and the highest RWHR (0.92), whereas InternVL3-38B achieved 0.74, 93%, and 0.42, respectively, illustrating a recognition–abstention trade-off. Doctrine-MCQ accuracy ranged from 0.87 to 0.99 with easy distractors and 0.70 to 0.91 with hard distractors. Free-form responses entailed 0.23–0.29 of reference sentences, while 52–59% of generated sentences were unsupported.
  • Recognition depends on the image: Substituting frames from another intervention reduced accuracy to 0.04–0.14 for four models, below the 0.20 uniform-choice reference, while Phi-3.5-Vision retained 0.33. Restricting the 204 multi-frame items to the first frame reduced accuracy by 7.8–12.7 percentage points for four models, with significant exact McNemar tests, while Qwen2-VL-7B decreased by 2.5 points (p = 0.42).
  • Pipeline versus direct generation: On 288 audit-confirmed candidates, direct generation by open VLMs produced questions naming another listed concept in 16–31% of cases, versus 0.4% for pipeline-generated questions. When models selected the concept themselves, 11–42% of questions remained inconsistent with the audited concept, whereas Qwen2.5-VL-72B produced no such errors when the audited concept was provided. Pipeline pairs satisfied concept-consistency and answer-support criteria for 95–99% of questions versus 72–78% for direct generation by Qwen2.5-VL-72B; among concept-consistent questions, at least 92% of answers in both groups satisfied both criteria.

Methodology in Plain English

The researchers first assembled source material. Using search terms developed with a TC3-trained physician, they collected 415 public YouTube videos of TC3 instruction and field casualty care, then split them into shots with PySceneDetect at its default threshold (27), dividing shots longer than 60 s into windows of at most 30 s and sampling frames roughly every 4 s. Filtering for poor exposure, blank content, blur, and near-duplicates left 32,987 frames in 17,055 windows. Qwen2.5-VL-72B-Instruct then judged each window in temporal order, assigning an action phase and a usefulness score from 0 to 5; windows with an active procedure or post-procedure result scoring at least 3 were kept with up to four qualifying frames, yielding 8,012 windows.

In parallel, they built a doctrine corpus of 9,320 chunks from 11 sources, including the January 2024 CoTCCC guidelines, the December 2021 guidelines for medical personnel, ATP 4-02.11, the Joint Trauma System Clinical Practice Guidelines, and Joint Publication 4-02. Documents were extracted with PyMuPDF and split at headings into chunks of approximately 500 tokens (maximum 800), with each chunk retaining source, section path, position, and page range; a prose filter removed 1,428 of the 5,142 chunks derived from guideline and doctrine documents.

Each retained window received one of 12 observable intervention or documentation labels organized by the MARCH sequence plus a documentation category. High- and medium-confidence assignments (5,051 of 8,012 windows) formed the answerable pool; limiting each video to six candidates produced 1,515 items for review. Labels were verified in stages: a stricter perception prompt, a blinded visual review with Claude Opus 4.

Authors’ abstract

Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constructed from public instructional and field TC3 videos and authoritative TC3 documents. It contains 581 items spanning 11 concepts, with 1,860 questions covering intervention recognition, doctrine, clinical reasoning, procedural guidance, and refusal when visual information is insufficient. Doctrine-based answers preserve verbatim source passages and character offsets. Construction combines visual annotation, passage retrieval, entailment checks, and verification across model families. Equipment boxes, anatomical labels, temporal segments, and source metadata accompany the question-answer pairs. Automated audits and ratings by two physicians and two medical students characterize annotation quality, with human ratings available for 88 retained items. The dataset provides a resource for adapting vision-language models to TC3, studying the connection between visual evidence and clinical knowledge, and evaluating recognition, doctrine recall, and abstention.

Read the original paper