Skip to content
AI.info

Research

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Overview Research area: Natural Language Processing / LLM evaluation, specifically health reasoning over longitudinal wearable time series. Technical level: Intermediate. The paper is a benchmark-and-

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
arXiv
2609.05405
Published
2026-09-04
Authors
Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda

AI summary

Overview

Research area: Natural Language Processing / LLM evaluation, specifically health reasoning over longitudinal wearable time series.

Technical level: Intermediate. The paper is a benchmark-and-evaluation study, so no new model architecture or training method is proposed, but readers need some familiarity with LLM evaluation protocols (chain-of-thought prompting, accuracy with confidence intervals, positional bias analysis) and with physiological signals such as heart-rate variability, VO2max, and sleep staging.

One-sentence scope: The paper introduces WearableQA, a 4,084-question multiple-choice benchmark built from the real longitudinal wearable, blood-biomarker, and demographic records of 200 users, and uses it to evaluate 14 proprietary and open-source LLMs on data versus health reasoning and single- versus cross-signal reasoning.

What This Paper Is About

Existing health and time-series benchmarks for LLMs mostly use static text, examination-style questions, or synthetic and simulated signals, so it is unclear whether a model can actually reason over a real person's daily wearable history. WearableQA was built to close that gap by grounding every question in genuine measurements from 200 real users, each with up to 500 days of daily records, while preserving the device noise and inter-individual variability that real wearables contain. The goal is a benchmark that both discriminates between models and diagnoses where their reasoning breaks down.

Key Contributions

  1. A real-user benchmark. WearableQA contains 4,084 10-option multiple-choice questions derived from the longitudinal wearable records of 200 real users, together with blood biomarkers, demographic attributes, and cohort-relative context.

  2. A dual-grounding construction framework. Every question is grounded either in peer-reviewed literature findings or in statistically validated patterns mined from a large real-world cohort, with ground truth re-derived deterministically from the underlying measurements rather than asserted. The paper reports 11 anchor papers from venues including Nature Medicine and PNAS, and population-grounded patterns must pass effect-size, robustness, and authenticity gates.

  3. A diagnostic taxonomy of 16 question types. Questions are organized along two complementary axes — data reasoning versus health reasoning, and single-signal versus cross-signal reasoning — plus a separate grounding axis recording whether ground truth came from literature or from population patterns.

  4. A comprehensive evaluation. The authors evaluate 14 proprietary and open-source LLMs, showing the benchmark is discriminative (accuracy from 19.6% to 72.9% against a 10% chance baseline) and diagnostic of distinct reasoning limitations.

Main Findings

  • Wide spread across models, and far from solved. Accuracy ranges from 19.6% (Llama-3.2-3B, near the 10% random baseline) to 72.9% (Gemini-3.1-Pro). Most models fall below 60%. Gemini-3.1-Pro leads the next-best model, Claude-Opus-4.6 (60.2%), by 12.7 points.

  • Data reasoning is harder than health reasoning for almost every model. With the exception of Gemini-3.1-Pro, all models score higher on health reasoning than on data reasoning. The gap is especially large for GPT-4o (53.5% health vs. 25.3% data), Gemma-4-26B-A4B (59.8% vs. 33.8%), and Mistral-Small-3.1 (47.9% vs. 20.0%). Gemini-3.1-Pro is the only model to reverse the trend, at 75.4% data vs. 67.8% health.

  • Cross-signal reasoning is a persistent weak point. Most proprietary models do better on single-signal than cross-signal questions: GPT-5.4 (59.1% vs. 45.7%), Gemini-2.5-Pro (58.1% vs. 45.3%), Claude-Sonnet-4 (59.8% vs. 48.5%). Open-weight models generally score lower on both subsets, with the strongest open model, Gemma-4-26B-A4B, reaching 41.0% single-signal and 43.5% cross-signal. Gemini-3.1-Pro is comparable on both (72.4% vs. 73.2%).

  • Per-question-type variation is substantial. For data reasoning, "strongest pair" and "recovery time" are hardest, while "signal summary" and "excursion count" are easier — Gemini-3.1-Pro scores 68.3% on strongest pair, 94.4% on excursion count, and only 35.9% on recovery time. Health tasks such as risk assessment, multimodal phenotype, and fitness prediction exceed 80% for strong proprietary models, but "cross-signal prediction" ranges from 8.3% (Llama-3.1-8B) to only 40.5% (Claude-Opus-4.6).

  • Chain-of-thought helps, but unevenly. CoT yields substantially larger gains for proprietary models — Claude-Opus-4.6 gains up to 27 points on single-signal and up to 14 points on cross-signal reasoning. Open-weight models gain only modestly, averaging +2.3 points on single-signal and +3.0 points on cross-signal.

  • Positional bias is severe in smaller models under direct answering. Ground-truth answers are balanced across the 10 option positions. Llama-3.2-3B selects a single option for 51.5% of its direct-answer responses; CoT reduces that maximum to 13.5%. Claude-Opus-4.6 goes from 14.5% to 10.6%. Proprietary models are relatively uniform.

  • Models lean on prior knowledge. Comparing definitional signal pairs (related by construction, e.g. steps and active energy expenditure) against empirical pairs (e.g. resting heart rate and stress level) within the same strong-coupling band, several models do far better on definitional pairs: Gemini-2.5-Pro (+51.5 points), GPT-5.4 (+48.2), and GPT-4o (+36.0). This suggests models recognize familiar relationships more reliably than they recover comparably strong user-specific associations from the data.

  • Input format barely matters; tool access matters a lot. For GPT-5.4, changing text serialization (col 50.3, csv 51.5, markdown 49.9, markdown+stats 51.9) yields only marginal changes over the row-wise baseline (51.2). Image representations hurt (image-grid 36.6, image-grouped 34.1), and csv+chart reaches 51.0. Giving the model Python access (agentic, fixed-repr) raises accuracy by 20.1 points to 71.3%; letting it also choose among representations (agentic, multi-view) reaches 69.7% with no further benefit.

  • The benchmark genuinely requires the wearable data. For Claude-Opus-4.6, removing longitudinal history drops history-dependent accuracy from 75.9% to 39.8% while window-scoped accuracy is nearly unchanged (58.1% to 58.0%). Removing all wearable time series drops window-scoped accuracy to 13.7% and overall accuracy from 60.2% to 17.3%.

Methodology in Plain English

The authors started with a large cohort of in-the-wild wearable data and sampled 200 users with high coverage across taxonomies. Each user contributes a multivariate daily time series of 16 wearable metrics spanning cardio-fitness, activity and energy, and sleep domains (plus average stress level), a 17-biomarker blood panel, and demographic attributes (age, sex, BMI, ethnicity). Cohort-reference percentiles for five core metrics are computed over a larger cohort of real-world users than the 200-question cohort.

Questions come from two grounding sources. Literature-grounded questions are drawn from peer-reviewed consumer-wearable studies that the authors verified against primary records (DOI, PMID, PMCID) and read at full-text depth; each retained relationship needed support from at least two independent, same-direction studies, and study-specific thresholds were deliberately not imported because they are cohort-dependent. Population-grounded questions are mined directly from the cohort over a 28-day window and must pass three gates: effect size (a coupling must exceed |ρ| ≥ 0.5 with consistent sign across both halves of the window, and winning options must beat the runner-up by ≥ 0.15), robustness (stability under bootstrap resampling of days in at least 80% of 30 resamples, leave-one-out removal of any single day, and baseline-window variation, with near-boundary cases rejected), and authenticity (a cross-user null test retaining patterns only at FDR ≤ 0.20, i.e. positive predictive value ≥ 80%).

A "discover-then-label" procedure then turns each reasoning objective into an answer by running a deterministic program over the individual's real measurements, built from a shared library of reusable computation primitives (for example, lagged correlation over lags up to ±3 days followed by argmax over signal pairs, or threshold flags on clinical biomarkers). Representing each instance as an explicit computation graph guarantees identical inputs yield identical answers. Distractors differ by reasoning type: data-reasoning distractors come from the same computation (opposite trend direction, adjacent window, or an incorrect magnitude drawn from the cohort distribution), and an item is kept only if the gold option uniquely satisfies the computation; health-reasoning distractors are clinically plausible but inconsistent interpretations of the same relationship.

Validation has two parts. Computational validity means every answer can be independently reproduced. Quality control includes duplicate removal, answer-choice balancing, distractor verification, removal of any context that would reveal the ground truth, review by three models (GPT-5.4, Gemini-3.1-Pro, Claude-opus-4.6), and human verification. The authors also audit for shortcuts; they report that an earlier version contained highly predictable metric pairs (such as active burn with steps) that models could answer without the time series, and that some answer choices appeared only for certain ground-truth categories, allowing inference from the options alone — both issues led them to revise the generation procedure and repeat verification.

For evaluation, each question shows demographics, wearable time series, blood panel, cohort reference distribution, and the options (A–J), with all models queried at temperature 0 and scored by exact match. The default setting uses a row-wise rendering of the most recent 500 days and chain-of-thought prompting; a direct-answer condition differs only in the trailing answer-format directive, so context blocks are byte-identical. Accuracy is reported with 95% Wilson confidence intervals, and unparseable or empty responses count as incorrect. Definitional versus empirical pairs are restricted to the same "strong-sync" band so relationship strength is matched.

Why This Matters

Impact on research. Prior health benchmarks such as MedQA, MedMCQA, PubMedQA, and HealthBench are largely text-based and static; ECG-QA involves waveforms but not longitudinal aggregation; time-series benchmarks such as TimeSeriesExam rely on procedurally generated synthetic signals; and PHIA, which does examine wearable data, evaluates against simulated user trajectories and focuses on retrieval and aggregation. WearableQA differs by using real-user longitudinal measurements jointly with blood biomarkers and demographics, and by separating "can the model compute the number" from "can the model interpret it physiologically."

Real-world applications:

  • Personal health assistants. The paper's motivation is that LLM-based health assistants are emerging, and they need to interpret a user's own noisy daily record rather than generic health knowledge.
  • Lifestyle and behavior assessment. Steps, exercise, sleep regularity, and stress are the raw inputs the benchmark tests models on for trend, anomaly, and recovery reasoning.
  • Clinical triage and risk framing. Health-reasoning question types include prognostic framing, recommendations, differential reasoning, phenotyping, and concordance between two signals.
  • Model and product selection. Because the benchmark shows a 19.6%–72.9% spread and separates data from health reasoning, it can inform which model to deploy for which wearable-analysis task, and reveals that agentic Python access can add 20.1 points for GPT-5.4.

Industry relevance. The work is a collaboration involving Meta and academic institutions (KAIST, Korea University), and the dataset and code are released publicly, so it is positioned as shared infrastructure for evaluating consumer-wearable AI.

Future Directions

  • Closing the cross-signal gap. Cross-signal prediction remains the hardest health-reasoning type (8.3% to 40.5%), and the authors find that cross-signal reasoning stays difficult across model scales. What training or tooling would change this is not established by the paper.

  • Better use of tools and representations. Agentic Python access improved GPT-5.4 substantially while multi-view selection did not, so the design space of how models should be given computation over wearable series remains open.

  • Reducing reliance on priors. The definitional-versus-empirical results show models often answer from memorized signal relationships rather than the user's measurements; how to train or prompt models to prefer the observed data is an open question.

  • Extending and maintaining the benchmark. The authors report that shortcut auditing already required regeneration of some question types, so continued adversarial auditing and expansion to further signal types, biomarkers, and populations is a natural continuation. The paper does not report the effects of model scale within families beyond the models evaluated, nor any ablation of the population-grounded FDR threshold.

Target Audience

This paper is most useful for researchers and engineers building or evaluating LLM systems for health and wearable applications, for benchmark designers interested in deterministic ground-truth construction and shortcut auditing, and for clinical or physiological informatics groups who want to know what current models can and cannot derive from real longitudinal sensor data. Readers with a background in ML evaluation and basic physiology will get the most out of it; general readers can still follow the headline result that the highest-performing model reaches 72.9% while the weakest sits at 19.6% against a 10% chance baseline, and that most models never exceed 60%.

Authors’ abstract

Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.

Read the original paper