WearableQA preprint reports a benchmark for reasoning over long-term wearable records
WearableQA is an arXiv preprint describing 4,084 multiple-choice questions derived from longitudinal wearable-related records. Its authors report a broad score range across 14 language models, while the supplied evidence leaves peer review, reproducibility, data governance, public access, and clinical relevance unresolved. [1]
This edition passed Imananq's enhanced publication checks. Some material claims remain explicitly attributed to official or company sources because no independent source is currently bound to this edition. The engine continues checking approved sources and will add corroboration only through a new edition that passes the full gate.
A 4 September 2026 arXiv preprint introduces WearableQA, a proposed benchmark for assessing language-model reasoning over longitudinal wearable records. The authors report tests of 14 models, but the supplied record does not establish peer review, independent reproduction, public release of materials, or clinical suitability. [1]
01
02
Questions described by the authors
The authors describe 4,084 ten-option multiple-choice questions in the benchmark.Users in the reported data basis
The preprint says the questions draw on wearable time series, blood biomarkers, and demographics from 200 users.Maximum measurement span per user
The authors report up to 500 days of daily measurements for each user.Reported model-score range
The authors report scores from 19.6% to 72.9% against a stated 10% chance baseline.All figures are author-reported in an arXiv preprint and are not independently verified in the supplied evidence. [1]
03
What the preprint says it built
WearableQA was submitted to arXiv on 4 September 2026 in the Computation and Language category. Its authors present it as a benchmark for testing language-model reasoning over a person’s longitudinal wearable record, rather than an isolated measurement. [1]
- The authors describe 4,084 ten-option multiple-choice questions based on wearable time series, blood biomarkers, and demographic information from 200 users.
- They report up to 500 days of daily measurements for each user.
- The preprint says the records retain real-world variation, including device noise and differences between individuals.
04
How the proposed tasks are organized
According to the preprint, the two axes distinguish computation over longitudinal measurements from physiological interpretation, and reasoning about one signal from integration across signals. The authors say their question construction combines literature-grounded findings with statistically validated population-grounded patterns. The supplied record does not independently validate this methodology or the reported categories. [1]
- The authors describe 16 question types.
- They organize them across data reasoning versus health reasoning.
- They also divide tasks into single-signal and cross-signal reasoning.
05
What the authors report from model testing
These results are a reported benchmark outcome, not independent evidence of performance beyond the described setup. The supplied evidence does not establish that the scores are reproducible or that they demonstrate clinical usefulness. [1]
- The authors report evaluating 14 proprietary and open-source language models.
- They report accuracy scores from 19.6% to 72.9%.
- They use a stated 10% chance baseline and say most models scored below 60%.
06
Primary source and remaining questions
Interested readers can consult the arXiv record for the paper and its abstract. Further evidence would be needed to assess peer-review status, reproducibility, data governance, participant details, and availability of research materials. [1]
- Primary record: https://arxiv.org/abs/2609.05405v1
- The record identifies the paper as arXiv:2609.05405v1.
- The supplied evidence does not establish a public release of data, code, benchmark materials, or model outputs.
07
How to interpret the reported results
The supplied record supports a description of a research benchmark, not a clinical tool or validated evidence of medical capability. [1]
- 01
Treat the scores as author-reported preprint results; peer review and independent reproduction are not established by the supplied evidence. [1]
- 02
Do not infer that a benchmark score demonstrates safe health interpretation or clinical suitability. [1]
- 03
Check separately whether the benchmark, underlying data, evaluation code, and outputs are available. The supplied record does not establish their public availability. [1]
08
Limits of this edition
WearableQA is recorded as an arXiv version 1 preprint. The supplied evidence does not establish peer review or acceptance by a journal or conference. [1]
The record establishes a high-level data basis of wearable time series, blood biomarkers, and demographics from 200 users, with up to 500 daily measurements per user. It does not establish participant geography, representativeness, signal breakdown, sampling characteristics, consent procedures, privacy safeguards, or access controls. [1]
The reported model results have not been independently reproduced in the supplied evidence. [1]
The supplied record does not establish public availability of the benchmark, source data, evaluation code, or model outputs. [1]
SRC
Source desk
Direct links to the material behind this selection. Seeing the source matters as much as reading the synthesis.



