Skip to content
D-CSIL

AI News · 2026-10-05 · 4:00 PM CT

A chatbot just beat web search at triaging lung symptoms

TL;DR

Researchers embedded a GPT-4o-based respiratory chatbot called LungDiag inside WeChat and tested it against ordinary web search in a preregistered randomized trial of 2,400 adults with no medical training. The chatbot group reached 70% overall accuracy versus 55% for web search — with the biggest gap on preliminary diagnosis. Published in Nature Health on October 5, 2026.

A person coughing into their elbow at a desk
Photo: Edward Jenner / Pexels

The test

Between April and June 2025, the team recruited 2,400 adults with no formal medical education across 24 healthcare systems in China. Participants were randomized one-to-one to either the LungDiag chatbot or a browser-based web search control — with dedicated AI and generative AI search features deliberately disabled in the control arm so the comparison would be fair.

The study was single-blind, prospective, non-interventional and multicentre, and it was preregistered — which matters in a field where retrospective evaluations of language models have been criticized for flexible reporting. Each participant worked through simulated clinical vignettes of respiratory conditions, scored on four domains: identifying risk factors and aetiology, reaching a preliminary diagnosis, and choosing the right triage level.

Where the chatbot won

In the compliant-case analysis of 2,176 questionnaires (1,088 per arm), LungDiag reached a model-based overall accuracy of 70.0% versus 55.4% for web search — a 14.6 percentage-point risk difference with a risk ratio of 1.26 and a p-value below 0.001.

The gap was largest on preliminary diagnosis: 81.2% accuracy for the chatbot group against 54.5% for the search group, nearly 27 points. Gains on triage were real but more modest — 53.4% versus 47.1% — suggesting the tool sharpens diagnostic reasoning more easily than it resolves genuinely hard urgency calls.

The honest limits

The authors are upfront about what this trial is not: simulated vignettes are not sick patients, and the study does not show the chatbot reduces delays in seeking care or improves real clinical outcomes. Questions of medical liability, regulation, and the black-box nature of large language models all still loom.

There are also practical tradeoffs. On binary high-urgency triage the chatbot hit 85.6% sensitivity but only 62.1% specificity — safer than missing emergencies, but it could push worried-well users toward already-strained emergency departments. And it cost about a minute per case: median completion time was 527 seconds versus 461 for search.

Why this one matters

LungDiag is not just a general chatbot pointed at medical questions. It is task-constrained to respiratory assessment and pairs GPT-4o with a respiratory knowledge layer derived from the team's earlier multi-centre electronic-health-record work — a retrieval-augmented setup that grounds answers in curated clinical resources. It ran as a WeChat mini-program, on the app more than a billion people already open every day.

The trial was led by corresponding author Jianxing He at the First Affiliated Hospital of Guangzhou Medical University's State Key Laboratory of Respiratory Disease, with funding from the Guangzhou National Laboratory, the Guangdong Natural Science Foundation, and China's National Natural Science Foundation. The analysis code and de-identified participant data are publicly available on GitHub and Zenodo, so the numbers can be checked.