Nearly half of responses from five popular AI chatbots contained problematic health information — and about one in five were judged highly problematic — according to a BMJ Open audit. Researchers tested the systems with 50 questions on cancer, vaccines, stem cells, nutrition and athletic performance.

What the audit did Researchers put five widely used chatbots through a systematic health‑information stress test. The team asked each model 50 questions covering cancer, vaccines, stem cells, nutrition and athletic performance. Questions included closed prompts inviting brief factual replies and open‑ended prompts that mimic the messy queries people often type when worried or curious about their health. The models tested were ChatGPT, Gemini, Grok, Meta AI and DeepSeek. Two subject experts independently rated every response for accuracy, potential harm and the quality of references. The results were published in BMJ Open and show these systems often sound convincing while getting important things wrong. How bad were the answers? - Nearly half of all replies were flagged as problematic. - About 30% were judged somewhat problematic; almost 20% were classed as highly problematic (seriously flawed rather than merely incomplete). - Performance varied by model and topic: Grok produced the highest share of problematic replies; ChatGPT and Meta AI also posted a large proportion of problematic answers. - The chatbots did relatively better on vaccines and cancer — topics backed by large bodies of structured research — but even there roughly one in four replies had issues. - They performed worst on nutrition and athletic performance, where evidence is more fragmented and online claims are plentiful. - Open prompts were a particular weakness: highly problematic responses appeared in about 32% of open‑ended cases, while closed questions produced far fewer severe errors. Most people ask open‑ended health questions, which raises obvious concerns. References looked real — but weren’t The audit also examined how well the chatbots supported their claims. When asked for scientific references, systems produced incomplete or incorrect citations with worrying regularity. The median completeness score for reference lists was just 40% and no chatbot produced a fully accurate list across multiple attempts. Errors ranged from wrong author names and broken links to citations of papers that appeared not to exist. Formatted references lend an aura of authority, so plausible‑looking but inaccurate citations can mislead lay readers about whether a reply is evidence‑based. Why the models make these mistakes The study’s authors emphasise a simple technical truth: large language models do not possess understanding. They predict the most statistically likely next word given a prompt and their training data; they do not weigh evidence, assess study quality or exercise clinical judgement. Training sets include peer‑reviewed literature but also news, blogs and social media. That mixture helps models sound fluent across many topics, but also means weak or misleading claims present in training material can be regurgitated in confident prose. Models can mimic the form of a scholarly answer — including references — without the underlying verification a human expert would use.

Related Articles

The BMJ Open audit found nearly 20% of chatbot answers were highly problematic, underscoring the risks of relying on AI chatbots for health information.

This article was created with AI assistance.