A test of ChatGPT Health and Claude for Healthcare with data from Apple Health finds the chatbots provided questionable and inconsistent responses
ChatGPT now says it can answer personal questions about your health using data from your fitness tracker and medical records.
Washington PostGeoffrey A. Fowler
Context & Ripple Effects
ChatGPTHealth’s rollout made it possible for users to bring medical records and health-app data into the assistant, following reports that people were already uploading years of records despite privacy and accuracy concerns. This test examines whether that added personal context improves answers in practice.
The findings also extend a recurring record: earlier research found harmful or inaccurate medical responses across ChatGPT, GPT-4, Bard and Claude, while OpenAI has said millions use ChatGPT for health information daily. Earlier medical-answer testing across leading chatbots makes inconsistency with Apple Health data especially consequential.
First-order effects
People seeking guidance from ChatGPT Health or Claude for Healthcare may receive answers that are questionable or vary across the tools even when based on the same Apple Health information.
The test puts immediate pressure on both products’ claims to personalize health guidance from user-provided records and tracker data, rather than merely summarize it.
Second-order effects
Users and healthcare organizations have stronger reason to treat chatbot output derived from personal health data as a prompt for verification, not a dependable assessment; that limits the value of the medical-record and health-app import feature as a stand-alone guidance tool.
Competing healthcare AI products will be judged less on access to richer patient data than on consistency, safety and clear escalation when the data does not support a confident answer.
Third-order effects
As consumer assistants become an interface for sensitive personal records, reliability evaluation will need to cover how models reason over real-world longitudinal data, not just generic medical question benchmarks.
The pattern suggests that integrating health data can widen the gap between an assistant’s personalized presentation and its clinically dependable usefulness unless product safeguards and validation improve.
The trend: Consumer AI is moving from general health-information search toward personalized health-data interfaces, making consistency and safety central competitive constraints.
You can now connect ChatGPT to an Apple Watch. So I imported 29 mil steps and 6 mil heartbeats into the new ChatGPT Health. It graded my heart an F. Cardiologist @erictopol called it “baseless.” Any bot claiming to give health insights shouldn't be this clueless. Even in beta. 🧵 …
I'm sorry but ChatGPT Health is worthless It gave me an F because my RHR is higher than usual (it's not) and I don't do enough steps (I do) I have no idea what's wrong with it. I like ChatGPT a lot but this an absurd. I think OpenAI should pull it down for a bit. [image]
The performance of the newly released ChatGPT Health, via a thorough assessment by @geoffreyfowler with his health data, is very disappointing gift link https://www.washingtonpost.com/ ... [image]
“...when it comes to your fitness tracker and some health records, the new Dr. ChatGPT seems to be winging it. That fits a disturbing trend: AI companies launching products that are broken, fail to deliver or are even dangerous.”
The hardest thing about medicine is every single human being is a special, unique, unreplicatible case. — LLM search engines are uniquely designed to be particularly terrible at the one thing everyone needs from their medical provider: personalized care www.washingtonpost.com/t…
I asked @EricTopol to look at ChatGPT's analysis. His view: “This is not ready for any medical advice.” The bot leaned heavily on Apple Watch VO₂ max estimates — which independent studies show can run ~13% low on average — and treated fuzzy metrics like hard facts.
The more I used ChatGPT Health, the worse its answers got. When I asked it the same heart-health question repeatedly, its analysis changed. My grade bounced back and forth between F and a B. Same data, same body. Different answers. [image]
Spot on, @geoffreyfowler —this highlights the core risk of AI in medicine: opaque black boxes where more data might refine or degrade outputs unpredictably, with no transparency to verify. At Neurosimplicity, we prioritize rigor and reproducibility in neuroscience imaging,