Some US hospitals test if GPT-3 can cut the time staff spend replying to online queries; a study claims the first ChatGPT version replied better than doctors
Pilot program aims to see if AI will cut time that medical staff spend replying to online inquiries
Context & Ripple Effects
This pilot is an early attempt to put generative AI into the patient-message workflow rather than use it solely as a standalone information tool. Subsequent coverage shows that use case broadening from individual clinicians using ChatGPT for patient communication to AI-drafted MyChart replies used by thousands of clinicians.
The appeal is operational as well as editorial: a claimed quality advantage in one study gives hospitals a reason to test drafting assistance, but later reports of harmful and inaccurate medical-model responses show why message quality cannot be assumed from fluency alone.
First-order effects
- Participating hospitals can test GPT-3 as a drafting layer for online patient inquiries, with the immediate objective of reducing staff time spent composing replies.
- The study's claim that an early ChatGPT performed better than doctors on replies raises the profile of AI-assisted patient communication, while leaving hospitals to assess whether that result holds in their own workflow.
Second-order effects
- Patient-portal and clinical-software providers gain pressure to embed reply-drafting capabilities into existing communications tools; the later MyChart deployment illustrates how this use case can move into a dominant work surface.
- Clinicians' work shifts from writing every response from scratch toward reviewing, editing, and deciding when an AI draft is appropriate, making reliability and tone central product requirements.
Third-order effects
- If these pilots scale, patient messaging could become a core workflow-native AI category, with competitive advantage accruing to tools integrated into clinical communication systems rather than general-purpose chatbots alone.
- The pattern also makes evaluation and oversight a durable constraint: evidence of inconsistent or harmful medical answers means adoption is likely to depend on whether systems can support safe human review, not just faster drafting.
The trend: Generative AI is moving from general chat interfaces into high-volume healthcare communication workflows, where its value depends on integration and review rather than autonomous advice.