Sources: the US FDA's AI tool Elsa has fabricated nonexistent studies, misrepresented research, and cannot access relevant documents to assist with review work
White Oak, Md. CNN — To hear health officials in the Trump administration talk, artificial intelligence has arrived …
Context & Ripple Effects
The FDA moved from a pilot to an agencywide generative-AI rollout, with a stated aim of speeding scientific review work by the end of June. The planned all-center deployment made tool reliability consequential across the agency rather than a limited experiment.
The reported problems now test the premise behind the FDA's agencywide AI launch: automation can streamline reviewer workflows only if it can retrieve the relevant material and preserve the evidentiary record.
First-order effects
- FDA reviewers using Elsa must independently verify outputs and locate underlying documents, reducing the time-saving value the deployment was meant to provide.
- Reports of fabricated citations and mischaracterized research create an immediate quality-control issue for any review work touched by the tool, even where staff remain the final decision-makers.
Second-order effects
- FDA leaders face pressure to narrow Elsa's permitted uses or add stronger retrieval, citation-validation, and human-review controls before expanding it into higher-consequence workflows such as inspection prioritization.
- Vendors and other federal AI deployments will be judged less on broad productivity claims than on whether their systems can reliably ground answers in authorized agency records.
Third-order effects
- If similar failures recur, public-sector AI adoption is likely to shift from agencywide chatbot rollouts toward bounded, auditable workflows with explicit accountability for source access and output verification.
- For health regulation, the episode underscores that generative AI's operational fit depends on evidence traceability—not merely faster drafting—which could slow deployment in safety-sensitive functions.
The trend: Government AI programs are moving from rapid deployment targets toward the harder task of operational governance for systems used alongside evidence-based public decisions.