Sources: multiple federal agencies raised concerns about Grok's safety and reliability in recent months, before DOD approved Grok for use in classified settings
Warnings about xAI's safety and reliability preceded Pentagon decision to approve Grok for use in classified settings.
Wall Street Journal
Context & Ripple Effects
The reported approval follows the Pentagon’s stated plan to embed xAI’s Grok-family systems in GenAI.mil, a program intended to serve military and civilian users. A planned direct integration into GenAI.mil made the model’s reliability consequential beyond a limited pilot.
It also comes after a DOD official said xAI accepted an “all lawful use” standard for classified military systems, a condition Anthropic had declined. The new reporting exposes the tension between that procurement alignment and internal federal safety concerns.
First-order effects
DOD’s approval clears Grok for classified settings even as multiple agencies had reportedly raised safety and reliability concerns, making xAI an immediately usable supplier for that environment.
The agencies that flagged concerns and the Pentagon now face a practical need to align any safeguards, evaluation, and operational oversight around the approved deployment.
Second-order effects
The contrast with Anthropic’s refusal of the “all lawful use” standard makes vendor willingness to accept military-use terms a more visible differentiator in defense AI procurement.
Other vendors competing for defense work may face pressure to demonstrate both classified-system readiness and governance processes that can address reliability objections without halting deployment.
Third-order effects
If such approvals continue, classified AI adoption may evolve into a distinct governance track: model-risk concerns remain relevant, but they are weighed alongside operational access, contractual terms, and mission requirements.
That could make ongoing evaluation and accountability after approval more important, since initial clearance would not by itself resolve the reliability issues reported across agencies.
The trend: Defense AI procurement is increasingly sorting frontier-model providers by their willingness and ability to operate under government-defined deployment and use conditions.
Most of USG does not want to get stuck with Grok instead of Claude: “Demand from other agencies to use Grok has been anemic, people familiar with the matter said, except in a few cases where people wanted to use it to mimic a bad actor for defensive testing.” 🤦♂️
I have reported for several months that despite getting a OneGov deal, Grok has not yet passed safety reviews for the General Services Administration's government-wide AI resource. The agency's position has been that other federal agencies are using it at their own risk.
who would have thought that the AI that once inexplicably became MechaHitler for a week might not be the best AI to trust with classified national security work? [image]
Federal officials raise alarm about the safety of xAI's Grok chatbot - just one month after Hegseth decided to force the military to integrate Grok into the Pentagon networks. This chatbot has a +90% error rate when reporting the news due to hallucinations. Military slop. [image]
this story is insane. even before their concerns about “mechahitler and sexualized child imagery,” wsj says government insiders believe Xai's models pale in comparison to others great reporting https://www.wsj.com/... [image]
Demand for Grok within some fed agencies not that high, unless you're trying to act like a bad actor for defensive testing: www.wsj.com/politics/nat... [image]
TLDR: Anthropic is “too woke” per the White House. The USAi sandbox only offers Anthropic, Google, and Meta models. Grok has been questioned for reliability and safety by a range of civilian and military officials.