Sources: multiple federal agencies raised concerns about Grok's safety and reliability in recent months, before DOD approved Grok for use in classified settings
The approval follows reporting that Grok had already been used in government-data analysis through a customized deployment, including concerns that it lacked approval at DHS an earlier customized government-data use. That history makes the reported agency concerns relevant beyond a single Pentagon procurement decision.
Days before this report, DOD officials said the provider had accepted the military's “all lawful use” standard for classified systems the classified-use agreement. The new reporting exposes the tension between that access decision and internal confidence in the model's safety and reliability.
First-order effects
DOD can proceed with Grok in classified settings despite reported reservations from multiple federal agencies, placing responsibility for risk controls and operational use with the department and its users.
The reported concerns raise the stakes for Grok's planned availability through GenAI.mil to military and civilian personnel: reliability failures or safety incidents would now occur in a higher-consequence environment.
Second-order effects
Other frontier-model vendors competing for defense work face a clearer trade-off: meeting DOD's permitted-use requirements may be as decisive as model performance, while agencies may press for stronger validation before accepting access decisions.
Federal buyers and integrators will have greater reason to separate authorization to run a model in a classified environment from confidence that it is suitable for particular tasks, data, and user populations.
Third-order effects
If classified deployments continue to advance amid unresolved cross-agency concerns, federal AI governance is likely to become more explicitly risk-tiered: broad platform authorization paired with tighter limits and oversight for specific uses.
The episode tests whether government AI procurement converges on a state-compatible model-provider class defined by policy commitments as well as technical assurance; the durability of that approach depends on operational outcomes and oversight.
The trend: This is one instance of government-gated frontier AI deployment, where access to sensitive systems is increasingly decided through a mix of technical assurance, agency risk assessments, and providers' use-policy commitments.
Most of USG does not want to get stuck with Grok instead of Claude: “Demand from other agencies to use Grok has been anemic, people familiar with the matter said, except in a few cases where people wanted to use it to mimic a bad actor for defensive testing.” 🤦♂️
who would have thought that the AI that once inexplicably became MechaHitler for a week might not be the best AI to trust with classified national security work? [image]
Demand for Grok within some fed agencies not that high, unless you're trying to act like a bad actor for defensive testing: www.wsj.com/politics/nat... [image]
TLDR: Anthropic is “too woke” per the White House. The USAi sandbox only offers Anthropic, Google, and Meta models. Grok has been questioned for reliability and safety by a range of civilian and military officials.
this story is insane. even before their concerns about “mechahitler and sexualized child imagery,” wsj says government insiders believe Xai's models pale in comparison to others great reporting https://www.wsj.com/... [image]
I have reported for several months that despite getting a OneGov deal, Grok has not yet passed safety reviews for the General Services Administration's government-wide AI resource. The agency's position has been that other federal agencies are using it at their own risk.
Federal officials raise alarm about the safety of xAI's Grok chatbot - just one month after Hegseth decided to force the military to integrate Grok into the Pentagon networks. This chatbot has a +90% error rate when reporting the news due to hallucinations. Military slop. [image]