How the Pentagon uses AI tools from Vannevar Labs, which got a DOD deal worth up to $99M, to scan open-source intelligence, write intelligence reports, and more
In a test run, a unit of Marines in the Pacific used generative AI not just to collect intelligence but to interpret it. Routine intel work is only the start.” www.technologyreview.com/2025/04/11/ 1... Dr Heidy Khlaaf / @heidykhlaaf : Was happy to speak to MIT TR on the use of LLMs for military “sentiment” analysis. Besides LLM “hallucinations”, the data sources used are prone to manipulation that can directly influence an AI's output used to make sensitive and escalatory determinations. — www.technologyreview.com/2025/04/11/ 1... James O'Donnell / @jamesmodonnell : A group of US Marines sailed the Pacific last year while testing a new technology: generative AI that analyzes terabytes of data collected every day in 180 countries across 80 different languages. www.technologyreview.com/2025/04/11/ 1... @ainowinstitute : “Double-checking” AI doesn't cut it, says @heidykhlaaf.bsky.social in @technologyreview.com. When an AI model draws conclusions from opaque, massive data sets, human oversight becomes a myth; and for military intel, a dangerous myth.
Context & Ripple Effects
The Pentagon’s use of Vannevar’s system extends a defense-AI arc that has moved from experiments with LLMs trained on classified information and Project Maven’s targeting work to broader intelligence workflows. It also puts the DoD’s earlier call for greater supplier transparency around AI systems into a more operational setting.
First-order effects
- Vannevar Labs gains a DoD agreement worth up to $99 million and a concrete role in routine intelligence work: scanning open-source material and drafting reports for Pentagon users.
- Marine users can apply generative AI to both collection and interpretation of intelligence, compressing work that would otherwise require analysts to sort large multilingual data flows.
Second-order effects
- The operational use raises the burden on the Pentagon and Vannevar to validate outputs and source integrity, because hallucinations or manipulated open-source inputs can affect sensitive assessments.
- Defense AI vendors competing for intelligence workloads will be judged less on demonstrations alone and more on whether their systems can support analyst review, provenance, and safe deployment.
Third-order effects
- If this use broadens, generative AI is likely to become embedded first in the intelligence production pipeline rather than solely in discrete targeting programs, expanding the governance surface for military AI.
- The central policy question shifts from whether humans are nominally in the loop to whether they can meaningfully audit model-generated interpretations and the underlying data at operational speed.
The trend: This is part of the military’s shift from testing AI for bounded tasks toward integrating generative systems into everyday intelligence analysis under growing demands for accountability.