/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

The US military is running an eight-week exercise of five LLMs trained on classified info to test AI-enabled data for decision-making, sensors, and firepower

Matthew Strohmeyer is sounding a little giddy.  The US Air Force colonel has been running data-based exercises inside the US Defense Department for years. Twitter: @stbaasch Twitter: Stephan Baasch / @stbaasch : @Techmeme @KatrinaManson 😂

Bloomberg Katrina Manson

Context & Ripple Effects

This 2023 exercise is the seed of a procurement arc the related coverage traces forward: the Pentagon first had to find out whether LLMs survive contact with classified data, and only then could it buy. Within months it moved to institutionalizing the testing layer, signing a one-year contract with Scale AI to test and evaluate LLMs for military planning and decision-making.

The exercise also previews the access question that dominates the later coverage: the Pentagon's discussed plans for secure training environments for AI companies and the deals with AWS, Microsoft, Nvidia, Oracle, and Reflection AI to run AI tools on classified networks are the commercial end-state of what Strohmeyer's eight-week drill was prototyping. Project Maven remains the cautionary baseline — the flagship targeting-AI effort that already carries adversary data-poisoning concerns.

First-order effects

  • The five model providers in the exercise get the first structured read on how their systems perform on classified military data for decision-making, sensor, and fires use cases — evidence no commercial benchmark can supply.

Second-order effects

  • A testing market crystallizes around the exercise's methodology: Scale AI's Pentagon contract to test and evaluate military-planning LLMs turns evaluation itself into a funded defense product, and Vannevar Labs' open-source intelligence tooling shows adjacent AI workflows already monetizing inside the same apparatus.

Third-order effects

  • If the exercise's pattern holds, classified-network access becomes the decisive defense-AI moat — the secure-environment discussions and the AWS/Microsoft/Nvidia/Oracle/Reflection deals point to a structure where cloud and model vendors compete for cleared infrastructure rather than for commercial benchmarks, while data-poisoning risk, already flagged around Maven, becomes a core national-security concern.

The trend: Defense AI is moving from one-off experimentation to institutionalized procurement, with classified-data access and formal LLM evaluation becoming the Pentagon's primary levers for onboarding commercial model vendors.

Discussion

  • @stbaasch Stephan Baasch on x
    @Techmeme @KatrinaManson 😂