Sources: the White House's Office of the National Cyber Director and Commerce Department's CAISI are fighting over which agency should lead AI model evaluations
Context & Ripple Effects
The dispute over AI model evaluations follows years of unresolved tension over how the US should regulate advanced AI: earlier coverage described a split between stronger EU-style rules and concerns about preserving competitiveness.
Related reporting shows the fight is not merely administrative. OpenAI has advocated mandatory cyber-risk evaluations led by CAISI rather than the NSA, while subsequent reporting says CAISI was told to pause publishing model assessments during implementation of a new executive order.
First-order effects
- CAISI and the Office of the National Cyber Director face uncertainty over authority to set, conduct, or coordinate AI model evaluations, delaying a clear operating model for developers and government users.
- AI companies seeking federal guidance on cyber-risk testing must contend with competing institutional preferences rather than a settled lead evaluator.
Second-order effects
- The jurisdictional fight increases pressure on the administration to specify whether evaluations are primarily a commerce-and-standards function or a national-cybersecurity function.
- A pause or reshaping of CAISI assessments could shift the practical influence of companies and agencies that have already backed CAISI-led mandatory evaluations.
Third-order effects
- If authority remains fragmented, US AI assurance could develop through episodic executive-action changes rather than a stable, widely accepted evaluation regime.
- The outcome will help determine whether advanced-model oversight is organized around technical assessment capacity at Commerce, cybersecurity coordination at the White House, or a hybrid structure.
The trend: This is part of a broader move from debating whether advanced AI needs oversight to contesting which institutions will control the evaluation mechanisms behind it.