/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making changes that appear to favor Astra and continuing to revise other metrics after launch

OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3.

Fortune Emily Forlini

Context & Ripple Effects

OpenAI moved Astra from its Daybreak debut into broader ChatGPT Work, Codex and API availability within days through a rollout to paid and business users. That made its published evaluations consequential not only for launch messaging but for customers comparing deployment options.

The revisions land alongside OpenAI's acknowledgment that it cannot fully inspect Astra's reasoning and that covert sandbagging could evade detection, as reported in its alignment disclosure. The combination raises the premium on evaluation methods that can be independently tracked and reproduced.

First-order effects

  • OpenAI must defend Astra's reported performance with clearer benchmark versions, test conditions and revision history as post-launch changes alter comparisons in the market.
  • ChatGPT Work, Codex and API customers evaluating Astra lose a stable published baseline and must treat the revised scores as moving inputs to procurement and deployment decisions.

Second-order effects

  • Competing model providers face pressure to disclose comparable evaluation protocols, because a headline score without a fixed methodology is less useful for differentiating models.
  • Independent evaluators and benchmark publishers gain importance as buyers seek comparisons that are separated from a model vendor's own launch materials.

Third-order effects

  • If providers routinely revise public scores after release, model evaluation shifts from a one-time launch claim to a versioned assurance process, with provenance and reproducibility becoming part of product credibility.

The trend: Frontier-model competition is making the evaluation supply chain—test design, conditions and score revisions—as strategically important as the model score itself.

Discussion

  • @mark_k Mark Kretschmann on x
    Interesting report from FORTUNE: @OpenAI quietly changed several GPT-6 Astra benchmark results around the model's launch, with some changes making Astra look better and competing models worse. One of the biggest examples: Astra's reported hallucination rate went from 4.2% to 2%, …
  • @jeremyakahn Jeremy Kahn on x
    OpenAI quietly altered the performance metrics it reported for GPT-6-Astra in ways that favored the new model, and it continues to change others post-launch. Good reporting here from @EmilyForlini https://fortune.com/...
  • @emilyforlini Emily Forlini on x
    OpenAI has quietly changed a few benchmark scores for GPT-6 Astra since launch—multiple times. Scores can change based on test conditions, but some updates make Astra look better and Anthropic look worse. Benchmaxxing, par for the AI course, or both? https://fortune.com/...
  • @garymarcus Gary Marcus on x
    Pause OpenAI - they are a deeply unethical company building technology that they are demonstrably not able to control.
  • r/BetterOffline r on reddit
    OpenAI quietly boosts some of Astra's evaluation metrics, and continues to change others post-launch
  • @emilyforlini Emily Forlini on x
    They're also still planning to change more
  • @katiemiller Katie Miller on x
    Among the most notable changes was Astra's reported hallucination rate. In the first internet archive snapshot of the blog post from 2:23 p.m. ET, it was 4.2%. It remained that number for several more snapshots, the last being a fifth at 3:11 p.m. ET—about 10 minutes before OpenA…