OpenAI calls GPT-6 Astra the “world's best computer use model”; in tests, it booked DMV appointments and searched for apartments faster than the average person
OpenAI leaders think the company's next generation model, which excels at computer use and coding, may mark a major milestone in AI development.
Context & Ripple Effects
OpenAI moved Astra from an initial Daybreak-program launch into ChatGPT Work, Codex and API availability for multiple paid customer tiers. The company is framing its largest-ever training run at its Texas Stargate site as the foundation for stronger computer-use and coding performance.
The reported tests matter because they focus on completing web-based tasks rather than answering prompts. Public reaction questioned the evidence behind OpenAI's performance claims, underscoring that reproducible task evaluation will shape how enterprise buyers interpret the launch.
First-order effects
- OpenAI's Plus, Pro, Enterprise and Business users gain access to an agent positioned for browser-based work and coding through ChatGPT Work, Codex and the API.
- OpenAI's claims put its evaluation methods under scrutiny, as critics have challenged whether the company disclosed enough detail about the testing setup.
Second-order effects
- Anthropic faces more direct price-and-performance comparison after OpenAI matched Claude Fable 5.1's $10-per-million-input-token and $50-per-million-output-token rates.
- Enterprise buyers can compare vendors on completed workflow speed, not just model output quality, increasing pressure on providers to package agents with deployment and safety controls.
Third-order effects
- If computer-use evaluations become credible purchasing criteria, model competition will shift toward the cost and reliability of completed tasks across software interfaces.
- Astra's 100,000-plus-GPU training run points to AI development becoming more capital-intensive, concentrating advantage among labs able to pair large-scale compute with broad product distribution.
The trend: Foundation-model competition is moving from conversational capability toward workflow-native agents measured by their ability to complete real software tasks at a competitive cost.