Sources: Microsoft has been rationing GPU access for teams building AI tools since late 2022; the company plans to announce Office 365 GPT-4 tools on March 16
Aaron Holmes / The Information :
Context & Ripple Effects
The arc here runs from access to allocation. In November 2021 Microsoft invited select businesses to run GPT-3 as an Azure tool, positioning itself as an open gateway to OpenAI models even as OpenAI sold its own API. By February 2023, reporting had Microsoft readying its Prometheus model integration across Word, Outlook, and other Office apps.
What changed today: sources say Microsoft has been quietly rationing GPU access among its own teams since late 2022 so that enough capacity exists to ship the Office 365 GPT-4 tools it plans to announce on March 16. The scarcest input in AI is being steered from an open Azure storefront toward Microsoft's flagship product line — and later coverage shows that squeeze extending outward, with startups unable to secure Nvidia GPUs as cloud providers divert supply to internal teams and large customers like OpenAI.
First-order effects
- Microsoft's internal AI teams are now competing for a fixed pool of GPUs, with Office 365's GPT-4 push effectively winning priority over other product groups' experiments.
- Azure customers and partners who assumed GPT-series models would be broadly available on demand face tighter, slower access as Microsoft reserves capacity for its own launch.
Second-order effects
- Startups priced out of Microsoft's cloud are pushed toward rival providers — a dynamic later reporting makes explicit, with AI companies struggling to obtain Nvidia GPUs as Microsoft and other clouds divert supply inward.
- Capacity scarcity strengthens the business case for Microsoft's in-house Athena AI chip, tested by Microsoft and OpenAI staff, reducing dependence on Nvidia's constrained supply.
Third-order effects
- If hyperscalers systematically prioritize their own products and anchor customers like OpenAI over the open market, frontier compute allocation becomes a competitive moat rather than a commodity service — reshaping which startups can build at all.
- Sustained rationing pushes the whole industry toward custom silicon and long-term capacity commitments, converting AI infrastructure into utility-like infrastructure controlled by a few allocators.
The trend: Cloud providers are shifting from selling open access to frontier AI models toward rationing scarce compute in favor of their own products and marquee partners.