Mustafa Suleyman said Microsoft cut the cost of running PowerPoint’s image model by about 85%. The company achieved that saving by replacing OpenAI image models with MAI, even while Microsoft remained deeply dependent on OpenAI.

Key takeaways

  • Microsoft replaced OpenAI image models with MAI in PowerPoint and Bing after MAI met product-specific quality thresholds; Mustafa Suleyman said PowerPoint’s model-running cost fell about 85%.
  • Control of productivity-software endpoints lets Microsoft route routine requests by quality, latency, and cost, reserving frontier models for workloads that require them.
  • Microsoft is preserving OpenAI as a strategic source of frontier capability while adding MAI, Mistral, and Anthropic as credible alternatives for pricing, capacity, and deployment negotiations.
  • High-frequency Copilot usage and long-term infrastructure commitments make inference margin—not just model leadership—a central competitive advantage.
  • OpenAI and Anthropic are countering distributors’ routing power by building controlled endpoints in ChatGPT and Claude, where they determine how models, context, and connected services are used.

Microsoft kept the alliance intact while putting two embedded workloads—PowerPoint and Bing—through a dispatch test. When MAI met each product’s quality bar at lower cost, it got the volume. By owning the software endpoint, Microsoft could treat frontier quality as a routing threshold and choose a lower-cost model that cleared it.

Microsoft turned dependency into a routing table

OpenAI supplied scarce frontier capability; Microsoft supplied cloud infrastructure, enterprise distribution, and software endpoints. Making OpenAI the default gave Microsoft a short path from research to product while generative AI entered general-purpose software.

Quarterly coverage volume: MicrosoftCoverage of Microsoft by quarter, 2024 Q4 to 2026 Q3: from 113 to 157 articles per quarter, peaking at 238.peak 2381572024 Q42026 Q3
Quarterly coverage · Microsoft · 2024 Q4–2026 Q3 · current quarter projected

As Microsoft embedded generative AI in recurring workflows, it began evaluating each request against quality, latency, and cost. The partnership still provided frontier capability without guaranteeing that OpenAI would serve every generation.

Four months after MAI-Image-2 placed third on Arena, Microsoft was serving MAI in production. For PowerPoint, third place cleared the deployment bar.

Mustafa Suleyman said MAI lowered PowerPoint model-running costs by about 85%

Once MAI cleared PowerPoint’s quality bar, Microsoft could capture that saving on every generation routed to its own model.

Default Copilot features multiply small cost gaps

Microsoft bears a marginal compute cost every time a user invokes a model. In April 2026, it made Copilot’s agentic capabilities generally available and enabled them by default in Word, Excel, and PowerPoint. At that scale, each routing decision changes the product’s unit economics.

OpenAI and Anthropic projected in investor materials that inference costs would exceed half their revenue. Training creates the model, but repeated use determines whether adoption produces attractive margins or widespread bills.

Copilot’s added features trigger more model calls. Microsoft can send routine work to lower-cost models and reserve frontier models for difficult tasks, reducing expense while preserving a higher-capability option.

OpenAI engineers reportedly found a way to more than halve inference costs. If deployed, that saving would let OpenAI contest MAI on price as well as image quality.

Microsoft gains leverage from every added supplier

Microsoft is assembling several suppliers rather than replacing one exclusive partner with another. Its multibillion-dollar agreement with Mistral brings Mistral models into Foundry, Copilot Studio, and Azure Local. The agreement also pairs that integration with European data-center construction. For a planned security product, Microsoft reportedly intended to use models from Anthropic, OpenAI, and Microsoft itself.

Mistral, Anthropic, OpenAI, and MAI give Microsoft credible alternatives when it negotiates inference prices, capacity, and deployment terms.

Google pushed its migration from Assistant to Gemini across most Android devices beyond its previous end-of-2025 target and into 2026. During that delay, Microsoft moved two production workloads onto MAI.

OpenAI and Anthropic are building their own endpoints

OpenAI has expanded ChatGPT to keep more model calls inside a surface it controls. Product recommendations with merchant links bring commercial discovery into the interface, while apps inside ChatGPT bring services including Booking.com, Canva, Spotify, and Zillow into the same surface.

Inside ChatGPT, OpenAI decides when its model is invoked, what context surrounds the request, and where a completed task can lead.

Anthropic also controls more of the workflow through Claude. Claude’s voice mode can use Opus and Sonnet while connecting with Gmail, Slack, Canva, and Notion. Anthropic can keep the user’s context and next action inside its own interface.

Microsoft must earn a return on financed capacity

Alphabet, Microsoft, Amazon, Meta, and Oracle are anchoring AI in racks, power connections, cooling systems, fiber, and financed buildings. A study estimated that their off-balance-sheet debt had grown roughly eightfold since 2022 and exceeded an estimated $1.35 trillion of on-balance-sheet debt.

estimated off-balance-sheet debt across the five companies

AMD’s Helios rack-scale AI system, scheduled to ship to Microsoft and other buyers later in 2026, carries an estimated cost above $5 million per system. Microsoft and other buyers will have to fill that capacity with high-frequency work while controlling the cost of each request.

Cloud operators use long-duration commitments in the market for contracted AI capacity to secure future compute. Those contracts lock in costs before customer demand arrives, increasing the pressure to extract more useful work and margin from each rack.

A production MAI model gives Microsoft a credible alternative in price negotiations with OpenAI. Enterprise customers can receive MAI by default in PowerPoint without changing how they request an image.

Behind PowerPoint’s familiar image button, Microsoft has already flipped the model switch—and cut running costs by about 85%.

From OpenAI commitment to multi-model control

  • 2023-11-20 — Satya Nadella said Microsoft remained committed to OpenAI while announcing that Sam Altman, Greg Brockman, and OpenAI staff would join a new advanced AI research team.
  • 2026-07-22 — Microsoft and Mistral agreed to integrate Mistral models into Foundry, Copilot Studio, and Azure Local.
  • 2026-07-24 — Microsoft replaced OpenAI image-generation models with its own MAI models in PowerPoint and Bing.

Frequently asked questions

Did Microsoft end its partnership with OpenAI?

No. Microsoft remains deeply dependent on OpenAI for frontier capability, but the partnership no longer guarantees that OpenAI will serve every generation inside Microsoft products.

Why did Microsoft deploy MAI if it ranked third on Arena?

PowerPoint did not require the top-ranked general image model; it required a model that cleared its product-specific quality bar. Once MAI did so at lower cost, Microsoft could give it the workload.

What does the reported 85% saving cover?

Mustafa Suleyman said MAI reduced the cost of running PowerPoint’s image model by about 85%. The figure refers to model-running cost, not PowerPoint’s total operating cost or customer pricing.

Which AI model suppliers is Microsoft using?

The piece identifies OpenAI, Anthropic, Mistral, and Microsoft’s own MAI models. Mistral models are being integrated into Foundry, Copilot Studio, and Azure Local, while a planned security product reportedly involved Anthropic, OpenAI, and Microsoft models.

How can OpenAI and Anthropic reduce their dependence on software distributors?

They can own more of the user workflow. ChatGPT and Claude provide direct interfaces with connected apps and services, allowing their developers to control model invocation, context, and the user’s next action.