Anthropic has committed to buy up to 2GW of AMD MI450 capacity from 2027. Investor materials put inference costs above half of revenue at both Anthropic and OpenAI. The larger its API business grows, the more money flows into the machinery beneath it.

Key takeaways

  • Inference is the economic pressure point: costs reportedly exceed half of revenue at Anthropic and OpenAI, so lower API prices must be matched by efficiency gains or thinner margins.
  • Anthropic’s commitment to buy up to 2GW of AMD MI450 capacity from 2027 is a bet that scale, supplier leverage, and system co-design can reduce its marginal cost per token.
  • The SK Hynix request shows that accelerator access alone is insufficient; high-bandwidth memory, packaging, networking, and serving software jointly determine usable capacity.
  • Hardware gains become API margin only when models, caching, batching, routing, scheduling, and failure recovery produce more reliable completed work per dollar.
  • Moving downstack creates a new risk: long-term capacity and custom-chip investments can destroy returns if demand, model architecture, or deployment timing misses expectations.

An API customer brings revenue and a stream of inference expense. If that expense grows too quickly, a better model can produce weaker economics. Anthropic must improve both the endpoint and the machinery that serves it.

Cheaper capability makes the serving stack the competitive frontier

Claude Opus 5 captures the mechanism. Anthropic says the model approaches Fable 5 performance at half the price and made it the default on Claude Max. A near-frontier model can undercut frontier pricing without abandoning much of the capability customers value.

Anthropic prices Opus 5 at $5 per million input tokens and $25 per million output tokens, matching Opus 4.8, while saying the newer model uses fewer tokens for comparable work. Fewer tokens, prompt caching, and better tool use reduce the paid infrastructure required to finish a task.

Inference scales with usage rather than remaining a fixed research expense. Investor materials reported that inference costs exceeded half of revenue at both Anthropic and OpenAI. When a variable cost sits above half of revenue, a price cut requires a corresponding efficiency gain, a lower margin, or both.

Open models such as Google’s Gemma 3 and Cohere’s Command A have sharpened this pressure by returning attention to marginal cost and cost of goods sold. A sufficiently capable alternative can make a premium endpoint earn its price through reliability, integration, governance, or total task economics.

Buyers first use cheaper prediction to obtain the same answer for less money. They then redesign applications around the lower price: expanding context, calling models more often, and delegating longer workflows. Lower token prices enlarge the market for efficient infrastructure because more applications become viable.

Falling unit prices can expand demand while exposing weak serving margins. A vendor that passes every gain to customers becomes a high-growth conduit for hardware revenue; one that holds prices loses workloads to cheaper alternatives. Enterprise buyers respond by routing routine work to cheaper models, reserving premium endpoints for failures and high-stakes tasks, and negotiating on completed-task cost rather than list price.

Anthropic is buying options on its cost base

Anthropic’s request that SK Hynix supply memory for its own chip development sits within a broader procurement system. It accompanies a commitment to buy up to 2GW of AMD MI450 capacity beginning in the first half of 2027, with AMD separately committing to invest up to $5 billion in Anthropic.

Up to 2GW of AMD MI450 capacity committed from 2027
SK Hynix share of the concentrated HBM market

The two moves hedge different risks. Anthropic can explore custom silicon while securing future capacity, diversifying suppliers, and gaining bargaining leverage. The AMD agreement leaves the company materially dependent on outside accelerators, but at a scale large enough to influence commercial terms.

Anthropic is also widening its routes into facilities. Fluidstack raised $830 million at a $7.5 billion valuation after partnering with the lab to help build AI data centers. Amazon’s total investment in Anthropic reached $8 billion in 2024. Hyperscaler capital, neocloud development, accelerator commitments, and memory discussions give Anthropic several paths to the infrastructure beneath its API.

Anthropic is using contracts, specifications, allocation rights, and partner competition to influence its cost base before owning major parts of it. The question is whether it can shape the cost, availability, and configuration of the components that determine its marginal token cost.

HBM can constrain a model vendor even when accelerator supply appears secure. SK Hynix held more than 52% of the market amid competition from Samsung and Micron, and an accelerator without adequate high-bandwidth memory is unusable capacity. A relationship with the dominant supplier gives Anthropic supply visibility and design leverage whether or not a proprietary accelerator ships.

An Anthropic chip still faces design, software, packaging, manufacturing, and deployment risks. The company may instead use the memory relationship to support externally designed systems or improve its position with accelerator partners. That option has value when concentrated suppliers otherwise determine the available configuration and margin.

Model labs integrate through workload-led co-design

Model labs integrate vertically by supplying workload knowledge: context patterns, attention behavior, quantization tolerances, cache reuse, tool-call structure, and throughput requirements. Semiconductor partners translate those requirements into accelerators, memory systems, packaging, networking, and racks.

OpenAI and Broadcom provide the clearest adjacent example. They developed Jalapeño, an LLM-optimized inference chip, from design through manufacturing tape-out in nine months. Early testing was said to show better performance per watt than the current state of the art. Their strategic asset was the ability to turn knowledge of a specific workload into silicon before that knowledge became stale.

Custom silicon must still overcome Nvidia’s estimated 70% share of AI-chip sales. Mature software, broad compatibility, developer familiarity, and deployed capacity can outweigh a theoretical chip-level gain. A specialized accelerator that is difficult to schedule or poorly supported by serving software becomes an expensive demonstration of why systems matter.

Anthropic can focus co-design on workloads with enough volume and predictability to justify specialization. Savings must exceed the engineering cost, execution risk, and flexibility lost. Merchant hardware wins below that threshold; above it, Anthropic pays a recurring tax for generality it no longer needs.

Serving software decides whether silicon becomes API margin

An API business captures hardware savings only when they survive the trip through model architecture, kernels, quantization, caching, batching, routing, scheduling, and failure recovery. Faster chips have little economic value if utilization falls, requests queue unpredictably, or the model consumes more tokens to complete the same work.

Opus 5’s emphasis on fewer tokens and prompt-cache-friendly tool use points directly at this layer. The completed task is the useful unit. A model that finishes comparable work with fewer generated tokens can be cheaper at an unchanged list price, while cached context and compatible request batches reduce the infrastructure those tokens consume.

Inferact, founded by the creators of vLLM to commercialize the open-source inference engine, raised a $150 million seed round at an $800 million valuation. OpenAI engineers reportedly found a way to more than halve inference costs. A major financing round for an inference engine and a lab-level cost reduction above 50% put serving software on the same strategic map as model design.

Agentic workloads strengthen the incentive. Interactive latency dominates when a person waits on each response. Longer autonomous workflows put more weight on throughput, scheduling flexibility, retry behavior, and total task cost. Infrastructure can trade some immediacy for higher utilization when the serving system knows which jobs permit that trade.

Vendors should optimize for cost per useful task. Models, accelerators, and data centers are inputs; the commercial output is a reliable completion such as changed code, assembled research, an executed tool sequence, or an advanced business process. A vendor that measures only tokens optimizes the meter. A vendor that controls the system can optimize the work.

Infrastructure control can improve margins or destroy returns

By moving downstack, Anthropic accepts capital and utilization risk in exchange for less supplier exposure. Long-duration capacity contracts can secure supply and lower unit costs when demand arrives as expected. They can also turn an incorrect demand estimate into years of fixed obligations.

Custom-chip programs carry the same trade-off. Specialization can improve performance per watt and reduce dependence on a dominant supplier. It can also strand engineering work if model architectures change, manufacturing slips, or software fails to support the device. Committed memory still needs accelerators, networking, power, and customer demand around it.

Anthropic must excel at coordination. The lab has to align model releases, token efficiency, hardware availability, HBM supply, data-center delivery, and customer demand closely enough to keep expensive capacity productive. Capacity must arrive when models and customers can use it, and the serving system must keep that capacity occupied without degrading completed work.

The 2GW AMD commitment makes utilization a central operating question: how much paid work can Anthropic route onto secured capacity as it arrives? Better procurement terms improve unit economics only when the company can match infrastructure supply to customer demand.

Restricting model supply cannot repeal substitution

Closed-model vendors can also defend endpoint pricing by limiting the availability or reproduction of alternatives. Anthropic and OpenAI have reportedly lobbied Washington over restrictions on Chinese open-source models, while clashing with other technology companies over the proposal. OpenAI, Anthropic, and Google have also shared information through the Frontier Model Forum to detect adversarial distillation attempts that violate their terms.

Those measures can affect access, compliance costs, and the ease of copying a particular model. The serving-cost contest continues under those rules. DeepInfra already supports more than 190 open models, making inference infrastructure a distribution channel for broad model supply. Google’s Gemma 3 and Cohere’s Command A were reported to run on one or two Nvidia H100s, lowering the hardware threshold for credible alternatives.

Policy may narrow the field while customers continue comparing price-performance within it. A closed provider must still justify its premium through superior economics, reliability, integration, or governance. Anti-distillation protects an artifact; efficient production protects its economics.

Anthropic’s 2GW commitment follows directly from inference consuming more than half of revenue. The API may win the customer; memory allocation, accelerator terms, and serving software determine what that customer is worth. The durable advantage is the system that turns changing models into dependable work at lower marginal cost—even after the model itself stops being scarce.

From rapid model iteration to infrastructure procurement

  • 2026-07-25 — Anthropic released Opus 5, its fourth model release in less than two months, and said it cost half as much as Fable 5.
  • 2026-07-25 — Anthropic said Opus 5 classifiers were expected to intervene around 85% less often than Fable 5 classifiers and made Opus 5 the default on Claude Max.
  • 2026-07-26 — Anthropic’s request for SK Hynix to supply memory chips for its own chip development was reported as confirmed, extending its optimization effort below the model layer.

Frequently asked questions

Why is Anthropic committing to so much AMD capacity?

The agreement secures up to 2GW of MI450 capacity beginning in the first half of 2027. It can diversify Anthropic beyond dominant accelerator suppliers and improve commercial leverage, but only if customer demand keeps the capacity productive.

Why does Anthropic want memory from SK Hynix?

High-bandwidth memory can constrain deployment even when accelerators are available. Working with SK Hynix could give Anthropic better supply visibility and design influence for its own chip effort or externally designed systems.

How can serving software lower Anthropic’s inference costs?

Caching, batching, quantization, routing, scheduling, and better failure recovery can raise utilization and reduce the infrastructure needed for each completed task. The piece argues that silicon-level gains matter only if this software layer converts them into reliable API output.

Does controlling infrastructure mean Anthropic must manufacture its own chips?

No. Anthropic can shape its cost base through specifications, allocation rights, procurement contracts, partner competition, and workload-led co-design without owning fabrication or even shipping a proprietary accelerator.

What is the biggest financial risk in the 2GW strategy?

Utilization risk. Capacity contracts lower unit costs when demand arrives on schedule, but a bad forecast can leave Anthropic paying for years of underused infrastructure.