By August 2026, Anthropic had confirmed an in-house silicon team and a “multi-chip approach” for Claude, though it had not shown a production chip. Investor documents reviewed by the Wall Street Journal put inference costs above half of revenue at both Anthropic and OpenAI. Those documents establish pressure on serving economics; they do not establish why Anthropic formed the team.

Key takeaways

  • Investor materials reviewed by The Wall Street Journal in April 2026 put inference costs above half of revenue at both Anthropic and OpenAI.
  • Anthropic reported that its Rust-based C-compiler experiment incurred about $20,000 in API costs.
  • On August 22, 2026, Anthropic confirmed that it was laying groundwork for custom semiconductors and had adopted a multi-chip approach.
  • Anthropic’s February 6 compiler experiment used 16 Claude Opus 4.6 agents in parallel across nearly 2,000 sessions.
  • Anthropic said first- and third-party traffic on Mythos-class models would be retained for 30 days.

Business Insider reported on August 5 that Anthropic plans to co-design hardware and models as demand for Claude rises. That work could lower expense, add capacity, improve performance or diversify supply. The available evidence supports a narrower conclusion: Anthropic is adding a hardware-design option while recurring inference already dominates its cost structure, but the company has not said which problem its first chip will solve.

Agents turn projects into compute schedules

Inference recurs every time a customer uses a trained model. The April 2026 investor materials reviewed by the Wall Street Journal put that expense above half of revenue, but they did not disclose Anthropic’s cost by model, customer or workload.

Anthropic’s own engineering work shows how agent use can accumulate. On February 6, the company described using 16 Claude Opus 4.6 agents in parallel to build a 100,000-line Rust-based C compiler over nearly 2,000 sessions.

API cost reported for Anthropic’s C-compiler project

That project is one engineering experiment, not a benchmark for ordinary customer workloads. It does not establish a typical number of model calls per task. It does show that an agent project can generate many paid sessions without requiring a fresh human prompt for every step.

When software continues through a sequence of decisions and tool calls, Anthropic must decide which model serves each step and how much latency, capacity and context the step warrants. The company has not disclosed those routing economics. The useful measure would be cost per completed agent task after routing and caching, rather than the price of one isolated response.

Anthropic’s chip plan is still an option

Anthropic’s “multi-chip approach” leaves open how it will divide work among custom silicon, external accelerators and cloud capacity. Business Insider said the company intends to co-design hardware and models, but its August report did not identify a production schedule, manufacturing partner or first workload.

Anthropic also hired Amir Salek, who ran Google’s TPU business until 2022, for its compute team. The same August reporting described Anthropic’s custom server chip as early-stage and said the company had held preliminary manufacturing discussions with Samsung. Anthropic has not shown that it can replace capacity from hyperscalers or GPU suppliers.

On May 2, The Information reported that Anthropic was in early talks to buy inference chips from UK-based Fractile when they become available in 2027. The report described Fractile as a possible addition to suppliers including Google, Amazon and Nvidia; it did not connect those talks to Anthropic’s internal design program.

Together, the reports point to supply diversification as well as custom design. Anthropic could compare model behavior, routing and hardware architecture inside one decision loop, then reserve custom silicon for a recurrent workload. The public record does not identify that workload, its volume or the savings required to repay the design effort.

Software can move the boundary before silicon ships

On July 1, The Information reported that OpenAI engineers had told colleagues they found a software method capable of more than halving inference costs. The report concerns OpenAI, and it does not establish that Anthropic has achieved the same result.

It does show why Anthropic’s make-or-buy decision remains movable. A software optimization can spread across hardware already installed, while a custom chip requires design work and manufacturing commitments before it serves a request. If software removes much of a workload’s cost first, Anthropic has less reason to internalize that workload in silicon.

Anthropic must compare those paths workload by workload. Custom hardware offers tighter coordination when recurring volume is stable; external suppliers spread development costs across more customers and products. Anthropic has disclosed neither the comparison nor the workloads selected for internal silicon.

Financing remains prospective

By August 2026, Anthropic’s bankers had reportedly told prospective investors that an IPO could raise more than $100 billion at a $2 trillion valuation. Those figures remain prospective and unconfirmed. They describe the capital under discussion, not money Anthropic can already deploy.

Broadcom was separately reported to be in talks for more than $60 billion of debt for an AI-chip financing deal expected to benefit Anthropic and other companies. Those talks remained incomplete. OpenAI and Broadcom had also reportedly discussed financing roughly $18 billion of initial custom-chip production, conditional on Microsoft buying about 40% of the chips.

In the reported OpenAI proposal, Microsoft would convert uncertain demand into contracted purchases that lenders and suppliers could finance. No public report identifies an equivalent anchor buyer for Anthropic’s chip effort. Its design could fail, manufacturing could slip and reserved capacity could arrive after the workload changes.

Access rules constrain model routing

Anthropic’s serving controls govern access as well as placement. On August 22, the company said Mythos 5 was in public beta through Claude Security for Enterprise users and that it was working with providers to embed the model in defensive tools. Anthropic did not offer unrestricted direct access.

A June 10 CyberScoop report said Anthropic’s red-team tests found no universal jailbreaks for Fable 5 and that first- and third-party traffic on Mythos-class models would be retained for 30 days. Those rules define who may submit requests and what records must persist. The cited sources do not link them to Anthropic’s silicon program or claim they reduce inference costs.

They still matter to the serving stack because Anthropic’s routers must enforce model eligibility, enterprise permissions and retention requirements before placing a workload. Access policy changes which requests can run; the evidence does not show that it makes those requests cheaper.

Frequently asked questions

When could Anthropic publicly file for an IPO?

Sources said Anthropic was preparing to file publicly as soon as late August 2026. That timing was reported as prospective, not as a completed filing or an announced IPO date.

Could enterprises keep Anthropic’s required 30-day data retention in their own cloud environments?

Anthropic was reportedly planning a change that would let enterprises retain required 30-day data on their own cloud systems. The evidence characterizes the change as rumored rather than confirmed.

Would buying Fractile inference chips mean Anthropic had produced its own custom chip?

No. The reported Fractile discussions concerned a potential purchase of externally supplied inference chips when they become available in 2027; reporting did not connect that potential supply deal to Anthropic’s internal chip-design effort.

The half-revenue figure puts Anthropic’s August chip announcement in sharper relief. Its compiler experiment shows one way agent work can accumulate a substantial serving bill, while the multi-chip plan gives Anthropic more control over how such work runs. A team, preliminary supplier talks and a co-design strategy still do not establish a cheaper production system. Anthropic now has the meter inside its design loop; the lower bill remains unproven.