In 2026, IBM signed a $240 million multiyear agreement with Together AI to run open-model inference on Nvidia HGX B300 systems inside IBM Cloud. At the same time, IBM’s own infrastructure business was shrinking. The company is trying to remain the enterprise doorway while other suppliers take larger roles behind it.

Key takeaways

  • IBM and Together AI signed a $240 million multiyear agreement on August 12, 2026, to build an AI inference cluster on IBM Cloud using Nvidia HGX B300 systems.
  • IBM reported $3.8 billion in infrastructure revenue on July 23, 2026, down 7% year over year.
  • AWS began offering models from Anthropic, Stability AI, AI21 Labs and AWS itself in 2023.
  • Inferact raised a $150 million seed round at an $800 million valuation after its founders created the vLLM inference engine.
  • OpenAI launched its Deployment Company with more than $4 billion in initial investment.

The contract buys relevance before it buys leadership

The agreement divides the service three ways: Nvidia supplies the computing systems, Together AI contributes specialized inference capability, and IBM provides the cloud environment and enterprise route to market.

The phrase “IBM Cloud” can obscure what the agreement assembles. It joins GPU systems, model-serving software, contracted capacity, security requirements and customer support under one commercial surface. The customer does not need each layer to carry the same logo. The customer needs the layers to arrive as one accountable service.

IBM has an immediate reason to assemble that service now. The company reported Q2 revenue of $17.2 billion, up 1% from a year earlier and below estimates, while infrastructure revenue fell 7% to $3.8 billion and Z mainframe revenue fell 42%. CEO Arvind Krishna said customers were shifting spending toward chips. With Together AI, IBM can follow that spending without designing and operating every AI layer itself.

The multiyear term also places the deal in the emerging market for contracted AI capacity, where companies secure usable compute through durable agreements instead of treating it as an interchangeable hourly utility. The duration gives IBM and Together AI time to turn the cluster into operating infrastructure.

Enterprise AI is being unbundled. Cloud incumbents can win by combining trusted distribution, governed deployment, durable compute access and specialist inference, even when outside companies supply the models and hardware.

Neutrality was winning before IBM arrived

In 2023, AWS began offering access to models from Anthropic, Stability AI, AI21 Labs and AWS itself, explicitly positioning its service as a neutral generative-AI platform. IBM joined that established pattern three years later.

The pattern has since spread from model catalogs to complete AI systems. In 2026, Tencent Cloud, DigitalOcean and Alibaba Cloud added OpenClaw support as a service, productizing software they had not developed.

A model-neutral cloud still chooses which systems enter its catalog, which hardware serves them, which security policies govern them and which enterprise contracts wrap them. Outside suppliers expand the catalog while the cloud retains control of the transaction.

Inference turns optimization into its own business

Open weights lower one barrier and expose another. Enterprises can access more models, but every production request still consumes hardware, energy and operating attention. Inference economics reward the operator that extracts more useful work from the same installed systems, even when competitors can obtain the same model.

OpenAI and Anthropic reported inference costs exceeding half of revenue in profitability projections supplied to investors. Serving efficiency therefore governs how much revenue survives after customers use the product. OpenAI engineers later reportedly found a way to more than halve inference costs, illustrating how sharply the economics can change beneath an unchanged endpoint.

An inference operator can turn a small utilization improvement across a large request stream into more margin than a nominally superior but expensive model. The operator determines how much of each inference dollar survives.

Investors have begun funding inference as a business of its own. Inferact raised a $150 million seed round at an $800 million valuation after its founders created vLLM, the open-source inference engine the startup was formed to support and commercialize. That valuation priced the machinery that makes many models economical to serve.

The more interchangeable models become, the more leverage can move to the systems that serve them.

The agreement assigns Together AI the job of turning HGX capacity into reliable inference. IBM can then place that capability inside a commercial relationship enterprises already know how to buy.

Deployment makes the surrounding system more valuable

Model access mattered most while companies were experimenting. Production use introduces workflows, permissions, data dependencies, security reviews and people who must absorb the model’s output into an operating process. European technology groups including SAP, Capgemini, Sopra Steria and OVHcloud reported stronger AI demand as enterprise customers shifted from experimentation to deployment. The purchasing question changes from “Which model can do this?” to “Who will make this work inside our organization?”

Major AI platforms have responded by spending on implementation capacity. AWS created an AI-focused forward-deployed-engineer organization backed by $1 billion. OpenAI launched a Deployment Company with more than $4 billion in initial investment to help organizations build and deploy AI systems. Both moves fund the gap between a functioning model and a functioning business process.

IBM enters that gap with enterprise relationships. Krishna identified cybersecurity fears as a top customer priority in July, giving IBM a concrete issue around which to package infrastructure, governance and support. An enterprise selecting an open model still needs someone to define access, integrate the system and remain responsible when the deployment crosses organizational boundaries.

Deployment teams practice managed model adoption by deciding who may invoke a model, where its output travels and which human process changes because it exists. Open-model availability gives those teams more components; it does not redesign the organization for them.

IBM’s door matters only if someone is accountable behind it

In 2023, cloud providers faced pressure because traditional infrastructure had not been designed for large-scale AI. Hyperscalers began rebuilding it, while on-premises hardware providers found an opening. A software interface never erased the physical divide beneath AI infrastructure.

IBM now faces competitors on every side of that divide. Hyperscalers already distribute third-party models and attach deployment teams to them. Direct inference providers can sell their expertise without IBM in the middle. Private-cloud vendors can bring systems closer to customers that want greater control: HPE and Nvidia introduced a co-developed turnkey private-cloud offering for generative-AI workloads in 2024.

The agreement confirms supply, but customer demand remains unproven. IBM customers can still choose AWS, private-cloud systems or direct providers. IBM’s infrastructure and mainframe declines also make the move partly defensive. Pressure can force a company toward the right architecture without granting it an advantage once it arrives.

IBM must now make the pieces behave like one system. If a customer encounters separate capacity limits, security obligations and support boundaries, the partnership is only a resale chain. If IBM can place one deployment contract around those boundaries while Together AI improves the inference operation beneath it, IBM Cloud becomes more than the building where another company’s equipment runs.

Frequently asked questions

Which open models will the IBM Cloud cluster serve?

Neither the agreement description nor the reported terms name specific open models. The announcement commits to open-model inference, not to a published model catalog.

When will customers be able to use the IBM–Together AI service?

No customer-availability date is disclosed. The agreement is described as multiyear, but it does not specify a launch schedule.

How large is the cluster or how many Nvidia systems will IBM deploy?

The reported agreement identifies Nvidia HGX B300 systems but does not disclose GPU counts, total capacity, power requirements or reserved customer allocation.

What pricing, service-level or support terms will customers receive?

No pricing, SLA, support-boundary or capacity-guarantee terms are disclosed. Those details will determine whether customers experience the offering as one managed service or as separate supplier relationships.

From model-neutral catalogs to IBM’s inference deal

  • 2023 — AWS began offering models from Anthropic, Stability AI, AI21 Labs and AWS itself as a neutral generative-AI platform.
  • 2024 — HPE and Nvidia introduced a co-developed turnkey private-cloud offering for generative-AI workloads.
  • July 23, 2026 — IBM reported Q2 revenue of $17.2 billion, while infrastructure revenue fell 7% to $3.8 billion and Z mainframe revenue fell 42%.
  • August 12, 2026 — IBM and Together AI signed a $240 million multiyear agreement for an IBM Cloud inference cluster using Nvidia HGX B300 systems.

On the old blueprint, IBM’s name sat on the machine at the center. On this one, Nvidia supplies the HGX B300 systems, Together AI runs the inference, open models fill the cluster—and the IBM logo moves from the engine to the door.