IBM and Together AI sign a $240M, multiyear deal to build an AI inference cluster on IBM Cloud, using Nvidia's HGX B300 systems, to support open-source models
IBM (IBM.N) and startup Together AI have signed a $240 million multi-year agreement to build a large-scale artificial intelligence cluster …
Context & Ripple Effects
IBM has been assembling enterprise AI offerings since launching the watsonx suite, and later moved to make Anthropic’s Claude available in its business-focused development environment. The Together AI agreement adds dedicated inference infrastructure on IBM Cloud to that broader enterprise-AI push.
Together AI entered the deal after reports that it was pursuing a major funding round amid sharply higher annualized revenue, making the IBM commitment a concrete infrastructure partnership alongside its growth plans.
First-order effects
- IBM and Together AI will build an IBM Cloud inference cluster using Nvidia HGX B300 systems, creating capacity intended to support open-source models under their multiyear $240 million agreement.
- Together AI gains a committed cloud infrastructure partner, while IBM gains a route to offer inference capacity tied to Together AI’s open-model focus.
Second-order effects
- IBM Cloud’s AI offering becomes more directly comparable with cross-provider capacity arrangements such as Microsoft’s access to Oracle’s Nvidia GPU capacity, increasing pressure on cloud providers to secure differentiated AI infrastructure partnerships.
- Nvidia benefits from the cluster’s use of HGX B300 systems as IBM and Together AI turn model-serving demand into a long-duration hardware deployment.
Third-order effects
- The deal reinforces a market structure in which cloud providers pair owned infrastructure with specialized AI companies to commercialize inference, rather than relying solely on internally developed models and services.
- As enterprise AI stacks combine model access, development tools and dedicated serving capacity, the value of an integrated cloud offer may increasingly depend on the economics and availability of inference.
The trend: AI infrastructure is shifting toward multiyear cloud-specialist partnerships that package dedicated GPU capacity around inference and open-model services.