IBM and Together AI sign a $240M, multiyear deal to build an AI inference cluster on IBM Cloud, using Nvidia's HGX B300 systems, to support open-source models
Context & Ripple Effects
IBM has been building out an enterprise AI portfolio around its own open-source Granite models and, more recently, access to Anthropic’s Claude in IBM software. The Together AI agreement adds dedicated inference infrastructure on IBM Cloud for open-source models rather than another application-layer model partnership.
The commitment also fits a wider pattern of long-duration compute procurement, from Microsoft’s use of Oracle’s Nvidia GPU capacity to OpenAI’s AWS agreement. IBM and Together AI are making that infrastructure commitment around inference workloads and Nvidia’s HGX B300 systems.
First-order effects
- IBM Cloud gains a multiyear, $240 million inference-cluster deployment, while Together AI gets IBM-hosted Nvidia HGX B300 capacity for serving open-source models.
- Nvidia supplies the HGX B300 systems at the center of the deployment, extending its role from hardware vendor to the foundation of IBM and Together AI’s inference offering.
Second-order effects
- IBM’s enterprise AI portfolio can pair model access and cloud infrastructure more tightly, putting pressure on cloud rivals to match specialized inference capacity for open-source workloads.
- Together AI’s customers gain another route to deploy open-source models on dedicated cloud infrastructure, making service performance and availability a more direct basis for competition among inference providers.
Third-order effects
- Long-term inference commitments are turning cloud AI capacity into a productized service layer: model providers, cloud operators, and GPU vendors are increasingly bundled into a single delivery stack.
- If this pattern persists, competition for enterprise AI workloads will shift beyond access to models toward who controls durable inference capacity and its underlying hardware supply.
The trend: AI infrastructure is moving from general-purpose cloud capacity toward long-term, vertically assembled inference platforms built around specific model ecosystems and GPU systems.