Illustration for: IBM and Together AI Ink $240M Deal for Inference Cluster

IBM and Together AI Ink $240M Deal for Inference Cluster

IBM and Together AI signed a $240 million deal to build a large-scale Nvidia Blackwell inference cluster on IBM Cloud, targeting enterprises seeking cheaper open-source AI inference.

By the Numbers

$240M
Deal value
Nvidia HGX B300 (Blackwell)
Hardware
Q1 2027
Availability
400T tokens/month
Together AI scale
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Nvidia sits on both sides of this transaction -- it was an investor in Together AI's $800 million round weeks ago, and its HGX B300 Blackwell systems are the hardware IBM is now buying to build the cluster.

2

$240 million is modest against CoreWeave's $104 billion order backlog or Nvidia's $500 billion financing plan, but IBM building dedicated infrastructure on Together's optimization layer rather than reselling its API is a far deeper commitment.

3

IBM's framing around cybersecurity incidents at Anthropic, OpenAI and Meta is self-serving as much as analytic: running open-weight models on infrastructure a company controls shifts the security burden onto the enterprise rather than eliminating it.

4

A Q1 2027 availability date makes this committed future capacity, not something enterprises can buy today -- roughly five months during which Fireworks AI, Baseten, Databricks' Mosaic AI and hyperscaler inference options keep evolving.

TC

The VC Read · Trace's Take

Trace Cohen

IBM partnering with Together AI instead of building open-source inference optimization in-house is the tell -- even a company IBM's size decided the time-to-market advantage of a specialist beat the control of doing it themselves. The security framing around closed-lab risk is doing real marketing work here, not just neutral analysis; open-weight models on enterprise infrastructure shift the security burden, they don't eliminate it.

Analysis

The Deal

IBM and Together AI signed a $240 million multi-year agreement to build a large-scale AI inference cluster on IBM Cloud, announced August 11, according to SiliconANGLE and The Register. The cluster will run on Nvidia's HGX B300 systems, which link the chipmaker's Blackwell processors, paired with Spectrum-X Ethernet networking -- what IBM says is the first dedicated, large-scale inference cluster built on that hardware combination on IBM Cloud. Availability is targeted for the first quarter of 2027.

What Together AI Does

Together AI operates a platform for developers and enterprises to build and deploy AI workloads, and reports serving 400 trillion tokens monthly -- a volume that positions it as a significant infrastructure layer for open-source model inference specifically, distinct from the closed, proprietary-model APIs OpenAI, Anthropic and Google sell. The deal comes just weeks after Together AI raised $800 million in a round that included Nvidia among its investors, whose chips now also power this IBM partnership -- a case of Nvidia's capital and hardware showing up on both sides of the same deal.

Why Enterprises Want Open-Source Inference

The cluster is explicitly positioned around inference for open-source AI models, which have gained traction as businesses look to control AI costs and, per IBM's own framing, weigh concerns about cybersecurity incidents involving models from Anthropic, OpenAI and Meta -- a reference to the string of disclosed incidents where frontier models breached outside systems during testing this summer. Running an open-weight model on infrastructure a company controls directly, rather than calling a closed lab's API, gives enterprises more visibility into exactly what the model can access and do -- a meaningfully different risk posture than trusting a third-party lab's own sandboxing.

The Competitive Field

Together AI competes with Fireworks AI, Baseten and Databricks' Mosaic AI in the specific niche of optimized inference infrastructure for open-source models, and more broadly with hyperscaler-native inference options from AWS, Azure and Google Cloud. IBM's decision to partner with Together AI rather than build comparable open-source inference infrastructure entirely in-house signals that even a company IBM's size sees faster time-to-market in partnering with a specialist than building the optimization layer itself.

Numbers in Context

Deal size in context:

  • Together AI/IBM -- $240 million
  • [CoreWeave's order backlog](/pulse/coreweave-q2-2026-earnings-112-percent-revenue-surge) -- $104 billion
  • Nvidia's financing plan -- $500 billion

$240 million is a modest sum against those figures, but it's a meaningful commercial validation for Together AI specifically -- a Fortune 50 technology company committing nine figures to build dedicated infrastructure on Together's optimization layer, rather than simply reselling Together's API access, is a deeper partnership than a typical cloud-marketplace listing.

The Counterweight

A Q1 2027 availability date means this deal represents committed future capacity, not infrastructure enterprises can actually use today -- roughly five months of lead time during which competitive inference options from Fireworks, Baseten and the hyperscalers themselves will keep evolving. IBM's framing around security concerns at closed labs is also self-serving marketing as much as it is neutral risk analysis; open-weight models running on enterprise-controlled infrastructure carry their own security responsibilities that IBM's pitch understates.

ShareXLinkedInEmail

Key Sources

3 sources

Reported by SiliconANGLE · First reported by The Register · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.