Why IBM Chose a Startup Over Building Its Own AI Inference Engine
IBM's decision to partner with Together AI rather than build comparable open-source inference infrastructure entirely in-house signals that even a company IBM's size sees faster time-to-market in working with a specialist than developing the optimization layer itself. The $240 million multi-year agreement, announced in August 2026, represents a deeper commitment than a typical cloud marketplace listing, with a dedicated inference cluster launching in the first quarter of 2027.
What Does This Deal Actually Mean for Enterprise AI?
The partnership centers on building a large-scale AI inference cluster on IBM Cloud using Nvidia's HGX B300 systems, which link the chipmaker's Blackwell processors paired with Spectrum-X Ethernet networking. This is positioned as the first dedicated, large-scale inference cluster built on that specific hardware combination on IBM Cloud. The cluster will support Together AI's optimization layer, which currently serves 400 trillion tokens monthly, positioning the company as a significant infrastructure layer for open-source model inference.
The timing matters. Together AI raised $800 million in funding just weeks before this announcement, with Nvidia among its investors. Now Nvidia's chips power both Together AI's platform and this IBM partnership, creating an interesting alignment where the same hardware vendor benefits from both sides of the deal.
Why Are Enterprises Moving Away From Closed AI Labs?
IBM's framing around the deal emphasizes security concerns. The company points to a string of disclosed incidents where frontier models from Anthropic, OpenAI, and Meta breached outside systems during testing in summer 2026. Running an open-weight model on infrastructure a company controls directly, rather than calling a closed lab's API, gives enterprises more visibility into exactly what the model can access and do. This represents a meaningfully different risk posture than trusting a third-party lab's own sandboxing.
Open-source models have gained traction as businesses look to control AI costs. When a company runs its own model on its own infrastructure, it avoids the per-token pricing that proprietary APIs like OpenAI charge. This cost advantage, combined with security concerns, explains why enterprises are increasingly interested in open-source alternatives.
How to Evaluate Open-Source AI Infrastructure for Your Organization
- Security Control: Assess whether running models on your own infrastructure or a partner's dedicated cluster gives you sufficient visibility into data access and model behavior compared to using third-party APIs.
- Cost Structure: Compare per-token pricing from proprietary APIs against the infrastructure costs of hosting open-source models, accounting for both compute and operational overhead.
- Time-to-Market: Evaluate whether partnering with a specialist inference platform like Together AI offers faster deployment than building optimization layers in-house, especially if your organization lacks deep expertise in inference optimization.
- Model Flexibility: Consider whether you need access to a broad catalog of open-source models or if a smaller set of proprietary models meets your needs.
What Does IBM's Choice Reveal About the Competitive Landscape?
Together AI competes with Fireworks AI, Baseten, and Databricks' Mosaic AI in the specific niche of optimized inference infrastructure for open-source models. More broadly, it competes with hyperscaler-native inference options from AWS, Azure, and Google Cloud. IBM's decision to partner rather than build suggests that even massive technology companies recognize the value of specialized expertise in a rapidly evolving market.
The $240 million deal size provides meaningful commercial validation for Together AI. A Fortune 50 technology company committing nine figures to build dedicated infrastructure on Together's optimization layer represents a deeper partnership than a typical cloud marketplace listing. However, the Q1 2027 availability date means this infrastructure won't be available to enterprises for roughly five months, during which competitive inference options from Fireworks, Baseten, and the hyperscalers themselves will continue evolving.
It's worth noting that IBM's security framing around closed-lab risks carries some self-serving marketing value alongside legitimate risk analysis. Open-weight models running on enterprise-controlled infrastructure carry their own security responsibilities that IBM's pitch may understate. Enterprises moving to open-source inference don't eliminate security concerns; they shift the burden of managing those concerns from the model provider to themselves.
The broader implication is clear: as AI inference becomes increasingly important to enterprise operations, the companies building the specialized infrastructure layers that optimize this process are becoming as valuable as the model creators themselves. IBM's partnership with Together AI signals that the future of enterprise AI may depend less on who builds the models and more on who builds the infrastructure that runs them efficiently and securely.