Together AI's $800M Raise Signals Open-Source Models Are Now Enterprise Infrastructure
Together AI just raised $800 million in Series C funding, reaching an $8.3 billion valuation and crossing a critical threshold: open-weight AI models are no longer experimental alternatives, but the infrastructure enterprises are betting their workloads on. The San Francisco company reported annual bookings exceeding $1.15 billion in its most recent quarter, a figure that places it alongside established enterprise software businesses. The question now confronting every engineering and procurement team is whether the cost premium of closed frontier models from companies like OpenAI remains justified.
What's Driving the Shift Away From Closed AI Models?
Together AI does not build its own foundation models. Instead, it builds the cloud infrastructure that lets enterprises run open-weight models from companies like DeepSeek, Nemotron, MiniMax, Kimi, and GLM on NVIDIA GPU clusters through an OpenAI-compatible API. The company's bet is straightforward: as open-weight model quality converges with proprietary frontier models, the value in AI infrastructure migrates away from whoever controls the model and toward whoever can serve it most cheaply and reliably at scale.
The funding round's composition suggests this bet is landing with investors. Aramco Ventures, the venture arm of Saudi Arabia's state oil company, led the round with participation from NVIDIA, Vista Equity Partners, General Catalyst, Emergence Capital, Schneider Electric's SE Ventures, and others. Open-weight model usage on Together AI's platform tripled over the past twelve months, according to the company.
"Customers building with open models routinely achieve cost reductions of six to twenty times versus equivalent closed-model APIs," stated CEO Vipul Ved Prakash.
Vipul Ved Prakash, CEO at Together AI
Decagon, an enterprise AI company and named Together AI customer, reported cutting its inference costs sixfold after switching to open models. The cost advantage climbs even higher in specific scenarios: configurations involving batch inference and heavily repeated workloads can achieve sixty times cost reduction, where Together AI's inference engine reuses cached computation across similar queries.
How Does Together AI Achieve These Cost Savings?
The cost differential is not purely the result of running cheaper open-weight model weights. It depends on a proprietary inference optimization engine called ATLAS, which stands for Adaptive-Learning Speculator System, that accelerates token generation through a technique called adaptive speculative decoding. Understanding this technology matters because it explains why Together AI's infrastructure advantage is defensible.
Standard speculative decoding, introduced in academic research in 2023, addresses a fundamental bottleneck in how large language models generate text. Autoregressive decoding produces one token at a time, each requiring a full forward pass through the model. This creates a memory-bandwidth bottleneck, meaning GPUs that could be running compute sit idle waiting on memory access. Speculative decoding breaks this sequential chain by using a lightweight draft model to propose several tokens ahead simultaneously; the full target model then verifies the entire block in a single parallel pass.
ATLAS extends this approach in a structurally important way. Rather than a static draft model trained once and deployed unchanged, ATLAS runs two cooperating speculators simultaneously: a heavyweight static speculator trained on a broad corpus, which provides reliable performance regardless of workload, and a lightweight adaptive speculator that continuously updates from real-time production traffic. A confidence-aware controller selects between them at each decoding step and adjusts lookahead dynamically.
The result is a system that gets faster as usage accumulates. In a fully adapted scenario on Arena Hard benchmarks, Together AI reports ATLAS reaching 500 tokens per second on DeepSeek-V3.1, up from 105 tokens per second on an FP8 baseline, roughly a 4x speedup. The company claims this outperforms Groq's Language Processing Unit for adapted workloads. However, independent benchmarking of ATLAS's full-adapted performance against competitors has not yet been published, so teams should run their specific production workloads before locking in infrastructure commitments.
How to Evaluate Together AI for Your Enterprise Workload
- Benchmark Your Specific Use Case: Run your actual production queries and workloads on Together AI's platform before committing to long-term contracts. Company-reported performance figures may not reflect your particular inference patterns, model choices, or traffic characteristics.
- Calculate Your True Cost Baseline: Document your current spending with closed-model APIs like OpenAI, including per-token costs, volume discounts, and any premium pricing for guaranteed latency or priority access. Compare this directly to Together AI's pricing for equivalent throughput.
- Assess Infrastructure Guarantees: Verify that Together AI can commit to the service level agreements your enterprise requires, including uptime guarantees, latency bounds, and capacity reservations for peak traffic periods.
- Evaluate Model Availability: Confirm that the open-weight models you need are available on Together AI's platform and that the company plans to support new models as they emerge in your domain.
What Does 500 Megawatts of Compute Capacity Mean for Enterprise Customers?
The $800 million in equity capital is only part of the story. Together AI separately secured commitments for more than 500 megawatts of compute capacity to be built independently by investors to support the company's expected growth. The company plans to grow its infrastructure footprint roughly fiftyfold over the next five years from that base.
Five hundred megawatts is not a startup-scale number. A single large hyperscaler data center campus typically draws between 100 and 500 megawatts. Together AI is securing power and physical capacity at a scale that would allow it to guarantee supply to enterprise customers planning large, multi-year AI workloads, the kind of guarantee that was previously only available from AWS, Azure, and Google Cloud.
This infrastructure commitment becomes a competitive moat. An enterprise that needs to run ten billion tokens per month at predictable latency with a contractual service level agreement cannot depend on spot GPU markets or under-capitalized providers. Together AI's 500 megawatt commitment is, in effect, a supply-guarantee argument to the enterprise segment that has historically stayed with hyperscalers precisely because only hyperscalers could credibly make that promise.
Why Is Saudi Arabia's Oil Company Leading an AI Infrastructure Round?
The identity of the lead investor deserves analytical attention equal to the dollar figure. Aramco Ventures, the venture arm of Saudi Aramco, the world's largest oil company by revenue, leading an open-weight inference cloud round is not a typical Silicon Valley funding event. It reflects a pattern that crystallized further on the same day Together AI's round was announced: Abu Dhabi's MGX closed its $49 billion AI fund on July 1, exceeding its $45 billion target, with capital from investors spanning the Gulf, North America, Asia, and Europe.
" }Gulf sovereign and sovereign-adjacent capital has moved from an occasional AI-deal participant to a structural pillar of global AI funding in roughly eighteen months. This shift signals that geopolitical actors view AI infrastructure, particularly open-weight model serving, as strategically important. The round's valuation represents a 2.5x step-up from Together AI's February 2025 Series B, which valued the company at $3.3 billion. Prosperity7 Ventures managing director Abhishek Shukla said the company is "headed towards the public markets," the clearest indication yet that Together AI is in the category of businesses where an initial public offering is a matter of timing rather than viability.
Abhishek Shukla