Logo
FrontierNews.ai

Together AI's Model Catalog Now Spans 200 Options, But Pricing Reveals the Real Competition

Together AI's inference platform now hosts over 200 open-source models with pricing that spans nearly two orders of magnitude, from $0.05 to $9.00 per million tokens. The shift reflects a fundamental change in how enterprises choose between speed, cost, and capability when running artificial intelligence workloads in the cloud.

What's Driving the Explosion in Model Choices?

The open-source AI landscape has fractured into competing ecosystems, and Together AI's catalog reflects that reality. Where enterprises once chose between a handful of large language models, or LLMs (AI systems trained on vast amounts of text), they now face a decision matrix that includes model size, origin, pricing tier, and specialized capabilities. Together AI, which positions itself as a neutral host for open-weight models, has become a clearinghouse for this diversity.

The most significant additions in 2026 include models from Chinese AI labs: Kimi K2.6 at $1.20 per million input tokens and $4.50 per million output tokens, GLM 5.2 at $1.40 and $4.40, and MiniMax M2.7 at $0.30 and $1.20. These models compete directly with DeepSeek's offerings, which have become the platform's flagship. DeepSeek V3.1 costs $0.60 for input and $1.70 for output tokens, while specialized reasoning models like DeepSeek R1 command premium pricing at $3.00 and $7.00.

How Should Enterprises Choose Between Cost and Capability?

  • Budget-First Deployments: GPT-OSS 20B at $0.05 per million input tokens serves as the platform's price floor, making it viable for cost-sensitive applications where speed matters less than throughput. GPT-OSS 120B offers a middle ground at $0.15 and $0.60.
  • General-Purpose Workloads: DeepSeek V3.1 and Llama 3.3 70B represent the sweet spot for most enterprises, balancing capability with reasonable pricing. These models handle summarization, classification, and content generation without requiring specialized reasoning.
  • Reasoning and Complex Tasks: DeepSeek V4 Pro at $2.10 and $4.40, or DeepSeek R1 at $3.00 and $7.00, target workloads requiring multi-step problem-solving, code generation, or mathematical reasoning. The higher cost reflects the computational overhead of reasoning tokens.

The pricing structure itself tells a story about inference economics in 2026. Input tokens, which represent the text an enterprise sends to the model, typically cost less than output tokens, which represent the model's response. This reflects the asymmetry in computational cost: generating new text is more expensive than processing existing text. For models like DeepSeek V4 Pro, output tokens cost roughly double the input rate, a ratio that holds across most of Together AI's catalog.

Context window size, which determines how much text a model can process at once, also influences pricing decisions. GLM 5.2 and newer Qwen models support extended context windows up to 256,000 tokens, equivalent to roughly 200,000 words. This capability commands a premium but enables use cases like full-document analysis or long conversation histories that would be impossible with smaller context windows.

Why Are Chinese Models Reshaping the Inference Market?

The emergence of Kimi, GLM, and MiniMax on Together AI's platform signals a shift in competitive dynamics. These models, developed by Chinese AI labs, offer aggressive pricing and capabilities that rival or exceed Western alternatives. Kimi K2.6 and GLM 5.2 both support 256,000-token context windows, matching or exceeding the capabilities of comparable Western models at similar or lower price points.

Together AI's role as a neutral host means it does not favor any particular model or region. The platform offers fine-tuning capabilities, including LoRA fine-tuning, which allows enterprises to customize models for specific tasks without retraining from scratch. It also provides dedicated deployments for organizations requiring guaranteed capacity or isolation, and maintains an OpenAI-compatible software development kit, or SDK, that lets developers switch between models with minimal code changes.

The breadth of Together AI's catalog, now spanning 200 models, reflects a market reality: there is no single best model for all use cases. Enterprises must now evaluate trade-offs between inference speed, output quality, cost, and specialized capabilities like reasoning or extended context. The pricing data from July 2026 shows that this market has matured enough to support models at nearly every price point, from ultra-cheap commodity inference to premium reasoning workloads. For enterprises building AI applications, the challenge is no longer finding a model, but choosing the right one for each specific task.