OpenAI's New Speed Tier Proves Inference Chips Are Becoming AI's Biggest Battleground
OpenAI has unveiled Ultrafast, a new API tier that runs its flagship GPT-5.6 Sol model 14 times faster than standard speeds, reaching around 750 output tokens per second on specialized hardware from chipmaker Cerebras. The limited preview, which opened on August 13 to select customers, marks a quiet but significant shift in how the AI industry measures progress. For years, the race focused on building smarter models. Now, the real competition is about making those models answer faster.
The speed improvement comes from an unusual piece of engineering. Cerebras builds processors the size of a dinner plate, cut from a single silicon wafer, which allows an entire AI model to sit on one chip rather than being split across multiple Nvidia graphics processing units (GPUs) that must constantly shuffle data between them. Eliminating that internal data traffic is what collapses the delay between a user's prompt and the model's response.
Why Does Speed Matter More Than Raw Intelligence Now?
Until now, anyone wanting genuinely real-time responses had to accept a trade-off: drop down to a smaller, less capable model and get speed, or stick with the smartest model and wait for the answer. Ultrafast is designed to dissolve that choice entirely. OpenAI says it can now deliver frontier-grade reasoning and near-instant answers at once, rather than forcing users to pick between them.
This combination matters most for AI agents, the autonomous software systems the entire industry is chasing. An AI agent that pauses for thirty seconds before each action feels like a demo. One that responds in the time it takes to hold a conversation starts to feel like an actual product. Speed, in other words, is quietly becoming a feature as important as raw cleverness.
"Speed completely changes the call experience for complex work and unlocks synchronous experiences for users that were previously limited by intelligence," said one early tester of Ultrafast.
Early tester, Ultrafast beta program
OpenAI is aiming Ultrafast squarely at time-sensitive work, including incident response and debugging, financial research and fraud detection, real-time customer support and voice interactions, and e-commerce applications. Early testers such as Jane Street, Podium, Basis, and Rogo describe the change as qualitative rather than incremental.
How Is the Inference Chip Market Reshaping AI Competition?
The broader shift reflects a fundamental change in AI economics. As raw model capability begins to plateau, the contest is moving toward who can run those models fastest and most cheaply. This race has lifted inference specialists like Groq and turned latency into a selling point. Rivals such as SambaNova and a clutch of specialist clouds are chasing the same prize, and the market increasingly rewards whoever can make a given model answer soonest, not simply whoever trained the biggest one.
For Cerebras, the OpenAI deal is a marquee endorsement at a crucial moment. The company went public in one of the year's biggest listings but has since struggled to convince the market that wafer-scale ambition translates into durable profit. Powering OpenAI's fastest tier is precisely the kind of validation it needed. It is also a reminder that the exotic chip architectures once dismissed as science projects are now doing real work for the biggest names in AI.
For OpenAI, leaning on Cerebras is also a quiet step away from total dependence on Nvidia, consistent with its work on its own custom silicon. The move signals that the company is hedging its bets across multiple specialized hardware vendors rather than relying solely on Nvidia's dominant position.
What Barriers Are Stopping Other Chip Startups From Competing?
While Cerebras and Groq have gained traction, thousands of AI chip startups across Asia are struggling to compete. The fundamental problem is not talent or ambition; it is access to manufacturing capacity and specialized components. All of these startups are fabless, meaning they design chips but leave the manufacturing to specialist foundries like TSMC and Samsung Foundry. Getting an allocation at TSMC is typically not easy, and capacity is increasingly scarce as global AI compute buildout accelerates.
South Korea's FuriosaAI is one of the few startups in Asia to have secured production for advanced AI chips. The company's flagship AI inference chip, called RNGD, entered mass production in January 2026 on TSMC's 5-nanometer process. Its initial batch of 4,000 units was delivered with assembly partner ASUS, and the company plans to produce another 16,000 units this year, taking its total 2026 output to 20,000 chips. Each RNGD chip costs an estimated $10,000.
Even with strong engineering talent and significant funding, startups face multiple interconnected bottlenecks that go far beyond just securing foundry access:
- Foundry Capacity: TSMC's pure-play foundry market share rose from 69 percent in the fourth quarter of 2024 to 73 percent in the first quarter of 2026, giving the company enormous leverage over which startups get access to advanced manufacturing nodes.
- High-Bandwidth Memory Supply: AI chips need high-bandwidth memory (HBM) so that data can move quickly enough through the chip to run large AI models. SK Hynix, Samsung, and Micron dominate this market, and securing their support is a prerequisite for many startups.
- Advanced Packaging: TSMC's CoWoS advanced packaging, which puts AI chips and high-bandwidth memory together into one chip package, is the best-known example of such specialized assembly. But TSMC can only produce so many CoWoS packages at a time, creating another bottleneck.
- Customer Validation: Startups must begin courting potential customers during chip design, not after engineering is complete, to ensure there is actual market demand for their products.
"HBM is so important. Those are the conditions that you need to win," said Alex Liu, senior vice president of product and business at FuriosaAI.
Alex Liu, Senior Vice President of Product and Business, FuriosaAI
Bengaluru-based Agrani Labs, led by former Intel and AMD executives, is among the few startups in India and Southeast Asia trying to build cutting-edge AI chips. The company has raised $8 million from Peak XV Partners, and its CEO is reportedly in talks to raise more than $100 million. But even with that level of funding and world-class chip design talent, securing TSMC capacity remains uncertain.
For startups that do manage to navigate these constraints, there is a compelling commercial argument. FuriosaAI emphasizes that as a non-US, non-China technology company, it offers geopolitical neutrality. This reduces political risk for governments and enterprises increasingly concerned about semiconductor supply chain security. As AI chips have become central to geopolitical competition, governments are seeking their own alternatives to US and Chinese suppliers.
What Does This Mean for the Future of AI Hardware?
The inference chip market is bifurcating. On one side, companies like Cerebras and Groq are winning high-profile deals with major AI labs by solving the speed problem. On the other side, startups across Asia are fighting for scraps of foundry capacity and memory supply, even when they have superior engineering talent and ambitious funding. The winners will be those who can navigate not just chip design, but the entire supply chain ecosystem.
OpenAI's Ultrafast tier is unlikely to be cheap. Running the top model at 14 times the speed on specialist hardware will probably remain a premium option for latency-obsessed cases rather than a default setting. Even so, the message is plain enough. In the next phase of the AI race, being clever will not count for much if you are also slow.