Why Cheaper AI Models Are Actually Driving Up Demand for Computing Power
When Moonshot AI released its Kimi K3 large language model in mid-July, Wall Street initially feared a repeat of the "DeepSeek moment",the concern that cheaper, more efficient AI models would kill demand for expensive computing hardware. Instead, the opposite is happening. Major investment banks now argue that Kimi K3 is not terminating compute demand but accelerating it, triggering a structural shift in how the computing power sector is valued and traded.
Why Did Wall Street Initially Panic About Kimi K3?
The market's initial reaction on July 17 made intuitive sense. If a Chinese open-source model could approach frontier capabilities at a lower cost, why would companies continue investing billions in computing infrastructure? This logic mirrors the "DeepSeek shock" from early 2025, when DeepSeek's efficient model sparked fears that the entire AI infrastructure boom might be overblown.
But the latest research from UBS, Nomura, Bank of America Merrill Lynch, and Citi delivered a starkly different verdict. These analysts pointed out a fundamental flaw in the panic logic: they were confusing efficiency gains with scale expansion. Kimi K3 represents something different from DeepSeek's approach, and the implications for computing demand are the opposite of what markets initially feared.
What Is the Jevons Paradox, and How Does It Apply to AI?
The key to understanding why cheaper AI models drive up compute demand lies in a 19th-century economic principle called the Jevons Paradox. When steam engines became more efficient, coal consumption actually increased because more industries could afford to use them and more applications became viable. The same dynamic is now playing out in artificial intelligence.
When high-quality AI models become cheaper and more accessible, developers and enterprises deploy more applications and process more tokens, ultimately driving compute consumption higher, not lower. This theory is being validated by real-world data. Within 48 hours of Kimi K3's launch, user requests exceeded the cluster's capacity limits, forcing Moonshot AI to urgently suspend new consumer subscriptions to focus entirely on expanding compute capacity.
"We believe competition and innovation in the global large model market will not stop. As we move closer to Artificial General Intelligence, the application of generative AI on both the consumer and enterprise sides will continue to expand. Frontier AI labs and hyperscale cloud platform companies are likely to continue investing during this phase to maintain their competitive position," noted Duan Bing, Asia-Pacific tech team analyst at Nomura.
Duan Bing, Asia-Pacific Tech Team Analyst at Nomura
How Does Kimi K3 Differ From DeepSeek's Approach?
Multiple Wall Street institutions explicitly distinguished between Kimi K3 and DeepSeek R1 in their research notes. DeepSeek R1 primarily demonstrated "efficiency gains," whereas Kimi K3 highlights the rigid demand for "scale expansion." This distinction is crucial for understanding why K3 will intensify, not reduce, computing demand.
Kimi K3's architectural characteristics inherently elevate pressure on inference, memory, networking, and storage simultaneously. The model boasts a 1-million-token context window, meaning it can process roughly 1 million words at once, capable of handling longer texts, larger codebases, and more complex enterprise documents. It also supports always-on inference for long-chain reasoning and agent tasks, and possesses native vision capabilities covering video, images, game development, and front-end design across multimodal tasks.
The model's Mixture of Experts architecture contains 896 experts, with 16 experts activated per token. According to Moonshot AI, overall scaling efficiency improved by approximately 2.5 times compared to Kimi K2. On the Artificial Analysis Intelligence Index, K3 scored 57, ranking third to fourth globally on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More strikingly, K3 topped the Frontend Code Arena leaderboard created by the University of California, Berkeley, with a score of 1,679, becoming the first open-source model to surpass all proprietary overseas models on this authoritative benchmark.
How Will US AI Companies Respond to Kimi K3?
Bank of America Securities semiconductor analyst Vivek Arya argued that the response from US frontier AI labs would be "not less compute, but more." If Chinese open-source models continue to close the gap, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration. Arya also highlighted a key context: media reports indicate Google's Gemini 3.5 Pro has been delayed by several months, with coding performance failing to meet internal targets, making frontier leadership "increasingly difficult to defend".
"The response from US frontier AI labs would be not less compute, but more," stated Vivek Arya, semiconductor analyst at Bank of America Securities.
Vivek Arya, Semiconductor Analyst at Bank of America Securities
What Are the Practical Implications for Computing Infrastructure?
Kimi K3's large-scale deployment creates immediate pressure across the entire computing supply chain. Currently, inference compute demand in China accounts for over 60% of the total, with daily token call volume reaching 140 trillion. This demand expansion directly drives continued increases in hardware requirements for DDR5 memory and enterprise solid-state drives (eSSD), lifting the entire compute power supply chain.
UBS analyst Timo Arcuri's team noted that open-source models are typically more memory-intensive than proprietary frontier models because their context windows are longer, and KV cache requirements continue to grow in absolute terms even after quantization, making open-source model deployment more dependent on high-bandwidth memory (HBM) and storage. Citi analyst Peter Lee emphasized that Kimi K3's inference-side memory requirements are no less than other frontier models, with the expansion of KV cache volume directly driving up demand for server DDR5 and enterprise solid-state drives. Furthermore, K3's large-scale deployment requires "super-node" cluster configurations with over 64 GPUs.
Steps to Understanding the Computing Demand Surge
- Recognize the Jevons Paradox Effect: When AI models become more efficient and cheaper, they enable new applications and use cases that were previously unaffordable, driving overall compute consumption higher rather than lower.
- Distinguish Between Efficiency and Scale: DeepSeek R1 demonstrated efficiency gains through better algorithms, while Kimi K3 drives demand through architectural scale, requiring more memory, longer context windows, and larger cluster configurations.
- Monitor Supply-Side Constraints: GPU delivery lead times have stretched to 2027, and cloud providers have completed multiple rounds of collective price increases, indicating that supply-side bottlenecks will persist and intensify.
- Track Competitive Responses: US frontier AI labs are likely to respond to Chinese model advances by investing in larger-scale training and inference, further accelerating compute demand across the industry.
What Does This Mean for Computing Prices and Availability?
The surge in compute demand is not an isolated phenomenon; severe supply-side lag further intensifies industry tailwinds. Rents for mainstream high-end GPUs such as Blackwell, H100, and H200 continue to rise overseas, with institutions forecasting the compute price hike cycle will extend into the first half of 2027. Domestic and international cloud providers have completed a second round of collective price increases, with AWS, Google Cloud, Alibaba Cloud, and Tencent Cloud successively raising AI compute prices, fully validating the fundamental supply-demand imbalance in high-end compute.
China's compute leasing market surged 62% year-on-year in the first quarter, reflecting the intensity of demand. The entire supply chain, including memory, chips, and networking, is benefiting from this expansion. With CME compute power futures set to launch, the industry's valuation framework faces a significant overhaul, signaling that compute power is transitioning from a cyclical trading asset to a long-term core asset class with commodity-like attributes and utility-like characteristics.
"Another Jevons Paradox," titled Citi semiconductor analyst Peter Lee's July 19 report, explaining that when high-quality models become cheaper and more accessible, developers and enterprises deploy more applications and process more tokens, ultimately driving compute consumption higher.
Peter Lee, Semiconductor Analyst at Citi
The market's initial panic about Kimi K3 reflects a common misunderstanding about how technological progress affects resource consumption. Rather than rendering computing infrastructure obsolete, cheaper and more capable AI models are accelerating the adoption of AI applications across industries, creating a virtuous cycle of demand that will sustain high computing prices and tight supply conditions well into 2027 and beyond.