Logo
FrontierNews.ai

Alibaba's Qwen3.8-Max Claims Second Place, But One Independent Test Says Otherwise

Alibaba's newest flagship AI model, Qwen3.8-Max, arrived with bold claims but minimal proof. Previewed on July 19, 2026, at the World AI Conference in Shanghai, the 2.4-trillion-parameter system claims to rank second only to Anthropic's Claude Fable 5 in capability. Yet Alibaba published no benchmark scores, no technical report, no model card, and no per-token pricing. The timing was strategic: the announcement came just days after Moonshot AI unveiled Kimi K3, a competing model with 2.8 trillion parameters that had just claimed the "largest ever" headline.

What Can We Actually Verify About Qwen3.8-Max?

The gap between Alibaba's claims and independently verified facts is substantial. Half the AI assistants tested did not even know the model existed when it launched, and those that did were already repeating specifications Alibaba never officially published. The company's "second only to Fable 5" ranking comes entirely from Alibaba's own internal evaluations, not from any independent leaderboard or third-party benchmark.

The one independent test that exists cuts against Alibaba's narrative. Trilogy AI's StackPerf benchmark ran Qwen3.8-Max-Preview and Kimi K3 through an identical 269-file codebase analysis under identical constraints. Kimi K3 scored 83 out of 100; Qwen3.8-Max scored 80. While a three-point difference on a single task proves little on its own, this is the only outside test of Qwen3.8-Max available, and it contradicts the vendor's claim of second-place dominance.

What Alibaba has disclosed versus what remains unknown reveals a pattern of strategic omission. The company claims the model handles text, images, video, and documents, and that it should outperform the earlier Qwen3.7-Max on coding, full-stack development, data analysis, and office work. However, none of these capabilities have been independently verified.

How to Access Qwen3.8-Max and What It Costs

  • Access Method: The preview model is available only through Alibaba subscription products, not as a standard pay-per-token API. You must subscribe to Token Plan, Qoder (Alibaba's agentic coding tool), or QoderWork to gain access.
  • Pricing Structure: Alibaba is offering aggressive discounts on the preview. The standard credit coefficient is reduced to 0.05x, a 90% discount, and during off-peak hours (22:00 to 08:00 Singapore time, which overlaps with 07:00 to 17:00 Pacific time), the coefficient drops to 0.01x, a 98% reduction.
  • API Compatibility: The preview endpoint supports both OpenAI and Anthropic API formats, allowing it to integrate directly into tools like Claude Code, Cursor, and OpenCode with only a configuration change.
  • Billing Entity: International Token Plan checkout lists Intelligent Cloud Computing (Singapore) Private Limited as the operator, not an alibabacloud.com property, so users should verify the contracting entity before entering payment details.

For developers seeking a cost comparison, Kimi K3 pricing offers a reference point: $3 per million input tokens and $15 per million output tokens. One early tester reported that four substantial coding runs on Qwen3.8-Max consumed about 6% of a weekly credit quota on the $18-per-month Standard tier.

Why Alibaba's Silence on Benchmarks Matters

Alibaba's decision to preview a flagship model without publishing a single benchmark is unusual and deliberate. The company's own track record contradicts this approach: both Qwen3.7 and Qwen3.6 arrived with detailed launch posts and full benchmark tables. Publishing nothing while claiming second place worldwide represents a strategic choice that raises questions about the model's actual performance.

The missing information extends beyond benchmarks. Alibaba has not disclosed the active-parameter count, a critical metric for understanding serving economics. In a mixture-of-experts (MoE) architecture, only a fraction of total parameters activate on each token. That fraction determines compute cost per token, while the 2.4-trillion total sets the memory required to load the model. Serving economics depend on both numbers, yet Alibaba has revealed only one.

The context window size also remains unofficial. Community developers have borrowed a 1,048,576-token configuration from Qwen3.7-Max's limits, but Alibaba has published nothing confirming this for the new model. Similarly, Alibaba promises that Qwen3.8 will "go open-weight soon," but provides no date, license, or checkpoint name. For a company whose Qwen family built its reputation on open models, this vague commitment requires skepticism.

How Qwen3.8-Max Compares to Its Predecessor and Rivals

The verified baseline for comparison is Qwen3.7-Max, the flagship this preview replaces. Alibaba's published benchmark card for Qwen3.7-Max lists 92.4 on GPQA Diamond (a knowledge benchmark), 80.4% on SWE-bench Verified (a software engineering benchmark), a one-million-token context window, and $1.25 per million input tokens on a limited-time 50% discount, with a list price of $2.50. Any improvement from Qwen3.8 over these numbers remains unproven.

Against Kimi K3, the competitive pressure is clear. Moonshot's model claims 2.8 trillion parameters and has promised to release weights by July 27, 2026. The StackPerf blind test showed Kimi K3 scoring 83 out of 100 on architecture analysis, while also demonstrating stronger tool use in the evaluation. Kimi K3 made 53 tool calls versus Qwen3.8-Max's 44, though none of Qwen's calls failed, and the blind review scored Qwen's tool use at 9 to Kimi's 8.

One important caveat applies to all current assessments: Qwen3.8-Max exists only as a preview endpoint. Alibaba explicitly warns that this infrastructure will eventually be removed or replaced by the formal model, so developers should test on it but not build production systems around it. The company's promise of open weights and a formal release remains undated and unspecified.

The broader pattern emerging from this launch reflects a shift in how Chinese AI labs compete with Western counterparts. Alibaba is claiming frontier-level performance without the transparency that typically accompanies such claims, betting that aggressive pricing and API compatibility will drive adoption before independent verification arrives. Whether that strategy succeeds depends on whether developers trust the vendor's internal evaluations or wait for third-party proof.