Alibaba's Qwen3.8-Max Beats GPT-5.6 Sol and Claude Fable 5 on Key Benchmarks
Alibaba's 2.4-trillion-parameter Qwen3.8-Max tops OSWorld, Parametric CAD Bench, and OmniDocBench, while trailing GPT-5.6 Sol on Terminal-Bench. Open weights ship next week.
Alibaba shipped Qwen3.8-Max today, a 2.4-trillion-parameter mixture-of-experts model that activates just 95 billion parameters per request and claims wins over GPT-5.6 Sol and Claude Fable 5 on several major benchmarks. Alibaba’s shares jumped 6% on the news.
The numbers back up the noise. Qwen3.8-Max tops OSWorld-Verified at 86.1, Parametric CAD Bench at 91.5, and OmniDocBench 1.5 at 92.1 — all vision-heavy, real-world tasks rather than synthetic leaderboards. On Terminal-Bench 2.1 it scores 86.6, ahead of Claude Opus 4.8 and Claude Fable 5 (84.6) but behind GPT-5.6 Sol’s 88.8.
Coding is where the jump from its predecessor is starkest. DeepSWE 1.1 moves from 21.6 to 56.6. FrontierSWE climbs from 40.7 to 73.5. JobBench goes from 31.3 to 53.4. Those aren’t incremental gains — they’re the kind of jump that suggests Alibaba retrained rather than fine-tuned.
The model supports a 1M-token context window, outputs up to 131K tokens, and can burn through a 262K-token reasoning budget on hard problems. Five built-in tools ship on the Responses API out of the box: code_interpreter, web_search, web_extractor, t2i_search, and i2i_search — positioning Qwen3.8-Max as an agent-first release, not just a chat model.
Pricing undercuts the US labs by a wide margin: $2 per million input tokens, $6 per million output, and $0.25 for cached input. GPT-5.6 Sol and Claude Fable 5 both charge multiples of that. Alibaba is betting that “good enough at a fifth of the price” wins enterprise contracts even when it doesn’t top every chart.
Open weights for both Qwen3.8-Max and a smaller Qwen3.8-27B variant are promised next week — a move neither OpenAI nor Anthropic has matched at this capability tier. It arrives just weeks after Moonshot’s 2.8-trillion-parameter Kimi K3 open-weight release, underscoring how fast Chinese labs are compressing the gap to the frontier, and doing it in the open.
For teams building agents, the calculus just shifted again: Qwen3.8-Max isn’t the smartest model on every task, but it may be the smartest model per dollar with tool-use built in from day one.