Models

Alibaba Qwen3.8 Max Matches Claude Opus 4.8

Alibaba's new Qwen3.8 Max model has caught up to Claude Opus 4.8 on key benchmarks, but its thorough reasoning process significantly drives up operational costs and hallucination rates.

The Decoder5 days agoModels
Illustration generated for this story

Alibaba has updated its flagship AI model with the release of Qwen3.8 Max, which achieved a score of 56 on the Artificial Analysis Intelligence Index. This represents a notable 10-point increase over its predecessor, Qwen3.7 Max, which scored 46. The upgrade positions Qwen3.8 Max on par with Anthropic's Claude Opus 4.8 and ahead of GLM-5.2's score of 51. However, it still trails Moonshot AI's Kimi K3, which scored 57 on the same index.

On the GDPval-AA benchmark, which evaluates models on professional and work-related tasks, Qwen3.8 Max demonstrated a massive leap. It gained 468 Elo points to reach a score of 1,739, overtaking Kimi K3 at 1,685. Currently, only Claude Opus 5 ranks higher on this specific test with a score of 1,852. Despite these gains, the model achieves this performance through a highly resource-intensive process. Qwen3.8 Max requires 64 steps per task compared to the 14 steps needed previously, and its input token usage expanded 15-fold because the benchmark resends the entire chat history at every step.

For developers and enterprise users, this thoroughness translates to slower speeds and higher expenses. Alibaba actually reduced its base API pricing, cutting input tokens from $2.50 to $2.00 per million, output tokens from $7.50 to $6.00 per million, and cache hits from $0.50 to $0.25. Yet, because of the massive increase in token volume, running a single task on the Intelligence Index now costs $1.14. This is more than double the $0.53 cost of Qwen3.7 Max. In comparison, the higher-performing Kimi K3 costs just $0.86 per task, while GLM-5.2 runs at $0.57.

Furthermore, the new model exhibits some performance regressions. On the AA-LCR test, which measures a model's ability to synthesize information from long documents, Qwen3.8 Max dropped 2 points. More concerningly, its AA-Omniscience score—evaluating factual accuracy and the willingness to admit ignorance—fell by 10 points. While its accuracy rate remains steady at around 31 percent, its hallucination rate surged from 23 percent to 40 percent, meaning the model is now far more likely to guess blindly rather than state that it does not know the answer.

This is our own summary of reporting by The Decoder

More in Models