Aug 04, 2026
Qwen 3.8 Benchmarks: What's Actually Verified So Far (2026)
Distributed Inference
Cost Optimization
Qwen 3.8 launched with a claim it's second only to Claude Fable 5, and still not a single published score. Here's every real data point so far, and what to watch next.

Alibaba claims Qwen 3.8 is the second-best model in the world. It launched the model on August 3 with published pricing and disclosed specs, and there is still no benchmark table, no model card, and no score behind that claim.
Meanwhile, "qwen 3.8 benchmark" is one of the most-searched questions about the model, and most of what's circulating is either speculation or repackaged vendor marketing.
This post is the other thing: every verifiable data point that exists on Qwen 3.8-Max as of August 4, clearly sourced, and updated as real numbers land.
If you're not sure which model the scores even apply to, start with the difference between Qwen 3.8 and Qwen 3.8-Max.
TL;DR
- Official benchmarks: none. Alibaba's "second only to Fable 5" ranking rests on internal evaluations it has not published
- Independent data point #1: in a real-world architecture evaluation, Qwen 3.8-Max preview scored 80/100, just behind Kimi K3 at 83
- Independent data point #2: community tracking of Code Arena placed a stealth Qwen 3.8 preview in frontier range on coding preference, a fragile early signal, not a score sheet
- The verified family baseline: Qwen 3.7-Max's published numbers (GPQA Diamond 92.4, SWE-bench Verified 80.4, Terminal-Bench 2.0 69.7) are the floor 3.8 should clear
- Pricing: published at launch. $2 per million input tokens, $6 output, $0.25 cached input on the standard API. What's missing is a verified benchmark to divide it by
- Contrast: Kimi K3 has third-party rankings, published pricing, and open weights. Qwen 3.8 now has an API and a rate card, but still no proof
- What flips this post from 'wait' to 'evaluate': a published benchmark table, a model card, or the promised open weights
The Official Claim, and What's Behind It
Alibaba's exact words: Qwen 3.8 is "one of the most powerful models available today, comparable to leading frontier AI models, second only to Fable 5," referring to Anthropic's Claude Fable 5.
What's behind it, even after the August 3 launch: internal evaluations. No benchmark names, no scores, no prompts, no harness, no methodology. The claim is not implausible, the Qwen 3.7-Max generation posted genuinely competitive published numbers, but a ranking with no scores attached is marketing until the table ships.
For production teams the distinction matters: you can't size a deployment, estimate cost per task, or compare against your current model using a press quote.
What's Actually Been Tested
Two independent data points exist so far. Both are early, both are limited, and both are more informative than anything official.
A real-world architecture evaluation. An independent evaluator ran Qwen 3.8-Max preview and Kimi K3 through a production-style software architecture task: analyzing 269 files across two real projects and proposing system designs, blind-reviewed and fact-checked. Result: Kimi K3 83/100, Qwen 3.8-Max preview 80/100. Notable detail: none of Qwen's 44 tool calls failed, a good sign for agent reliability. One test, one workload type, but a real one.
Code Arena community tracking. Before the WAIC announcement, a stealth preview model widely believed to be Qwen 3.8 ran on Code Arena, where community analysis placed its coding-preference Elo in frontier range. This is the most fragile kind of evidence, unofficial identification, preference voting rather than task completion, but it is consistent with the architecture eval: 3.8 appears to sit near, not above, the current frontier.
The honest synthesis of both: Qwen 3.8-Max preview looks like a legitimate frontier-class model that trades blows with Kimi K3 rather than dominating it. "Second only to Fable 5" remains unproven.
The Verified Baseline: What Qwen 3.7-Max Actually Scored
The best calibration for 3.8 claims is the family's last published score sheet. Qwen 3.7-Max, released May 19 with a full benchmark table, posted: 92.4 on GPQA Diamond, 80.4 on SWE-bench Verified, 69.7 on Terminal-Bench 2.0, 91.6 on LiveCodeBench, and 90.4 on MRCR-v2 128k, competitive with Claude Opus 4.6 across most of the agentic suite. Full breakdown with pricing context in our Qwen 3.7 Max vs Claude Opus 4.6 comparison.
Those are the numbers 3.8-Max needs to clear for the generation jump to be real. When the official table drops, that's the first check worth doing: 3.8 vs 3.7's verified floor, not 3.8 vs the marketing.
Qwen 3.8 vs Kimi K3: The Evidence Gap
The sharpest way to see 3.8's evidence problem is next to its direct rival, launched the same week:
| Evidence | Kimi K3 | Qwen 3.8-Max |
| Official benchmark table | Published | None |
| Third-party index rankings | Artificial Analysis 4th of 189, Vals 2nd of 38 | Not yet indexed |
| Head-to-head independent test | 83/100 | 80/100 |
| Published API pricing | $3/$0.30 in, $15 out | $2 in / $6 out, $0.25 cached |
| Open weights | Released July 27 | Promised within days of the Aug 3 launch, including a 27B; not out yet |
K3 has receipts; 3.8 has a claim. That can flip in a single announcement, and this post will be updated the day it does. Full K3 breakdown in our Kimi K3 guide.
K3 isn't the only rival with receipts: the same evidence gap runs through our Qwen 3.8 vs GLM 5.2 comparison, where GLM's open weights make its published numbers checkable and 3.8's are not yet.
Qwen 3.8 Pricing: What We Know
Short version: pricing landed at launch. The standard API runs $2 per million input tokens, $6 per million output, and $0.25 for cached input, with the Token Plan subscription as the individual-use alternative. That undercuts Kimi K3's $3 in / $15 out on list price. What's still impossible is the math that matters: cost per solved task. Without a verified benchmark, a cheaper token is only a cheaper token.
For family context, Qwen 3.7-Max runs $1.25 per million input tokens and $3.75 per million output on Yotta AI Gateway today. Where 3.8 lands relative to that, and to Kimi K3's aggressive $0.30 cached-input rate, is the pricing question of the quarter. Now that list prices exist on both sides, the missing variable is quality, and we'll run the full cost breakdown when real benchmarks land.
How to Evaluate Qwen 3.8 Without Official Benchmarks
If you can't wait for the table, the playbook is the same one we recommend for every frontier release:
- Build your baseline now. Run your actual workload, your prompts, your tools, your evaluation criteria, against the models you can access today. One OpenAI-compatible API in front of Qwen 3.7-Max, GLM 5.2's family, DeepSeek, and Claude makes side-by-side runs a config change instead of four integrations.
- Get preview access if the workload justifies it. The Token Plan route works today for hands-on testing on subscription credits.
- Slot 3.8's standard API into the same harness now that it's open. Your own eval beats any leaderboard: vendor benchmarks, including the eventual official table, are directional at best. The defensive integration pattern means adding it is a routing rule, not a rewrite.
Frequently Asked Questions
Are there official Qwen 3.8 benchmarks? No. Qwen 3.8-Max launched August 3, 2026, and as of August 4 Alibaba has published no benchmark table, no model card, and no methodology behind its 'second only to Fable 5' claim.
Is Qwen 3.8 really the second-best model in the world? Unverified. The only independent data so far suggests it's frontier-class but trades blows with Kimi K3 rather than clearly beating it. The ranking rests on Alibaba's internal evaluations.
How does Qwen 3.8 compare to Kimi K3 on benchmarks? In the one independent head-to-head published so far, K3 edged the 3.8 preview 83 to 80 on a real-world architecture task. K3 also has third-party index rankings; 3.8 isn't independently indexed yet.
How much does Qwen 3.8 cost? $2 per million input tokens, $6 per million output, $0.25 for cached input on the standard API launched August 3. For family reference, Qwen 3.7-Max is $1.25 in / $3.75 out per million tokens.
When will official Qwen 3.8 benchmarks be released? Alibaba hasn't said. The signals to watch: a model card, a benchmark table, or the promised open-weight release, any of which would enable independent verification.
Should I wait for benchmarks before evaluating Qwen 3.8? Build your evaluation harness now against models you can already access, then slot 3.8's live API into the same harness. Your own workload results beat any leaderboard.
Bottom Line
Everything verified about Qwen 3.8-Max fits in one paragraph: it's a 2.4 trillion parameter multimodal MoE, it's frontier-class in the two independent tests that exist, it trailed Kimi K3 narrowly in the only head-to-head, and every stronger claim than that currently traces back to Alibaba's own press materials.
That's not a knock on the model. It's the state of the evidence, and evidence is what deployment decisions run on. This post gets updated the day the official table, model card, or weights land. Until then: the full 3.8-Max launch breakdown covers access and specs, and the Kimi K3 guide covers the rival that already showed its numbers.



