Aug 04, 2026
Qwen 3.8 Benchmarks: What's Actually Verified So Far (2026)
Distributed Inference
Cost Optimization
Qwen 3.8 launched claiming second only to Claude Fable 5. The only scores so far are Alibaba's own. Here's every real data point, and what to watch next.

Alibaba claims Qwen 3.8 is the second-best model in the world. It launched the flagship on August 3 with published pricing and disclosed specs, and for the API flagship there is still no benchmark table behind that claim. The family's official numbers arrived from the open side instead: the 27B's weights shipped August 13-14 with a published model card, and the text-only 2.4T checkpoint followed with card tables of its own, the first official numbers at the Max tier.
This post is the other thing: every verifiable data point that exists on the Qwen 3.8 family, clearly sourced, and updated as real numbers land.
If you're not sure which model the scores even apply to, start with the difference between Qwen 3.8 and Qwen 3.8-Max.
TL;DR
- Official benchmarks for the flagship API: still none. The open releases are different: the 27B's card published scores (61.7 SWE-bench Pro), and the 2.4T checkpoint's card added the first Max-tier tables (GPQA Diamond 92.6). Vendor numbers, but numbers
- Independent data point #1: in a real-world architecture evaluation, Qwen 3.8-Max preview scored 80/100, just behind Kimi K3 at 83
- Independent data point #2: community tracking of Code Arena placed a stealth Qwen 3.8 preview in frontier range on coding preference, a fragile early signal, not a score sheet
- The verified family baseline: Qwen 3.7-Max's published numbers (GPQA Diamond 92.4, SWE-bench Verified 80.4, Terminal-Bench 2.0 69.7) are the floor 3.8 should clear
- Pricing: published at launch. $2 per million input tokens, $6 output, $0.25 cached input on the standard API. What's missing is a verified benchmark to divide it by
- Contrast: Kimi K3 has third-party rankings, published pricing, and open weights. Qwen 3.8 now has an API, a rate card, and the 27B's open weights, but still no proof for the flagship
- What flips the flagship from "wait" to "evaluate": independent numbers. The open checkpoints have model cards now; an independent index run is the real flip
The Official Claim, and What's Behind It
Alibaba's exact words: Qwen 3.8 is "one of the most powerful models available today, comparable to leading frontier AI models, second only to Fable 5," referring to Anthropic's Claude Fable 5.
What's behind it, even after the August 3 launch: internal evaluations. No benchmark names, no scores, no prompts, no harness, no methodology. The claim is not implausible, the Qwen 3.7-Max generation posted genuinely competitive published numbers, but a ranking with no scores attached is marketing until the table ships.
For production teams the distinction matters: you can't size a deployment, estimate cost per task, or compare against your current model using a press quote.
What's Actually Been Tested
Two independent data points exist so far. Both are early, both are limited, and both are more informative than anything official.
A real-world architecture evaluation. An independent evaluator ran Qwen 3.8-Max preview and Kimi K3 through a production-style software architecture task: analyzing 269 files across two real projects and proposing system designs, blind-reviewed and fact-checked. Result: Kimi K3 83/100, Qwen 3.8-Max preview 80/100. Notable detail: none of Qwen's 44 tool calls failed, a good sign for agent reliability. One test, one workload type, but a real one.
Code Arena community tracking. Before the WAIC announcement, a stealth preview model widely believed to be Qwen 3.8 ran on Code Arena, where community analysis placed its coding-preference Elo in frontier range. This is the most fragile kind of evidence, unofficial identification, preference voting rather than task completion, but it is consistent with the architecture eval: 3.8 appears to sit near, not above, the current frontier.
The honest synthesis of both: Qwen 3.8-Max preview looks like a legitimate frontier-class model that trades blows with Kimi K3 rather than dominating it. "Second only to Fable 5" remains unproven.
The First Official Numbers: The 27B's Model Card
The August 13-14 weights release came with the Qwen 3.8 family's first published benchmark results, on the 27B: 61.7 on SWE-bench Pro on the coding side and 84.3 on OSWorld-Verified for computer-use tasks, alongside the spec sheet (28B dense, vision encoder, 262k native context, Apache 2.0). Two caveats. These are Alibaba's own model-card numbers, not independent replications, so the usual rule applies: validate before you build on them. And they say nothing about the flagship; a 28B's scores don't verify a 2.4T model's "second only to Fable 5" claim. What they do give the ecosystem is something checkable: open weights mean anyone can rerun these numbers, which is exactly the accountability the Max still lacks. Our 27B specs guide has the full confirmed picture.
The Verified Baseline: What Qwen 3.7-Max Actually Scored
The best calibration for 3.8 claims is the family's last published score sheet. Qwen 3.7-Max, released May 19 with a full benchmark table, posted: 92.4 on GPQA Diamond, 80.4 on SWE-bench Verified, 69.7 on Terminal-Bench 2.0, 91.6 on LiveCodeBench, and 90.4 on MRCR-v2 128k, competitive with Claude Opus 4.6 across most of the agentic suite. Full breakdown with pricing context in our Qwen 3.7 Max vs Claude Opus 4.6 comparison.
Those are the numbers 3.8-Max needs to clear for the generation jump to be real. When the official table drops, that's the first check worth doing: 3.8 vs 3.7's verified floor, not 3.8 vs the marketing.
Qwen 3.8 vs Kimi K3: The Evidence Gap
The sharpest way to see 3.8's evidence problem is next to its direct rival, launched the same week:
| Evidence | Kimi K3 | Qwen 3.8-Max |
| Official benchmark table | Published | None |
| Third-party index rankings | Artificial Analysis 4th of 189, Vals 2nd of 38 | Not yet indexed |
| Head-to-head independent test | 83/100 | 80/100 |
| Published API pricing | $3/$0.30 in, $15 out | $2 in / $6 out, $0.25 cached |
| Open weights | Released July 27 | Max-class 2.4T checkpoint landed ~Aug 12-13 (text-only, custom license). The 27B shipped Aug 13-14 under Apache 2.0 |
K3 has receipts; 3.8 has a claim. That can flip in a single announcement, and this post will be updated the day it does. Full K3 breakdown in our Kimi K3 guide.
K3 isn't the only rival with receipts: the same evidence gap runs through our Qwen 3.8 vs GLM 5.2 comparison, where GLM's open weights make its published numbers checkable and 3.8's are not yet.
Qwen 3.8 Pricing: What We Know
Short version: pricing landed at launch. The standard API runs $2 per million input tokens, $6 per million output, and $0.25 for cached input, with the Token Plan subscription as the individual-use alternative. That undercuts Kimi K3's $3 in / $15 out on list price. What's still impossible is the math that matters: cost per solved task. Without a verified benchmark, a cheaper token is only a cheaper token.
For family context, Qwen 3.7-Max runs $1.25 per million input tokens and $3.75 per million output on Yotta AI Gateway today. Where 3.8 lands relative to that, and to Kimi K3's aggressive $0.30 cached-input rate, is the pricing question of the quarter. Now that list prices exist on both sides, the missing variable is quality, and we'll run the full cost breakdown when real benchmarks land.
How to Evaluate Qwen 3.8 Without Official Benchmarks
If you can't wait for the table, the playbook is the same one we recommend for every frontier release:
- Build your baseline now. Run your actual workload, your prompts, your tools, your evaluation criteria, against the models you can access today. One OpenAI-compatible API in front of Qwen 3.7-Max, GLM 5.2's family, DeepSeek, and Claude makes side-by-side runs a config change instead of four integrations.
- Get preview access if the workload justifies it. The Token Plan route works today for hands-on testing on subscription credits.
- Slot 3.8's standard API into the same harness now that it's open. Your own eval beats any leaderboard: vendor benchmarks, including the eventual official table, are directional at best. The defensive integration pattern means adding it is a routing rule, not a rewrite.
Frequently Asked Questions
Are there official Qwen 3.8 benchmarks? For the flagship, no: Qwen 3.8-Max still has no published benchmark table, model card, or methodology behind its "second only to Fable 5" claim. The 27B does: its August 13-14 release included a model card with scores, the family's first official numbers.
Is Qwen 3.8 really the second-best model in the world? Unverified. The only independent data so far suggests it's frontier-class but trades blows with Kimi K3 rather than clearly beating it. The ranking rests on Alibaba's internal evaluations.
How does Qwen 3.8 compare to Kimi K3 on benchmarks? In the one independent head-to-head published so far, K3 edged the 3.8 preview 83 to 80 on a real-world architecture task. K3 also has third-party index rankings; 3.8 isn't independently indexed yet.
How much does Qwen 3.8 cost? $2 per million input tokens, $6 per million output, $0.25 for cached input on the standard API launched August 3. For family reference, Qwen 3.7-Max is $1.25 in / $3.75 out per million tokens.
When will official Qwen 3.8 benchmarks be released? Alibaba hasn't said, for the flagship. The 27B's came with its weights on August 13-14. For the Max, the signals to watch are a model card or benchmark table; the Max-class open weights have since landed as Qwen3.8-2.4T-A95B, text-only and 400GB+ even at 1-bit, so independent replication is now possible in principle but expensive in practice.
The 2.4T checkpoint's card adds a second set of official numbers: benchmark tables including GPQA Diamond 92.6, the first official figures at the Max tier. Same rule applies: card numbers are the vendor grading its own homework until an independent index picks the model up.
Should I wait for benchmarks before evaluating Qwen 3.8? Build your evaluation harness now against models you can already access, then slot 3.8's live API into the same harness. Your own workload results beat any leaderboard.
Bottom Line
Everything verified about Qwen 3.8-Max fits in one paragraph: it's a 2.4 trillion parameter multimodal MoE, it's frontier-class in the two independent tests that exist, it trailed Kimi K3 narrowly in the only head-to-head, and every stronger claim than that currently traces back to Alibaba's own press materials.
That's not a knock on the model. It's the state of the evidence, and evidence is what deployment decisions run on. The 27B's weights and model card landed August 13-14; this post updates again the day the flagship's official table does. Until then: the full 3.8-Max launch breakdown covers access and specs, and the Kimi K3 guide covers the rival that already showed its numbers.



