Aug 04, 2026
Qwen 3.8-Max: Specs, Pricing, Benchmark Status, and How to Access It (2026)
Distributed Inference
Cost Optimization
Qwen 3.8-Max is live: 2.4T parameters, 1M context, $2/$6 per million tokens. What's confirmed, what's still unverified, and how to use it today.

Alibaba previewed Qwen 3.8-Max on July 19, 2026, and officially launched it on August 3 with published specs and API pricing. It is the biggest model release of the summer: 2.4 trillion parameters, multimodal, and a claim that it trails only Anthropic's Claude Fable 5. Here is what is confirmed, what is still a vendor claim, and how to start using it today.
TL;DR
- Launched: August 3, 2026, following a July 19 preview at the World AI Conference in Shanghai
- Size: 2.4 trillion parameters, sparse Mixture-of-Experts, roughly 95B active per token
- Multimodal: text plus visual inputs confirmed. Coverage differs on the full list (video, documents, speech, image generation have all been reported); Alibaba has not published a spec sheet
- Context window: 1M tokens, up to 128k output tokens
- Performance claim: "second only to Fable 5," Alibaba's own words. No benchmark table published yet
- Open weights: promised within days of launch, for both Qwen 3.8-Max and a smaller Qwen3.8-27B. Not on Hugging Face yet
- Access today: standard API at $2 per million input tokens, $6 output, $0.25 cached, plus the Token Plan subscription
- API compatibility: OpenAI spec and Anthropic spec, same as the rest of the Max line
- Running Qwen in production now: Qwen 3.7-Max is live on Yotta AI Gateway
What Qwen 3.8-Max Is
Qwen 3.8-Max is the next flagship in Alibaba's Qwen line, previewed two months after Qwen 3.7-Max shipped. Where 3.7-Max was framed entirely around long-horizon agent workloads, 3.8-Max adds a second headline: native multimodality. Text plus visual inputs is confirmed; early coverage also reports video, document, speech, and image generation support, but Alibaba has not published a spec sheet, so treat the full modality list as unsettled.
The stated target workloads are coding, full-stack development, data analysis, and office workflows. That is a direct continuation of the agent positioning the 3.7 line established, now with visual inputs in scope.
The scale claim matters for context: at 2.4 trillion parameters, Qwen 3.8-Max is the second-largest publicly known model, behind Moonshot's Kimi K3 at 2.8 trillion, which launched as an open-weight release the same week. The timing is not a coincidence. The frontier race among Chinese labs is compressing release cycles.
At launch Alibaba disclosed the number that was missing at preview: roughly 95B active parameters per token, about 4 percent of the total. For a sparse MoE model that is the figure that drives serving cost and latency, and it puts 3.8-Max in a lighter serving class per token than Kimi K3's 104B active.
Qwen 3.8 Release Date
Alibaba previewed Qwen 3.8-Max on July 19, 2026, at the World AI Conference in Shanghai, then released it officially on August 3, 2026, with standard API access and published pricing. The preview-only phase is over.
Source: Qwen's official announcement
If you are seeing "Qwen 3.8" and "Qwen 3.8-Max" used interchangeably in coverage, the model shown at WAIC is Qwen 3.8-Max, the flagship tier.
The full split, including the incoming 27B, is in our guide to which Qwen 3.8 version you actually need.
Is Qwen 3.8 Open Source?
Not yet, but the promise now has a date attached.
Every Max-tier Qwen model so far has stayed closed. Qwen 3.7-Max is API-only, and the open-weight line continued separately with Qwen 3.6. This time is different: at launch Alibaba committed to releasing open weights within about a week, and not just for the flagship. A smaller Qwen3.8-27B is going open-weight alongside it, and we've broken down the hardware you'll need to run it.
Update, August 10: the promised week has now passed, neither model is on Hugging Face, and no license has been named. Alibaba has not given a new date. Until a repository and license exist, plan around the API. We will update this post the day the weights land either way.
If you need open weights today, the practical options remain Qwen 3.6, GLM 5.2, and Kimi K3.
Qwen 3.8-Max Benchmarks
Alibaba launched Qwen 3.8-Max without an official benchmark table, and that is worth saying plainly.
Alibaba's exact words: Qwen 3.8 is "one of the most powerful models available today, comparable to leading frontier AI models, second only to Fable 5," referring to Anthropic's Claude Fable 5. That ranking rests on internal evaluations. No benchmark table has been published and no model card exists. Outside of one early community test, every performance figure in circulation is a vendor claim.
For calibration, the verified scores of its predecessor: Qwen 3.7-Max posted 92.4 on GPQA Diamond, 80.4 on SWE-bench Verified, and 69.7 on Terminal-Bench 2.0, competitive with Claude Opus 4.6 on most of the agentic suite. If 3.8-Max improves on that baseline while adding multimodality, the ranking claim is plausible. Plausible is not verified.
The right move for production teams: wait for the benchmark table, then run your own workload against it. Vendor rankings, including this one, are directional at best.
We track every verified data point as it lands in our Qwen 3.8 benchmarks breakdown.
How to Access Qwen 3.8-Max
Qwen 3.8-Max is generally available through Alibaba's standard API at $2 per million input tokens, $6 per million output, and $0.25 for cached input. The Token Plan subscription and the Qoder platforms remain the subscription-style alternative.
The model speaks both the OpenAI API spec and the Anthropic API spec, consistent with the rest of the Max line, so existing client code should port with minimal changes.
For teams that want Qwen-class agent capability in production today, Qwen 3.7-Max is live on Yotta AI Gateway: one API key, OpenAI-compatible and Anthropic-compatible endpoints, routed alongside Claude, DeepSeek, GLM, and the rest of the catalog. We broke down the full deployment decision in how to run Qwen 3.7 in production, and the same logic will apply to 3.8 as access expands.
We also keep a live breakdown of Qwen 3.8 API access options, including Token Plan preview pricing, updated as availability changes.
Qwen 3.8-Max vs Qwen 3.7-Max
What actually changed, based on what has been disclosed:
| Factor | Qwen 3.7-Max | Qwen 3.8-Max |
| Announced | May 19, 2026 | July 19, 2026 (preview), launched August 3, 2026 |
| Total parameters | Undisclosed | 2.4 trillion (sparse MoE) |
| Active parameters | Undisclosed | ~95B per token |
| Modalities | Text-first | Text + visual inputs confirmed; full list unsettled |
| Context window | 1M tokens | 1M tokens |
| Availability | Generally available, published pricing | Generally available, published pricing |
| Open weights | No, confirmed closed | Promised within days, including a 27B variant |
| Benchmarks | Published, third-party comparable | Internal claims only |
The practical read: 3.7-Max is the Qwen you can build on today. 3.8-Max is the one to evaluate the moment real benchmarks appear. If cost is the deciding factor between frontier models, our Qwen 3.7 Max vs Claude Opus 4.6 pricing breakdown shows how the current generation compares.
Qwen 3.8-Max vs Kimi K3
The two biggest model announcements of July landed days apart, and they represent opposite bets.
Moonshot's Kimi K3 shipped first at 2.8 trillion parameters, the largest open-weight model ever announced, with weights released July 27.
Qwen 3.8-Max previewed days later at 2.4 trillion parameters then launched August 3 as a managed API at $2 in / $6 out per million tokens, undercutting K3's $3 in / $15 out, with open weights promised within days.
For production teams, the choice is the same one the GLM 5.2 vs Qwen 3.7-Max matchup posed: open weights give you deployment control, fine-tuning rights, and GPU-level cost management; a closed frontier API gives you the vendor's best model with zero infrastructure lift.
K3 already has real third-party results while Qwen 3.8's numbers remain vendor claims, we break down everything verified so far in our full Kimi K3 guide.
We break the two models down in full in our Qwen 3.8 vs Kimi K3 comparison.
Frequently Asked Questions
When did Qwen 3.8 release? Alibaba previewed Qwen 3.8-Max on July 19, 2026, at WAIC in Shanghai, and released it officially on August 3, 2026, with standard API access and published pricing.
Is Qwen 3.8 open source? Not yet. At launch Alibaba committed to releasing open weights within about a week, for both Qwen 3.8-Max and a smaller Qwen3.8-27B. Neither is on Hugging Face yet and no license has been named.
How big is Qwen 3.8-Max? 2.4 trillion parameters in a sparse Mixture-of-Experts architecture, with roughly 95B active per token. Active parameters are the number that drives serving cost.
How much does Qwen 3.8 cost? The standard API is $2 per million input tokens, $6 per million output, and $0.25 for cached input. The Token Plan subscription is the alternative for individual use.
Is Qwen 3.8-Max better than Qwen 3.7-Max? Alibaba says it is, claiming it trails only Claude Fable 5 overall. No official benchmark table has been published, so the claim cannot be verified yet. Qwen 3.7-Max's published scores remain the verified baseline for the family.
Can I use Qwen 3.8-Max through an API? Yes. Standard API access is live at published pricing, and the model supports both OpenAI and Anthropic API specs. The Token Plan and Qoder platforms remain the subscription route.
Is Qwen 3.8 bigger than Kimi K3? No. Kimi K3 is 2.8 trillion total parameters with 104B active; Qwen 3.8-Max is 2.4 trillion with roughly 95B active. Total parameter count says little about quality or serving cost on its own, especially for sparse MoE models where active parameters matter more.
How can I run Qwen models in production today? Qwen 3.7-Max and Qwen 3.6 Plus are both live on Yotta AI Gateway with OpenAI-compatible and Anthropic-compatible endpoints, and open-weight Qwen 3.6 can be self-hosted on Yotta GPU Pods.
Bottom Line
Two of the three open questions from launch week are now closed: API pricing is published and the specs are disclosed. What remains: an official benchmark table and the actual open-weight drop. Either one turns this from news into a deployment decision.
Until then, the Qwen you can ship on is 3.7-Max, live on Yotta AI Gateway today. Start there, and this post will be updated as 3.8 access expands.



