Sep 29, 2026
Qwen 4 Max vs Plus vs Flash vs 27B: The Four Tiers Explained (2026)
Cost Optimization
Alibaba named four Qwen 4 tiers on stage at Apsara and confirmed nothing else. Here's what each tier is for, what the Qwen 3.8 lineup tells you about it, and what to have ready.

Four names, zero spec sheets. What Max, Plus, Flash, and 27B will probably mean, based on how Alibaba has shipped every Qwen generation so far.
TL;DR
Alibaba's Qwen 4 announcement at the Apsara Conference on September 22, 2026 came with four tier names and nothing to attach to them. The official line is that Qwen 4 is in training and coming "very soon." On stage, Qwen lead Liu Dayiheng showed a lineup of Qwen 4 Max, Qwen 4 Plus, Qwen 4 Flash, and Qwen 4 27B. No parameter counts, no pricing, no context windows, no licenses, no dates. If you searched for any of those and landed here, that's the answer: they don't exist yet.
What does exist is the pattern. Alibaba has used the same tier names for three generations, and the roles have stayed stable. Max is the flagship API model. Plus is the cheaper hosted middle tier. Flash is the low-latency, high-volume model. The 27B is the one you download. This post walks through each, using Qwen 3.8 as the reference point, and tells you what to set up now so you can test on day one.
Qwen 4 tier status at a glance
| Tier | What Alibaba has said | Qwen 3.8 precedent | Open weights? |
| Qwen 4 Max | Named on stage, no specs | Qwen 3.8-Max: 2.4T MoE, ~95B active, 1M context, $2 / $6 per million | 3.8-Max weights arrived Aug 12-13, ten days after launch. Nothing said for Qwen 4. |
| Qwen 4 Plus | Named on stage, no specs | Qwen 3.7-Plus: hosted mid-tier, 1M context at a higher rate above 256K | Not in prior generations |
| Qwen 4 Flash | Named on stage, no specs | Qwen 3.8-Flash-Next: 125B total, 6B active, 51B n-gram layer, 262K native context | Flash-Next shipped as open weights under qwen-community-1.0 |
| Qwen 4 27B | Named on stage as the open-weights tier | Qwen 3.8-27B: dense 27B, Apache 2.0, 262K context, vision encoder | Yes, that's the point of the tier |
Everything in the middle column is inference from precedent, not a Qwen 4 fact. Treat it that way.
Qwen 4 Max
Max is the model Alibaba benchmarks against GPT and Claude. In the 3.8 generation it launched on August 3 as a 2.4 trillion parameter mixture of experts with roughly 95B active per token and a 1M token context window, priced at $2 per million input and $6 per million output on Alibaba Cloud's international region. Weights followed on Hugging Face about ten days later.
The open question for Qwen 4 Max is architecture. Alibaba has said Qwen 4 builds on the design previewed in Qwen 3.8-Flash-Next: very sparse activation, an n-gram embedding layer that stores a large chunk of the parameter count as cheap lookups, and sparse attention. If Max inherits that, the total parameter number could climb well past 3.8-Max's 2.4T while active parameters stay small. That's a guess. Alibaba said the 5 to 10 trillion parameter target applies to Qwen 4.5 and Qwen 5, not Qwen 4, so don't read that number into Max.
For an API buyer, Max is the tier to test first and the tier most likely to arrive first. The 3.8 cadence was flagship API, then flagship weights, then the smaller models.
Qwen 4 Plus
Plus is the tier most people skip in the coverage and most people end up using. It's the hosted middle option: cheaper than Max, stronger than Flash, and it's where Alibaba puts the model that most production traffic actually runs on. Qwen 3.7-Plus listed at $0.40 in and $1.60 out per million on the international region, with a 1M window that steps up to a higher rate above 256K.
Two things worth knowing. Plus has never shipped as open weights, and there's no sign Qwen 4 changes that. And Alibaba didn't ship a distinct 3.8-Plus at all as far as its pricing page shows; the 3.8 generation went Max, Flash, Flash-Next, 27B. So Qwen 4 Plus reappearing on the slide is itself a small piece of news. If you're currently on 3.7-Plus or 3.6-Plus, this is your upgrade path, and it's the tier whose price will matter most to your bill.
Qwen 4 Flash
Flash is the throughput tier: summarization, extraction, classification, routing, agent sub-steps, anything you run millions of times a day where a few points of benchmark score matter less than cost per call.
This is the tier we know the most about, because Alibaba already showed the architecture. Qwen 3.8-Flash-Next, released August 28, is described by Alibaba as an experimental preview of the Qwen 4 design. It runs 125B total parameters with 6B active per token, plus a 51B n-gram embedding table, a 262K native context that extends toward 1M, and multimodal input. Alibaba reported roughly 90% lower training cost than Qwen 3.7-Plus. That's a vendor number; nobody has replicated it.
The safe read is that Qwen 4 Flash is Flash-Next with the experimental label removed and the training run finished. If you want to know what Qwen 4 Flash will feel like, run Flash-Next today. We covered the specs and the license in the Qwen 4 Flash preview post. One caution on that license: Flash-Next ships under qwen-community-1.0, not Apache 2.0, and it's worth reading before you build on it commercially.
Qwen 4 27B
The 27B is the tier Alibaba described as open weights for local deployment, and it's the one drawing the most search traffic by a wide margin. Qwen 3.8-27B was a dense 27B under Apache 2.0 with a 262K context and a vision encoder, and it became the most-run open model of the summer because it fits one 80 GB card at BF16 and one 24 GB card at 4-bit.
The real question for Qwen 4 27B is whether it stays dense or picks up the sparse design. A dense 27B keeps the same hardware math as 3.8. A sparse 27B with an n-gram table changes what "27B" means for VRAM. Neither is confirmed. We're tracking that, the license, and the hardware tiers in the Qwen 4 27B post, which updates when the model card lands.
Which tier to plan for
Pick by the workload, not the headline.
If you're running agents, long-context coding, or anything where quality is the constraint, plan for Max and budget for it. If you're running a product with real traffic and a real bill, Plus is where you'll probably end up, and Max is what you'll benchmark it against. If you're doing high-volume, low-stakes calls, Flash. If you need to run it yourself, on your own hardware, with your own data, 27B is the only tier that's been confirmed as open weights.
The Qwen 3.8 lineup is the practical stand-in for all four until Qwen 4 ships. The Qwen 3.8-Max and Qwen 3.8-27B posts cover the two ends of it, and the Best Chinese LLMs roundup puts them next to DeepSeek, GLM, and Kimi.
How to prepare
Set up an eval on the Qwen 3.8 tier that matches your target. Max for Max, Flash-Next for Flash, 27B for 27B, 3.7-Plus for Plus. When Qwen 4 lands, you rerun the same eval and get a real delta instead of a benchmark table from a vendor slide.
If you're calling Qwen through an OpenAI-compatible endpoint, keep the model string in config so switching tiers is a one-line change. Qwen 3.8-Max and Qwen 3.8-27B are both live on the Yotta AI Gateway today, and new Qwen releases get added as they ship.
For the 27B, size your hardware to the 3.8 tiers (24 GB, 48 GB, 80 GB) and expect to revisit if the architecture changes.
Frequently asked questions
Did Alibaba officially announce the four Qwen 4 tiers?
The tier names were shown on stage at Apsara on September 22 by Qwen lead Liu Dayiheng and reported by multiple outlets. Alibaba's written release only says Qwen 4 is in training. Until a model card exists, treat the names as previewed and the specs as unknown.
When do the Qwen 4 tiers come out?
No date. "Very soon" is the only timing Alibaba has given. In the 3.8 generation, Max launched first, weights followed about ten days later, and the Flash and 27B tiers came within a month.
Will Qwen 4 Max be open source?
Unknown. Qwen 3.8-Max weights were published, so there's precedent. The 27B is the only tier described as open weights.
What's the difference between Qwen 4 Flash and Qwen 4 27B?
Flash is a hosted, high-throughput model built on the sparse architecture previewed in Flash-Next. The 27B is the tier meant for running locally. Flash-Next also shipped as open weights, but at 125B total it's a very different thing to host than a 27B.
Is there a Qwen 4 Coder?
Not announced. Alibaba has released Coder variants in past generations, but nothing was shown at Apsara.
Is Qwen 4 on Yotta?
Not yet. Qwen 3.8-Max and Qwen 3.8-27B are on the AI Gateway now, and Qwen 4 will be added when it's available.
Bottom line
Four names and a "very soon." That's the whole announcement. But the names aren't random: Alibaba has shipped Max, Plus, Flash, and an open 27B before, and the roles held. Test the 3.8 tier that matches your workload now, keep your model string in config, and you'll have a real comparison on launch day instead of a slide.
Track the rest of the wave in the Qwen 4 release date post, and see current Qwen models and rates on the AI Gateway.



