Aug 30, 2026
DeepSeek V4 Flash vs GLM 5.3 Flash vs Qwen 3.8-Flash-Next (2026)
Cost Optimization
Distributed Inference
Three open Flash-class models shipped in one month. Specs, prices, hardware floors, and licenses compared, with the honest gaps flagged.

August 2026 gave us three open Flash-class models in thirty days. Here's the three-way comparison nobody's benchmarks make easy.
Something changed this month. DeepSeek, Z.ai, and Alibaba each shipped the same play within weeks of each other: a lean-activation MoE, priced far below their own flagship, released with downloadable weights. DeepSeek V4 Flash came July 31, GLM 5.3 Flash on August 26, and Qwen 3.8-Flash-Next on August 28. Three labs, one bet: that most production traffic doesn't need a flagship, and whoever owns the cheap tier owns the volume.
If you're choosing between them, the specs are public, but the comparison is messier than the launch posts suggest. Here's the honest version.
The three side by side
| DeepSeek V4 Flash | GLM 5.3 Flash | Qwen 3.8-Flash-Next | |
| Released | July 31, 2026 | August 26, 2026 | August 28, 2026 |
| Total parameters | 304B | 320B | 125B + 51B n-gram |
| Active per token | ~13B | 18B | 6B |
| Modalities | Text only | Text, image, video | Text, image, video |
| Context | 1M | 1M advertised, 300K evaluated | 262K native, 1M extensible |
| License | MIT | MIT | qwen-community-1.0 |
| Weights on disk | 166.9 GB | ~306 GiB FP8 | Not yet firmly documented |
| Hardware floor | 2x H200 | 8-GPU node | Multi-GPU, configs still settling |
| API price (in/out per M) | $0.44/$1.32 peak, half off-peak | $0.15/$0.50 | Via Alibaba, standard rates |
Three details in that table decide most real choices, so let's take them in turn.
The benchmark comparison you can't actually make
Here's the uncomfortable part: there is no benchmark that all three models were measured on. DeepSeek reported its own suite, GLM reported Terminal-Bench 2.1 and DeepSWE, Qwen reported SWE-bench Pro and CoWorkBench. Every number is vendor-published, and the tests barely overlap, so any article ranking these three on "performance" is comparing announcements, not models.
What can be said with evidence: DeepSeek V4 Flash has been out a month and has the most independent scoring, including Artificial Analysis coverage, and it held up well enough that it embarrassed DeepSeek's own Pro preview. GLM 5.3 Flash's headline is matching its flagship on agent tasks at a tenth of the price, per Z.ai's own tables. Flash-Next's numbers are strong for a 6B-active model and completely unreplicated, and it's explicitly labeled experimental. If benchmark certainty matters to your decision, DeepSeek V4 Flash is the only one with a real track record today, and the honest move with the other two is testing on your own workload.
Hardware: three very different commitments
The self-hosting floors differ more than the parameter counts suggest. DeepSeek V4 Flash is the lightest lift because its checkpoint ships pre-compressed to 166.9 GB, which fits a 2x H200 pod. GLM 5.3 Flash needs roughly double the memory, with no two-GPU entry point at all, an 8-GPU node is the floor. Flash-Next holds roughly 176B parameters including its n-gram table, multi-GPU territory for certain, and tested configurations are still settling since it's four days old.
If self-hosting economics drive the decision, that ordering is the decision: DeepSeek V4 Flash is the only one of the three with a modest entry point.
Licenses: the detail that bites later
DeepSeek and Z.ai both shipped plain MIT, use the weights commercially, fine-tune them, no meaningful strings. Alibaba shipped qwen-community-1.0, a community license with conditions that permissive licenses don't carry. For a hobby deployment that difference is academic. For a product built on the weights, it's the kind of detail legal asks about after you've already committed, so read it first. It's the clearest structural difference between Flash-Next and the other two.
Which one to use
Choose DeepSeek V4 Flash if you want the proven option: a month of independent scoring, the cheapest hardware floor, and a mature serving stack. Its gap is vision, it's the only text-only model of the three.
Choose GLM 5.3 Flash if you need multimodal input with open weights today, or you're betting on the long-horizon agent story its tables claim. Budget for the 8-GPU floor or use the API, whose $0.15/$0.50 pricing is the most aggressive of the three.
Choose Flash-Next if you're evaluating where Qwen 4 is going, that's what it's for, Alibaba says so. It's an architecture preview, not a production commitment, and the license plus the experimental label say treat it that way.
And if the real answer is "run them side by side and measure," which it usually is: all three speak OpenAI-compatible interfaces, so a bake-off is routing configuration. DeepSeek V4 Flash is live on Yotta AI Gateway with flat pricing, alongside the GLM line, the Qwen 3.8 line, Kimi K3, and the rest of the catalog, one API key for the whole experiment.
Frequently asked questions
What is a Flash-class model?
The pattern all three follow: a sparse MoE with a small active-parameter count, priced well below the vendor's flagship, positioned for high-volume production traffic rather than the hardest tasks. This month all three major Chinese labs shipped one with open weights.
Which Flash model is best?
There's no shared benchmark across the three, so nobody honestly knows on capability. DeepSeek V4 Flash has the most independent evidence, GLM 5.3 Flash the strongest vendor claims and multimodality, Flash-Next the most interesting architecture and the most caveats.
Which is cheapest to run?
By API, GLM 5.3 Flash at $0.15 in / $0.50 out. By self-hosting, DeepSeek V4 Flash, its 166.9 GB checkpoint is the only one that fits a 2-GPU pod.
Are they all open source?
All three have downloadable weights, but licenses differ: DeepSeek V4 Flash and GLM 5.3 Flash are MIT, Flash-Next is qwen-community-1.0, which carries conditions worth reading before commercial use.
Which ones handle images and video?
GLM 5.3 Flash and Qwen 3.8-Flash-Next both take image and video input. DeepSeek V4 Flash is text-only.
Can I run any of them on a single GPU?
No. The smallest checkpoint of the three is 166.9 GB, and the MoE rule applies to all of them: every parameter needs memory even though few fire per token. Two datacenter GPUs is the absolute floor, and only DeepSeek V4 Flash reaches it.
Which are on Yotta?
DeepSeek V4 Flash and V4 Pro are live on Yotta AI Gateway, alongside the GLM line including the flagship GLM 5.3, and the Qwen 3.8 line. This post will note it as the newer Flash models join the catalog.
Bottom line
The Flash class is August's real story: three labs concluded independently that open, lean, cheap wins the volume tier, and shipped within a month of each other. The scoreboard is unsettled because the benchmarks don't overlap, but the structural differences are settled enough to choose on: DeepSeek for evidence and cheap hosting, GLM for open multimodality and API price, Flash-Next for a look at Qwen 4.
The full breakdowns behind this comparison: DeepSeek V4 Flash hardware requirements, GLM 5.3 Flash hardware requirements, and the Qwen 3.8-Flash-Next preview. And when you're ready to test any of them against your own workload, multi-GPU capacity by the hour and the Gateway catalog are one console away.



