Oct 11, 2026
Qwen 3.8 Flash-Next Hardware Requirements: Size, GPU, and VRAM (2026)
GPU Pods
Distributed Inference
Qwen 3.8 Flash-Next is 360 GB at full precision, 186 GB at FP8, and 135 GB at NVFP4. Every build size, which GPUs each one fits, the verified vLLM and SGLang configs, and what the GGUF builds need.

Flash-Next activates 6 billion parameters per token and stores 180 billion. The second number is the one your GPUs have to hold. Here's every build size and the hardware each one fits.
Qwen 3.8 Flash-Next's main checkpoint was downloaded about 1.8 million times in the last month on Hugging Face, and it's an easy model to size wrong. "Flash" in Alibaba's lineup means fast and cheap per token on the API. It doesn't mean small on disk.
When we wrote up the Flash-Next specs and benchmarks at release, the tested hardware configurations hadn't settled. They have now. vLLM and SGLang both publish verified recipes, Alibaba ships an official FP8 checkpoint, and there are two NVFP4 builds and a few hundred community quantizations. This post is the full breakdown.
The numbers that matter
| Spec | Qwen 3.8 Flash-Next |
| Backbone parameters | 125B total, 6B active per token (MoE, 512 experts, 10 routed plus 1 shared) |
| N-gram embedding table | 51B parameters |
| Multi-token prediction head | 4B parameters |
| Total stored | 180B, per the Hugging Face model card |
| Full precision checkpoint (BF16) | About 360 GB (335 GiB) |
| Official FP8 checkpoint | About 186 GB (172.78 GiB) |
| NVFP4 checkpoint | About 135 GB |
| Smallest community GGUF | About 70 GB (1-bit) |
| Context window | 262,144 tokens native, extensible to 1M with YaRN |
| Inputs | Text, image, and video |
| License | qwen-community-1.0 |
Two things in that table explain the rest. The first is the usual MoE rule: 6B active parameters make each token cheap to compute, but all 125B backbone parameters have to sit in memory, because the router can pick any expert on the next token. Low active parameters save compute, not VRAM.
The second is specific to this model. Fifty-one billion of the stored parameters are an n-gram embedding table, a large lookup structure that adds capacity with very little compute per token. It's the reason the model is bigger than its 125B headline, and it's also the piece you can move off the GPU, which is what makes the smaller configurations below possible.
The four official and near-official builds
| Build | Size | Who publishes it | What it's for |
| BF16 | About 360 GB | Alibaba (Qwen/Qwen3.8-Flash-Next) | Full precision, fine-tuning, making your own quantizations |
| FP8 | About 186 GB | Alibaba (Qwen/Qwen3.8-Flash-Next-FP8) | The default for serving on H100 and H200 nodes |
| NVFP4 | About 135 GB | RadixArk and NVIDIA | Blackwell GPUs, including single-GPU serving |
| GGUF | 70 GB to 354 GB | Community (llama.cpp, Ollama, LM Studio) | Workstations and unified-memory machines |
Alibaba describes the FP8 checkpoint as fine-grained FP8 quantization with performance nearly identical to the original. The RadixArk NVFP4 build quantizes only the routed experts to 4-bit and keeps attention, routers, embeddings, and the vision encoder at BF16, which is how it lands at 135 GB instead of the roughly 90 GB a flat 4-bit encoding would give. Its model card reports results close to BF16 on the tests it ran and notes that long agentic generations tend to run longer.
Which GPUs fit which build
Status reflects the published recipes as of mid-October 2026. "Verified" means the recipe's maintainers ran that configuration.
| Config | Total GPU memory | Build that fits | Status |
| 1x RTX 5090 | 32 GB | None | The smallest build is 70 GB |
| 1x H100 | 80 GB | 1-bit and 2-bit GGUF (70 to 78 GB) | Fits on paper with almost no room for context |
| 1x RTX PRO 6000 | 96 GB | 2-bit and 3-bit GGUF (78 to 93 GB) | Community builds on llama.cpp |
| 1x H200 | 141 GB | 4-bit GGUF (Q4_K_M, 119.6 GB) | Community build on llama.cpp |
| 1x B200 | 192 GB | NVFP4 (135 GB) | Verified in the SGLang cookbook, one GPU per replica |
| 1x B300 | About 275 GB | NVFP4 (135 GB) | Verified in the SGLang cookbook, one GPU per replica |
| 4x H100 | 320 GB | FP8, with the n-gram table in host RAM | Verified in the vLLM recipe |
| 4x H200 | 564 GB | FP8 | Used in vLLM's KV cache offload validation |
| 8x H200 | 1,128 GB | FP8 | Verified in both vLLM and SGLang |
| 2x GB300 | About 550 GB | FP8 or BF16 | vLLM's validated minimum |
| 4x GB300 | About 1.1 TB | FP8 or NVFP4 | vLLM's recommended config, also in the SGLang cookbook |
The practical reading, by what you have:
One Blackwell GPU. The NVFP4 build on a single B200 or B300 is the cheapest verified way to serve Flash-Next with a production engine. NVFP4 is a Blackwell format, so this path doesn't exist on H100 or H200.
A Hopper node. Four H100s is the verified floor, and it only works because vLLM moves the n-gram table to CPU memory. Eight H200s is the comfortable config, verified in both engines, with room for the full 262K context and real concurrency.
One workstation card. A 96 GB RTX PRO 6000 holds a 2-bit or 3-bit GGUF. That gets the model running. It won't give you a long context or much concurrency, and low-bit builds cost quality.
A consumer card. No. A 24 GB or 32 GB GPU can't hold any build. If one GPU is your constraint, the dense Qwen 3.8 27B is the Qwen to run, and Flash-Next vs the 27B covers what you give up.
The n-gram table trick
The 51B embedding table is why a 186 GB FP8 checkpoint can run on four 80 GB cards. Because it's a lookup table and not a set of layers every token passes through, an engine can keep it in system RAM and fetch rows as needed.
In vLLM this is one environment variable, VLLM_PLE_CPU_OFFLOAD=1, and the recipe sets it automatically on H100. Without it, a plain four-way tensor-parallel launch on 80 GB cards runs out of memory at startup. With it, plan for at least 51 GB of free host RAM plus headroom. The offload currently works on NVIDIA GPUs only.
vLLM's recipe reports about 1,430 output tokens per second on that 4x H100 config at 64 concurrent requests, on a synthetic 1,024-in, 256-out workload. That's the recipe's own number on its own test, so treat it as a ceiling to check against, not a promise.
SGLang has said it is adding cookbook recipes that take the same idea further, including a single RTX PRO 6000 with the table in system RAM. Those weren't on the cookbook page when we checked, so confirm there before you plan around them.
Engines and launch commands
Both major engines support Flash-Next, and Alibaba's model card lists TokenSpeed as well. The vLLM vs SGLang comparison covers the general choice. For this model the difference is mostly which hardware each one has verified.
SGLang has verified configs for 8x H200 with the FP8 checkpoint, and for B200, B300, and GB300 with NVFP4 at one GPU per replica. The model card's basic launch is:
python3 -m sglang.launch_server --model-path "Qwen/Qwen3.8-Flash-Next-FP8" --host 0.0.0.0 --port 30000The cookbook's verified configs add hardware-specific flags and a pinned Docker image, so copy the launch arguments from the cookbook page for your GPU.
vLLM needs version 0.29.0 or later and, per the recipe, its tagged Docker image; a plain PyPI install isn't supported for this model. The shape of the 8x H200 launch:
vllm serve Qwen/Qwen3.8-Flash-Next-FP8 --tensor-parallel-size 8 --enable-expert-parallel --moe-backend triton --gpu-memory-utilization 0.85 --max-num-seqs 256 --kv-cache-dtype fp8Three things from the recipe worth knowing before you start. On eight GPUs the FP8 checkpoint needs expert parallelism turned on; plain eight-way tensor parallelism isn't compatible with how the checkpoint is quantized. The recipe says to keep --max-num-seqs at 256 to avoid a cache-capacity error at startup. And the model has a built-in multi-token prediction head for speculative decoding, which the full recipe enables with three draft tokens. Copy the complete command, including the speculative decoding config, from the recipe for your engine version.
On context: both engines serve the native 262,144 tokens. The 1M window needs the YaRN configuration from the model card, and vLLM's recipe advises checking quality on shorter inputs first when you turn it on.
GGUF, Ollama, and Macs
The community builds are where most single-machine users land. Sizes from a widely used GGUF repository:
| Quantization | File size | Smallest single GPU that holds it |
| IQ1_S (1-bit) | 70.1 GB | 80 GB H100 |
| Q2_K (2-bit) | 80.9 GB | 96 GB RTX PRO 6000 |
| Q3_K_M (3-bit) | 92.0 GB | 96 GB RTX PRO 6000, with nothing to spare |
| IQ4_XS (4-bit) | 97.7 GB | 141 GB H200 |
| Q4_K_M (4-bit, the usual default) | 119.6 GB | 141 GB H200 |
| Q5_K_M (5-bit) | 134.7 GB | 141 GB H200, with little to spare |
| Q8_0 (8-bit) | 188.3 GB | 192 GB B200, with little to spare |
| BF16 | 354.0 GB | None. Needs a multi-GPU node |
Notice how little the low-bit builds shrink. Going from 4-bit to 1-bit only takes the file from about 120 GB to 70 GB. The n-gram table is the reason: lookup layers don't tolerate aggressive quantization the way dense layers do, so builders keep them at higher precision.
The file size is the floor, not the requirement. Context and cache sit on top of it, and the builder's own rule of thumb is to pick a file a gigabyte or two smaller than your VRAM for full GPU offload. You need a recent llama.cpp release for the architecture.
Ollama has a Flash-Next library page, including an MLX build of about 105 GB for Apple silicon. That needs a machine with more unified memory than the file size, which rules out most Macs.
The route most teams should take
Three things to weigh before you rent a node.
Flash-Next is a preview. Alibaba describes it as an experimental look at the architecture behind Qwen 4, and the Qwen 4 release tracker follows what that means. Building a production system on a preview model is a choice to make on purpose.
The license is qwen-community-1.0, not Apache 2.0. Read it before a commercial deployment. The specs post covers what's different about it.
And it's a node-scale model next to others in its class. GLM 5.3 Flash and DeepSeek V4.1 Flash have their own hardware breakdowns, and the three-way Flash comparison lines them up on price and benchmarks.
If you're self-hosting, 8-GPU H200 and B300 nodes by the hour cover the verified Hopper and Blackwell configs in the table above. If you want production Qwen behind an API today, Qwen 3.8-Max and Qwen3.8-27B are on Yotta AI Gateway behind one OpenAI-compatible key.
Frequently asked questions
How big is Qwen 3.8 Flash-Next? 180B stored parameters: a 125B backbone with 6B active per token, a 51B n-gram embedding table, and a 4B prediction head. On disk that's about 360 GB at BF16, 186 GB at FP8, and 135 GB at NVFP4.
How much VRAM does Qwen 3.8 Flash-Next need? It depends on the build. The NVFP4 build runs on one 192 GB B200. The FP8 build is verified on four 80 GB H100s with the n-gram table in host RAM, and on eight H200s without that. Community 4-bit GGUF builds are about 100 to 120 GB.
Can I run Qwen 3.8 Flash-Next on a single GPU? Yes, on a Blackwell GPU with the NVFP4 build, which SGLang has verified on one B200 and one B300. On a 96 GB RTX PRO 6000 only 2-bit and 3-bit GGUF builds fit. No 24 GB or 32 GB card can hold any build.
Can Qwen 3.8 Flash-Next run on 4x H100? Yes, with the FP8 checkpoint and the n-gram table offloaded to CPU memory. vLLM's recipe verifies that config and sets the offload automatically. Plan for at least 51 GB of free host RAM.
Is there a GGUF or Ollama version of Qwen 3.8 Flash-Next? Yes. The model card links several hundred community quantizations, GGUF files run from about 70 GB to 354 GB, and Ollama has a library page that includes an MLX build of about 105 GB.
Why is Flash-Next so large if only 6B parameters are active? Active parameters set the compute per token. Memory is set by everything stored: all 125B backbone parameters, because any expert can be picked next, plus the 51B n-gram table.
What is the n-gram embedding table? A 51B-parameter lookup table that adds capacity with very little compute per token. Because it's a lookup, engines can keep it in system RAM, which is what lets the model run on smaller GPU configurations.
What's the difference between the FP8 and NVFP4 builds? FP8 is Alibaba's official 8-bit checkpoint at about 186 GB and runs on Hopper and Blackwell GPUs. NVFP4 quantizes the routed experts to 4-bit for about 135 GB and runs on Blackwell GPUs.
Which engines support Qwen 3.8 Flash-Next? vLLM 0.29.0 and later, SGLang, and TokenSpeed, plus llama.cpp, Ollama, and LM Studio for the community builds.
Bottom line
Qwen 3.8 Flash-Next needs between 135 GB and 360 GB of memory for its weights in the builds made for serving, and the config depends on the build: one B200 or B300 for NVFP4, four H100s or eight H200s for FP8, and a multi-GPU node for BF16. A workstation card can load a low-bit GGUF, and a consumer card can't load anything.
The n-gram table is what makes this model unusual, both for its size and for the ways engines are finding to run it on less. If you're renting hardware to serve it, an 8-GPU H200 or B300 node covers the verified configs. If one GPU is the limit, start with the 27B.



