---
title: "Qwen 3.8 Flash-Next Hardware Requirements: Size, GPU, and VRAM (2026)"
slug: qwen-3-8-flash-next-hardware-requirements-gpu-memory-2026
description: "Qwen 3.8 Flash-Next is 360 GB at full precision, 186 GB at FP8, and 135 GB at NVFP4. Every build size, which GPUs each one fits, the verified vLLM and SGLang configs, and what the GGUF builds need."
author: "Yotta Labs"
date: 2026-10-11
categories: ["Inference"]
canonical: https://www.yottalabs.ai/post/qwen-3-8-flash-next-hardware-requirements-gpu-memory-2026
---

# Qwen 3.8 Flash-Next Hardware Requirements: Size, GPU, and VRAM (2026)

![](https://cdn.sanity.io/images/wy75wyma/production/84d37cd30e3dc846b9290a7716915dacaee092ed-1200x627.png)

*Flash-Next activates 6 billion parameters per token and stores 180 billion. The second number is the one your GPUs have to hold. Here's every build size and the hardware each one fits.*

Qwen 3.8 Flash-Next's main checkpoint was downloaded about 1.8 million times in the last month on Hugging Face, and it's an easy model to size wrong. "Flash" in Alibaba's lineup means fast and cheap per token on the API. It doesn't mean small on disk.

When we wrote up the [Flash-Next specs and benchmarks](https://www.yottalabs.ai/post/qwen-3-8-flash-next-specs-qwen-4-preview-2026) at release, the tested hardware configurations hadn't settled. They have now. vLLM and SGLang both publish verified recipes, Alibaba ships an official FP8 checkpoint, and there are two NVFP4 builds and a few hundred community quantizations. This post is the full breakdown.

## The numbers that matter

<!-- unsupported block: table -->

Two things in that table explain the rest. The first is the usual MoE rule: 6B active parameters make each token cheap to compute, but all 125B backbone parameters have to sit in memory, because the router can pick any expert on the next token. Low active parameters save compute, not VRAM.

The second is specific to this model. Fifty-one billion of the stored parameters are an n-gram embedding table, a large lookup structure that adds capacity with very little compute per token. It's the reason the model is bigger than its 125B headline, and it's also the piece you can move off the GPU, which is what makes the smaller configurations below possible.

## The four official and near-official builds

<!-- unsupported block: table -->

Alibaba describes the FP8 checkpoint as fine-grained FP8 quantization with performance nearly identical to the original. The RadixArk NVFP4 build quantizes only the routed experts to 4-bit and keeps attention, routers, embeddings, and the vision encoder at BF16, which is how it lands at 135 GB instead of the roughly 90 GB a flat 4-bit encoding would give. Its model card reports results close to BF16 on the tests it ran and notes that long agentic generations tend to run longer.

## Which GPUs fit which build

Status reflects the published recipes as of mid-October 2026. "Verified" means the recipe's maintainers ran that configuration.

<!-- unsupported block: table -->

The practical reading, by what you have:

**One Blackwell GPU.** The NVFP4 build on a single B200 or B300 is the cheapest verified way to serve Flash-Next with a production engine. NVFP4 is a Blackwell format, so this path doesn't exist on H100 or H200.

**A Hopper node.** Four H100s is the verified floor, and it only works because vLLM moves the n-gram table to CPU memory. Eight H200s is the comfortable config, verified in both engines, with room for the full 262K context and real concurrency.

**One workstation card.** A 96 GB RTX PRO 6000 holds a 2-bit or 3-bit GGUF. That gets the model running. It won't give you a long context or much concurrency, and low-bit builds cost quality.

**A consumer card.** No. A 24 GB or 32 GB GPU can't hold any build. If one GPU is your constraint, the dense [Qwen 3.8 27B](https://www.yottalabs.ai/post/qwen-3-8-27b-specs-hardware-requirements-how-to-run-2026) is the Qwen to run, and [Flash-Next vs the 27B](https://www.yottalabs.ai/post/qwen-3-8-flash-next-vs-qwen-3-8-27b-2026) covers what you give up.

## The n-gram table trick

The 51B embedding table is why a 186 GB FP8 checkpoint can run on four 80 GB cards. Because it's a lookup table and not a set of layers every token passes through, an engine can keep it in system RAM and fetch rows as needed.

In vLLM this is one environment variable, VLLM_PLE_CPU_OFFLOAD=1, and the recipe sets it automatically on H100. Without it, a plain four-way tensor-parallel launch on 80 GB cards runs out of memory at startup. With it, plan for at least 51 GB of free host RAM plus headroom. The offload currently works on NVIDIA GPUs only.

vLLM's recipe reports about 1,430 output tokens per second on that 4x H100 config at 64 concurrent requests, on a synthetic 1,024-in, 256-out workload. That's the recipe's own number on its own test, so treat it as a ceiling to check against, not a promise.

SGLang has said it is adding cookbook recipes that take the same idea further, including a single RTX PRO 6000 with the table in system RAM. Those weren't on the cookbook page when we checked, so confirm there before you plan around them.

## Engines and launch commands

Both major engines support Flash-Next, and Alibaba's model card lists TokenSpeed as well. The [vLLM vs SGLang comparison](https://www.yottalabs.ai/post/vllm-vs-sglang-which-inference-engine-should-you-use-in-2026) covers the general choice. For this model the difference is mostly which hardware each one has verified.

**SGLang** has verified configs for 8x H200 with the FP8 checkpoint, and for B200, B300, and GB300 with NVFP4 at one GPU per replica. The model card's basic launch is:

```bash
python3 -m sglang.launch_server --model-path "Qwen/Qwen3.8-Flash-Next-FP8" --host 0.0.0.0 --port 30000
```

The cookbook's verified configs add hardware-specific flags and a pinned Docker image, so copy the launch arguments from the cookbook page for your GPU.

**vLLM** needs version 0.29.0 or later and, per the recipe, its tagged Docker image; a plain PyPI install isn't supported for this model. The shape of the 8x H200 launch:

```bash
vllm serve Qwen/Qwen3.8-Flash-Next-FP8 --tensor-parallel-size 8 --enable-expert-parallel --moe-backend triton --gpu-memory-utilization 0.85 --max-num-seqs 256 --kv-cache-dtype fp8
```

Three things from the recipe worth knowing before you start. On eight GPUs the FP8 checkpoint needs expert parallelism turned on; plain eight-way tensor parallelism isn't compatible with how the checkpoint is quantized. The recipe says to keep --max-num-seqs at 256 to avoid a cache-capacity error at startup. And the model has a built-in multi-token prediction head for speculative decoding, which the full recipe enables with three draft tokens. Copy the complete command, including the speculative decoding config, from the recipe for your engine version.

On context: both engines serve the native 262,144 tokens. The 1M window needs the YaRN configuration from the model card, and vLLM's recipe advises checking quality on shorter inputs first when you turn it on.

## GGUF, Ollama, and Macs

The community builds are where most single-machine users land. Sizes from a widely used GGUF repository:

<!-- unsupported block: table -->

Notice how little the low-bit builds shrink. Going from 4-bit to 1-bit only takes the file from about 120 GB to 70 GB. The n-gram table is the reason: lookup layers don't tolerate aggressive quantization the way dense layers do, so builders keep them at higher precision.

The file size is the floor, not the requirement. Context and cache sit on top of it, and the builder's own rule of thumb is to pick a file a gigabyte or two smaller than your VRAM for full GPU offload. You need a recent llama.cpp release for the architecture.

Ollama has a Flash-Next library page, including an MLX build of about 105 GB for Apple silicon. That needs a machine with more unified memory than the file size, which rules out most Macs.

## The route most teams should take

Three things to weigh before you rent a node.

Flash-Next is a preview. Alibaba describes it as an experimental look at the architecture behind Qwen 4, and the [Qwen 4 release tracker](https://www.yottalabs.ai/post/qwen-4-release-date-what-is-known-how-to-prepare-2026) follows what that means. Building a production system on a preview model is a choice to make on purpose.

The license is qwen-community-1.0, not Apache 2.0. Read it before a commercial deployment. The [specs post](https://www.yottalabs.ai/post/qwen-3-8-flash-next-specs-qwen-4-preview-2026) covers what's different about it.

And it's a node-scale model next to others in its class. [GLM 5.3 Flash](https://www.yottalabs.ai/post/glm-5-3-flash-hardware-requirements-gpu-memory-2026) and [DeepSeek V4.1 Flash](https://www.yottalabs.ai/post/deepseek-v4-1-flash-hardware-requirements-gpu-memory-2026) have their own hardware breakdowns, and the [three-way Flash comparison](https://www.yottalabs.ai/post/deepseek-v4-flash-vs-glm-5-3-flash-vs-qwen-flash-next-2026) lines them up on price and benchmarks.

If you're self-hosting, [8-GPU H200 and B300 nodes by the hour](https://www.yottalabs.ai/pricing) cover the verified Hopper and Blackwell configs in the table above. If you want production Qwen behind an API today, Qwen 3.8-Max and Qwen3.8-27B are on [Yotta AI Gateway](https://www.yottalabs.ai/ai-gateway) behind one OpenAI-compatible key.

## Frequently asked questions

**How big is Qwen 3.8 Flash-Next?** 180B stored parameters: a 125B backbone with 6B active per token, a 51B n-gram embedding table, and a 4B prediction head. On disk that's about 360 GB at BF16, 186 GB at FP8, and 135 GB at NVFP4.

**How much VRAM does Qwen 3.8 Flash-Next need?** It depends on the build. The NVFP4 build runs on one 192 GB B200. The FP8 build is verified on four 80 GB H100s with the n-gram table in host RAM, and on eight H200s without that. Community 4-bit GGUF builds are about 100 to 120 GB.

**Can I run Qwen 3.8 Flash-Next on a single GPU?** Yes, on a Blackwell GPU with the NVFP4 build, which SGLang has verified on one B200 and one B300. On a 96 GB RTX PRO 6000 only 2-bit and 3-bit GGUF builds fit. No 24 GB or 32 GB card can hold any build.

**Can Qwen 3.8 Flash-Next run on 4x H100?** Yes, with the FP8 checkpoint and the n-gram table offloaded to CPU memory. vLLM's recipe verifies that config and sets the offload automatically. Plan for at least 51 GB of free host RAM.

**Is there a GGUF or Ollama version of Qwen 3.8 Flash-Next?** Yes. The model card links several hundred community quantizations, GGUF files run from about 70 GB to 354 GB, and Ollama has a library page that includes an MLX build of about 105 GB.

**Why is Flash-Next so large if only 6B parameters are active?** Active parameters set the compute per token. Memory is set by everything stored: all 125B backbone parameters, because any expert can be picked next, plus the 51B n-gram table.

**What is the n-gram embedding table?** A 51B-parameter lookup table that adds capacity with very little compute per token. Because it's a lookup, engines can keep it in system RAM, which is what lets the model run on smaller GPU configurations.

**What's the difference between the FP8 and NVFP4 builds?** FP8 is Alibaba's official 8-bit checkpoint at about 186 GB and runs on Hopper and Blackwell GPUs. NVFP4 quantizes the routed experts to 4-bit for about 135 GB and runs on Blackwell GPUs.

**Which engines support Qwen 3.8 Flash-Next?** vLLM 0.29.0 and later, SGLang, and TokenSpeed, plus llama.cpp, Ollama, and LM Studio for the community builds.

## Bottom line

Qwen 3.8 Flash-Next needs between 135 GB and 360 GB of memory for its weights in the builds made for serving, and the config depends on the build: one B200 or B300 for NVFP4, four H100s or eight H200s for FP8, and a multi-GPU node for BF16. A workstation card can load a low-bit GGUF, and a consumer card can't load anything.

The n-gram table is what makes this model unusual, both for its size and for the ways engines are finding to run it on less. If you're renting hardware to serve it, [an 8-GPU H200 or B300 node](https://www.yottalabs.ai/pricing) covers the verified configs. If one GPU is the limit, start with the [27B](https://www.yottalabs.ai/post/qwen-3-8-27b-specs-hardware-requirements-how-to-run-2026).
