Oct 07, 2026
Mistral Large 4 Hardware Requirements: GPU, Memory, and What to Plan For (2026)
GPU Pods
Distributed Inference
Mistral Large 4 is a 1.05 trillion-parameter open-weight model with weights due by the end of October. The memory math at each precision, which GPU nodes fit it, and what to have ready.

Mistral Large 4 is in public preview and the weights land by the end of October. Here's the memory math you can do today, and the hardware to line up before the download exists.
Mistral released Mistral Large 4 on October 6, 2026 as a public preview on its API, with a promise that matters to anyone who runs their own models: the weights are coming by the end of the month. Mistral's model docs list it at 1.05 trillion total parameters, which puts it in the same weight class as the biggest open-weight releases of the year.
The checkpoint isn't public yet, so nobody can give you a verified GPU count. What you can do is the arithmetic. Parameter count times bits per parameter gives you the size of the weights, and the size of the weights decides the hardware. This post walks through that math, maps it to real GPU nodes, and flags what is still unknown. It will be updated on this URL when the weights ship.
The numbers that matter
| Spec | Mistral Large 4 |
| Released | October 6, 2026 (public preview, API) |
| Total parameters | 1.05T per Mistral's model docs (the announcement rounds to 1 trillion) |
| Active parameters | 52B per the model docs, 49B per the announcement |
| Architecture | Granular Mixture-of-Experts, with a 1.6B vision encoder |
| Modalities | Text and image input |
| Context window | 1M tokens |
| Languages | 160+, including every official EU language |
| API model ID | mistral-large-4 |
| Open weights | Due by the end of October 2026 |
| License | Not announced |
| Checkpoint size and precision | Not published |
Two notes on that table. Mistral's own pages disagree slightly on the active parameter count, 49 billion in the announcement and 52B in the docs, so treat it as about 50B until the model card settles it. And the last two rows are the ones this whole post hinges on. Until Mistral says what precision the checkpoint ships in, every hardware figure below is an estimate from arithmetic, and it's labeled that way.
Why 50B active doesn't make it small
Mistral Large 4 is a Mixture-of-Experts model, which means only about 50B of its 1.05T parameters fire for any given token. That keeps the compute cost per token down. It doesn't help you with memory. Every expert has to sit in GPU memory and be ready, because the router picks a different set for each token. Low active parameters save compute, not VRAM.
The same rule shaped every large open release this year. Kimi K3 has 104B active parameters and still needs 1.56 TB for its weights. DeepSeek V4.1 Flash runs 8B to 16B active and still needs a full node. Mistral Large 4 will follow the same pattern.
The memory math, by precision
Here is what 1.05 trillion parameters weigh at each precision a lab might ship. The "plan for" column adds about 20% on top of the weights for activations, runtime overhead, and cache, which is the margin the vLLM recipe for DeepSeek V4.1 Flash used (510 GB of weights, about 614 GB to serve).
| Precision | Bits per parameter | Weights (estimate) | GPU memory to plan for |
| BF16, full precision | 16 | About 2.1 TB | About 2.5 TB |
| FP8 | 8 | About 1.05 TB | About 1.26 TB |
| Mixed FP8 and FP4 | About 4.8 | About 630 GB | About 760 GB |
| 4-bit (NVFP4 style) | 4 | About 525 GB | About 630 GB |
There's a reason to take the 4-bit row seriously. When Mistral released Mistral Large 3 in December 2025, a 675B model with 41B active parameters, it published an NVFP4 checkpoint alongside the full weights and said the model could run on a single 8x A100 or 8x H100 node with vLLM. If Mistral does the same for Large 4, the practical download is closer to 525 GB than 2.1 TB. That's a pattern from the last release, not a commitment for this one.
Which GPU configurations fit
Mapping those estimates onto the nodes people rent. "Fits" means the planning figure from the table above fits in aggregate GPU memory. None of these are tested configurations, because there is nothing to test yet.
| Config | Total GPU memory | 4-bit (about 630 GB) | FP8 (about 1.26 TB) | BF16 (about 2.5 TB) |
| 8x H100 80 GB | 640 GB | At the floor, no margin | No | No |
| 4x B200 | 768 GB | Fits | No | No |
| 8x H200 | 1,128 GB | Fits, with room for long context | No, the weights alone are about 1.05 TB | No |
| 4x B300 or GB300 tray | About 1.1 TB | Fits, with room for long context | No | No |
| 8x B200 | About 1.5 TB | Fits | Fits | No |
| 8x B300 | About 2.2 TB | Fits | Fits | No, the weights alone are about 2.1 TB |
| Multi-node cluster | 2.5 TB and up | Fits | Fits | Fits |
The practical reading: if Mistral ships a 4-bit checkpoint, this is a single-node model, and an 8x H200 node or a 4-GPU Blackwell tray is the config to price first. An 8x H100 node only clears it on paper. If the only checkpoint is FP8, you need an 8-GPU Blackwell node. If you want full BF16 weights, for fine-tuning or for making your own quantization, that's multi-node territory from the start.
There is no single-GPU or workstation path at any of these sizes. A 32 GB RTX 5090 holds about 6% of the smallest estimate.
What's still unknown
Four things decide the final answer, and Mistral hasn't published any of them yet.
Checkpoint precision. Covered above. It moves the requirement by 4x.
The license. Mistral Large 3 shipped under Apache 2.0. Mistral Large 2 before it shipped under a research license that restricted commercial use. Mistral calls Large 4 open-weight and talks about self-deployment on private cloud and on-premise, but the license text isn't out. Don't plan a commercial deployment around it until it is.
Serving engines. Mistral Large 3 was supported in vLLM, SGLang, and TensorRT-LLM. Nothing has been confirmed for Large 4. Expect the major engines to have support near release, and check the release notes before you build an image. Our vLLM vs SGLang comparison covers how to choose between the two for a model this size.
KV cache cost. The context window is 1M tokens. How much memory a long context costs depends on the architecture, and it varies a lot between models: it was a real cost on Kimi K3 and nearly free on DeepSeek V4.1 Flash. Until there are numbers, leave headroom.
What to have ready before the weights land
Disk. Plan for at least twice the checkpoint size on fast local storage. For a 4-bit checkpoint that's over 1 TB, and for BF16 it's over 4 TB.
A node you can get quickly. Know which config you'd rent, and where, before release day. On Yotta, 8-GPU H200 and B300 nodes by the hour cover every single-node row in the table above.
An evaluation set. You can build this today. The preview API serves the same model, so run your own prompts through it now and you'll know whether the weights are worth a node before they exist.
A baseline. Mistral reports 61.7% on DeepSWE v1.1 for Large 4. On the same benchmark, each vendor's own run puts GLM 5.3 at 66.9 and DeepSeek V4 Pro at 62.7. Those are different labs' harnesses and not a ranking, but they tell you which models to test side by side. The best open-source LLMs roundup covers the rest of the field.
The route that works today
Until the weights ship, the API is the only way to use Mistral Large 4. Mistral's announcement lists it at $1.36 per million input tokens and $4.18 per million output. Mistral's docs page currently shows $0.68 and $2.09 with the list prices struck through, so check the page for what's in effect when you read this.
At either price, self-hosting a trillion-parameter model has to earn its place. An 8-GPU node needs sustained volume before it beats a per-token rate in the low single dollars. The reasons to self-host are the ones Mistral itself leads with: data that has to stay on your infrastructure, deployments that need to be auditable, and fine-tuning.
For everything else, the usual split holds. Call an API for flagship traffic and self-host the smaller open models where flat cost wins. GPT-OSS 120B and Gemma 4 are the Western open-weight models that fit on far less hardware, and Yotta AI Gateway puts open models like GLM 5.3, DeepSeek V4 Pro, and Kimi K3 behind one OpenAI-compatible key.
Frequently asked questions
How big is Mistral Large 4? 1.05 trillion total parameters with about 50 billion active per token, plus a 1.6B vision encoder, per Mistral's model docs. The checkpoint size isn't published. By arithmetic it's about 525 GB at 4-bit, 1.05 TB at FP8, and 2.1 TB at BF16.
How much VRAM does Mistral Large 4 need? It depends on the precision Mistral ships. Estimated planning figures are about 630 GB of GPU memory for a 4-bit checkpoint, 1.26 TB for FP8, and 2.5 TB for BF16. These are estimates until the weights are released.
Can I run Mistral Large 4 on a single GPU? No. The smallest estimate is over 500 GB of weights, and the largest single GPUs hold under 300 GB.
Can I run Mistral Large 4 on 8x H200? If Mistral ships a 4-bit checkpoint, yes on paper: 1,128 GB of GPU memory against about 630 GB needed. An FP8 checkpoint would not fit, because the weights alone are about 1.05 TB.
When are the Mistral Large 4 weights released? Mistral says by the end of October 2026. No exact date or download location has been given.
Is Mistral Large 4 open source? Mistral describes it as open-weight. The license hasn't been announced. Mistral Large 3 used Apache 2.0, and Mistral Large 2 used a research license.
Can I run Mistral Large 4 with Ollama or llama.cpp? Not today, since there are no weights to convert. Even after release, a model this size needs hundreds of gigabytes of memory at any usable quantization, so it won't be a laptop model.
How much does the Mistral Large 4 API cost? Mistral's announcement lists $1.36 per million input tokens and $4.18 per million output. The docs page currently shows $0.68 and $2.09 with the list prices struck through.
How does Mistral Large 4 compare to Mistral Large 3 on hardware? Large 3 is 675B parameters and Mistral said its NVFP4 checkpoint runs on a single 8x H100 node. Large 4 is about 1.56 times the size, which pushes a 4-bit checkpoint past what eight H100s hold comfortably and onto H200 or Blackwell nodes.
Bottom line
Mistral Large 4 is a 1.05 trillion-parameter model, and the weights will need somewhere between about 630 GB and 2.5 TB of GPU memory depending on the precision Mistral ships. If Mistral offers a 4-bit checkpoint, as it did for Large 3, plan on an 8x H200 node or a 4-GPU Blackwell tray. If it's FP8, plan on an 8-GPU Blackwell node. Full precision means a cluster.
The license, the checkpoint format, and engine support are the three things to watch between now and the end of October. This post will be updated with the verified numbers when the weights land. In the meantime, test the preview API against your own workload, and have an 8-GPU node picked out for release day.



