Oct 05, 2026
Gemma 4 Hardware Requirements: GPU and Memory for 31B, 26B, 12B, and E4B (2026)
Cost Optimization
Google’s Gemma 4 comes in five sizes under Apache 2.0, from a phone model to a 31B that fits one H100. What each size needs in VRAM, which cards work, the Ollama and vLLM commands, and where it stands against other 2026 open models.

Five models, one license, and a 31B that Google built to fit a single 80 GB card. Here’s the memory math for each one.
TL;DR
Gemma 4 is Google’s open-weight family, released April 2, 2026 under Apache 2.0 with no user caps or revenue terms. Four sizes shipped at launch and a fifth followed: E2B (5.1B total, 2.3B effective) and E4B (8B, 4.5B effective) for phones and edge boxes; a 12B dense model; a 26B mixture-of-experts with 3.8B active per token; and a 31B dense flagship. The small ones run on almost anything. The 26B and 31B both fit a single 24 GB consumer card at 4-bit, and the 31B fits one H100 unquantized, which is the line Google led with.
What it isn’t: the best open model per GPU in late 2026. On the current Artificial Analysis index the 31B scores 15 where Qwen 3.8 27B scores 52. Gemma 4’s case is the license, the multimodal input (image on all sizes, audio on the small ones and the 12B), the 256K context, and the fact that Google’s own tooling supports it everywhere from a Raspberry Pi to vLLM.
The numbers that matter
| Model | Total params | Active per token | Context | Inputs | BF16 weights | 4-bit (Ollama) |
| Gemma 4 E2B | 5.1B | 2.3B effective | 128K | Text, image, audio | About 10 GB | 4.6 to 7.5 GB |
| Gemma 4 E4B | 8B | 4.5B effective | 128K | Text, image, audio | About 16 GB | 6.6 to 9.5 GB |
| Gemma 4 12B | 11.95B | 11.95B (dense) | 256K | Text, image, audio | About 24 GB | About 8 GB |
| Gemma 4 26B A4B | 25.2B | 3.8B (8 of 128 experts) | 256K | Text, image | About 50 GB | 16 to 19 GB |
| Gemma 4 31B | 30.7B | 30.7B (dense) | 256K | Text, image | About 61 GB | 19 to 20 GB |
BF16 sizes are parameter count times two bytes. The 4-bit column is the range of Ollama’s published download sizes for each tag. Add KV cache on top of either; Gemma 4 uses a mix of sliding-window and global attention layers, so the cache per token is moderate, but at 256K context on the big models you should measure on your engine rather than guess.
Why the 31B fits one H100
Google’s exact phrase: the 31B’s unquantized bfloat16 weights fit on a single 80 GB H100. At 61 GB of weights that leaves about 19 GB for KV cache and activations, which is enough for moderate context and a few concurrent requests. It’s the same arithmetic as gpt-oss-120b, which lands at 61 GB too, except Gemma gets there with a dense 31B at full precision rather than a 117B sparse model at 4 bits. In practice: an H100 is the floor for serving the 31B at BF16; an H200’s 141 GB is where the full 256K context and real batch sizes live.
Drop to FP8 and the 31B is about 31 GB, which fits a 48 GB L40S with room, and an RTX 5090’s 32 GB with almost none. Drop to 4-bit and it’s a 19 to 20 GB download that runs on an RTX 4090.
GPU tiers for the 31B
| GPU | Memory | BF16 (61 GB) | FP8 (31 GB) | 4-bit (20 GB) |
| H200 141 GB | 141 GB | Fits, full context and batch | Fits | Fits |
| H100 80 GB | 80 GB | Fits, moderate context | Fits | Fits |
| RTX Pro 6000 96 GB | 96 GB | Fits | Fits | Fits |
| L40S 48 GB | 48 GB | No | Fits | Fits |
| RTX 5090 32 GB | 32 GB | No | Tight, short context only | Fits |
| RTX 4090 24 GB | 24 GB | No | No | Fits |
The 26B A4B: the one to test first
The 26B is the interesting one for anyone who has to pay for tokens at volume. Only 8 of 128 experts fire per token, 3.8B parameters, so decode is fast and cheap while you still get 25B of stored knowledge. At 4-bit it’s a 16 to 19 GB download that runs on a 24 GB card; at BF16 it’s about 50 GB and wants an 80 GB card, same tier as the 31B. On Google’s own benchmarks it trails the 31B by two or three points on most tasks (MMLU Pro 82.6 vs 85.2, GPQA Diamond 82.3 vs 84.3, LiveCodeBench 77.1 vs 80.0) and by more on competitive coding (Codeforces 1718 vs 2150). Those are vendor numbers. If your workload is high-volume and the 26B clears your eval, it’s the cheaper model to serve by a wide margin.
The 12B is the quiet middle option: dense, audio input, 256K context, about 24 GB at BF16 so it fits a 32 GB RTX 5090 at full precision and a 12 GB card at 4-bit. E2B and E4B are for phones, Jetson boards, and laptops; a 4-bit E4B is under 10 GB and runs on a 16 GB machine with room to spare.
Engines and the launch command
vLLM supports the whole family. For the 31B on one 80 GB card:
vllm serve google/gemma-4-31B-itFor the 26B:
vllm serve google/gemma-4-26B-A4B-itFor local use, Ollama has every size as a tag, plus MLX builds for Apple Silicon:
ollama run gemma4:31bollama run gemma4:26bollama run gemma4:12bollama run gemma4:e4bllama.cpp, LM Studio, SGLang, Unsloth, NVIDIA NIM, and Transformers all support Gemma 4 as well. If vLLM is new to you, what vLLM actually is covers the engine and how to deploy vLLM in production with Docker covers running it as a service.
Where it stands in 2026
Google launched Gemma 4 claiming the 31B was the #3 open model on the Arena text leaderboard and that it beat models 20 times its size. Six months on, the independent picture is different. Artificial Analysis scores Gemma 4 31B at 15 on its Intelligence Index (v4.3.2); Qwen 3.8 27B, a dense model that needs less memory, scores 52, and GLM 5.3 Flash scores 42. Gemma 4 is a competent, well-supported family that has been passed on raw capability by the open models that shipped after it.
What it still has going for it, and it’s a real list: Apache 2.0 with no strings, from a US lab, which some procurement teams require; image input on every size and audio on three of them, which most open LLMs don’t offer at all; 256K context on the three larger models; first-class support in Google’s edge stack (LiteRT-LM, Jetson, Raspberry Pi) for the small ones; and the broadest framework coverage of any open family, because Google lined up every partner on day one.
For a capability-first pick on one GPU, read the Qwen 3.8 27B hardware guide. For the other US-lab, Apache 2.0 option that fits one 80 GB card, read GPT-OSS 120B hardware requirements. The best open-source LLMs roundup puts them all side by side.
The route most teams should take
Run the 26B at 4-bit on whatever 24 GB card you have and put it through your eval. If it passes, serve it there or on an L40S and you’re done. If you need the extra points, the 31B at BF16 on one H100 or H200 with vLLM is the production config, and FP8 on an L40S is the budget version. Pick the 12B or E4B only if audio input or an edge target is the requirement.
On Yotta, every one of those is a single-GPU Pod: choose the card in the console, deploy, run the vLLM line above. Rates are on the pricing page. If you’d rather not host, the AI Gateway serves current open models behind one OpenAI-compatible key, with new models added as they ship.
Frequently asked questions
How much VRAM does Gemma 4 31B need?
About 61 GB at BF16, so an 80 GB card is the floor. Roughly 31 GB at FP8, and 19 to 20 GB at 4-bit, which fits a 24 GB consumer GPU.
Can Gemma 4 run on an RTX 4090?
Yes. The 31B and 26B both run at 4-bit on 24 GB. The 12B runs at 4-bit with lots of room, and E4B and E2B run at any precision.
Which is better, Gemma 4 26B or 31B?
The 31B scores two to three points higher on Google’s benchmarks and more on competitive coding. The 26B activates 3.8B parameters per token, so it decodes faster and costs less to serve at the same memory tier. Test the 26B first.
Is Gemma 4 open source?
Open weights under Apache 2.0, with no monthly active user limit and no revenue-sharing clause. Training data and code are not released.
Is Gemma 4 better than Qwen 3.8 27B?
Not on independent benchmarks. Artificial Analysis scores the 27B at 52 and Gemma 4 31B at 15 on the current index. Gemma 4 wins on license terms, multimodal input, and edge support.
Does Gemma 4 support audio and video?
Audio input on E2B, E4B, and the 12B. Image input on every size. Video is handled as image frames on the larger models; check the model card for the current limits.
Is Gemma 4 on Yotta?
Every size runs on a Yotta GPU Pod today with vLLM or Ollama. Gateway availability for hosted Gemma 4 is listed in the console model catalog.
Bottom line
Gemma 4 is the easiest open family to fit to hardware you already have: a 31B that fits one H100 at full precision, a 26B that fits a 4090 at 4-bit, and small models that run on a phone, all under Apache 2.0. It is not the strongest open model you can run on a given GPU anymore. Choose it for the license, the modalities, and the edge story; benchmark it against a current open model on your own eval before you commit production traffic to it.



