---
title: "GPT-OSS 120B Hardware Requirements: GPU, Memory, and How to Run It (2026)"
slug: gpt-oss-120b-hardware-requirements-gpu-memory-2026
description: "OpenAI's open-weight 120B fits one 80 GB GPU, and the 20B runs in 16 GB. The exact memory math, which cards work, the vLLM and Ollama commands, and where it stands against 2026 open models.
"
author: "Yotta Labs"
date: 2026-10-03
categories: ["Inference"]
canonical: https://www.yottalabs.ai/post/gpt-oss-120b-hardware-requirements-gpu-memory-2026
---

# GPT-OSS 120B Hardware Requirements: GPU, Memory, and How to Run It (2026)

![](https://cdn.sanity.io/images/wy75wyma/production/d9e5969c98beb6b15180fecc802b8cb66fa00157-1200x627.png)

*A 117B model that runs on a single H100 sounds like a typo. It isn’t. Here’s why, and what you actually need.*

## TL;DR

gpt-oss-120b is a 117B parameter mixture-of-experts model with 5.1B active per token, shipped by OpenAI under Apache 2.0 with its MoE weights already quantized to MXFP4. That’s why the checkpoint is about 61 GB instead of 234 GB, and why OpenAI’s own line is that it runs within 80 GB of memory. One H100, one H200, one RTX Pro 6000, or one MI300X will hold it. gpt-oss-20b is the same design at 21B total and 3.6B active, about 13 GB on disk, and runs in 16 GB, which means a single consumer card or a laptop.

The catch isn’t memory, it’s age. These models are from August 2025 and the open field has moved. On the current Artificial Analysis index gpt-oss-120b scores 12 where Qwen 3.8 27B scores 52. If you need the best open model per GPU, that’s a different post. If you need an Apache 2.0 model from a US lab that fits one card, has a year of tooling behind it, and is the MLPerf reference workload, keep reading.

## The numbers that matter

<!-- unsupported block: table -->

Both figures in the download row are what you’ll actually pull. The Hugging Face safetensors for the 120B come in around 61 GB; Ollama’s build is 65 GB. Neither needs further quantization to fit an 80 GB card, which is unusual and is the whole point of OpenAI shipping in MXFP4.

## Why it fits one GPU

Two things stack. First, only 4 of 128 experts fire per token, so compute per token is that of a ~5B model even though you store 117B. Second, OpenAI quantized the expert weights, which are over 90% of the parameters, to MXFP4 at roughly 4.25 bits each. Everything else (attention, embeddings, the router) stays in higher precision, but it’s a small share of the total. 117B parameters at 4.25 bits lands near 61 GB. Compare that with a dense 70B model, which needs 140 GB at BF16 and still 35 GB at 4-bit.

That leaves roughly 19 GB on an 80 GB card for KV cache and activations. gpt-oss alternates sliding-window and full-attention layers and uses grouped-query attention with 8 KV heads, so KV cache per token is small by 2026 standards, on the order of a few GB for a full 128K sequence at BF16 (our estimate from the architecture; measure on your engine). Practically: one H100 serves the 120B fine at moderate context and concurrency. For long contexts and real batch sizes, an H200’s 141 GB is the comfortable answer.

## GPU tiers

<!-- unsupported block: table -->

For the 20B, anything with 16 GB works: RTX 4090, RTX 5090, a 16 GB Apple Silicon machine for local use. On a 24 GB card you get room for context; on an RTX Pro 6000 you can run several copies or a large batch.

## Engines and the launch command

vLLM has supported gpt-oss natively since the 0.10 line, with the MXFP4 kernels built in. On a current release it’s one line:

```bash
vllm serve openai/gpt-oss-120b
```

That gives you an OpenAI-compatible endpoint on port 8000. Reasoning effort is a system-prompt setting (low, medium, high), not a flag, so your clients pick it per request.

See [what vLLM actually is](https://www.yottalabs.ai/post/what-is-vllm-architecture-performance-and-why-teams-use-it-for-llm-inference) for the engine, [how to deploy vLLM in production with Docker](https://www.yottalabs.ai/post/how-to-deploy-vllm-in-production-with-docker) for the container setup, and [the vLLM OpenAI-compatible server](https://www.yottalabs.ai/post/vllm-openai-compatible-server) for the API surface.

For local or single-user use, Ollama is the shortest path:

```bash
ollama run gpt-oss:120b
```

```bash
ollama run gpt-oss:20b
```

llama.cpp, LM Studio, and Transformers all support both sizes as well. SGLang works too; for a single-model, single-GPU deployment the engine choice matters less than the GPU choice.

## gpt-oss-20b: the version most people should start with

The 20B is not a cut-down demo. It’s the same architecture, same 128K context, same reasoning modes, same license, at 21B total and 3.6B active. OpenAI positions it at roughly o3-mini level on its own benchmarks. It fits a 16 GB card, fine-tunes on consumer hardware, and a 24 GB GPU gives it headroom for real context. If your workload is classification, extraction, routing, or an agent sub-step, test the 20B first and only move up if quality demands it. The cost difference between a 16 GB card and an 80 GB card is the whole budget conversation.

## Where it stands in 2026

Be honest with yourself about this one. gpt-oss shipped in August 2025. Since then Qwen 3.8, DeepSeek V4.1, GLM 5.3, Kimi K3, and Gemma 4 have all landed, most under MIT or Apache 2.0, and the independent numbers reflect it: gpt-oss-120b sits at 12 on the current Artificial Analysis Intelligence Index (v4.3.2), against 52 for Qwen 3.8 27B, a dense model that needs less memory. On vendor benchmarks OpenAI claimed near parity with o4-mini at launch; nobody is comparing against o4-mini anymore.

What it still has: an Apache 2.0 license from a US lab, which matters for some procurement teams; a year of production tooling and quantized builds; a policy-classification variant (gpt-oss-safeguard, October 2025); and status as an official MLPerf Inference v6.0 workload since March 2026, so every hardware vendor publishes numbers on it. Hosted pricing has collapsed to the $0.04 to $0.15 per million input range across providers, so if you only want the API, self-hosting rarely wins on cost. Self-host it when you need the weights on your own hardware for data, latency, or licensing reasons.

If raw capability per GPU is the goal, read the [Qwen 3.8 27B hardware guide](https://www.yottalabs.ai/post/qwen-3-8-27b-specs-hardware-requirements-how-to-run-2026) and the [best open-source LLMs](https://www.yottalabs.ai/post/best-open-source-llms-2026) roundup before you commit.

## The route most teams should take

Test the 20B on whatever GPU you already have. If it clears your eval, you’re done, and you can serve it on a single 24 GB card. If it doesn’t, put the 120B on one H200 with vLLM and run the same eval; the H200 gives you the context and batch headroom an H100 doesn’t, and it’s still one card. Only go to multi-GPU if you need throughput, not to fit the weights.

On Yotta, that’s a single-GPU Pod: pick an H100, H200, or RTX Pro 6000 in the [console](https://console.yottalabs.ai/), deploy, and run the vLLM command above. Current rates are on the [pricing page](https://yottalabs.ai/pricing). For the hosted route, the [AI Gateway](https://yottalabs.ai/ai-gateway) serves current open models behind one OpenAI-compatible key, with new models added as they ship.

## Frequently asked questions

**How much VRAM does gpt-oss-120b need?**

About 61 GB for the weights plus headroom for KV cache. OpenAI’s floor is 80 GB. One H100, H200, RTX Pro 6000 (96 GB), or MI300X holds it.

**Can gpt-oss-120b run on an RTX 5090?**

No. 32 GB isn’t enough for 61 GB of weights. Two RTX 5090s don’t cover it either once you add cache. Run the 20B on a 5090; it needs 16 GB.

**Can I run gpt-oss-120b locally?**

Yes, if “locally” means a workstation with an 80 GB or 96 GB card, or a Mac with 96 GB or more of unified memory via Ollama or llama.cpp. On a typical gaming PC, the 20B is the local option.

**Do I need to quantize it?**

No. The MoE weights ship in MXFP4 already. Further quantization buys little memory and costs quality.

**Is gpt-oss open source?**

Open weights under Apache 2.0, which is as permissive as model licenses get. Training data and code are not released.

**Is gpt-oss better than Qwen 3.8 27B?**

Not on independent benchmarks. Artificial Analysis scores the 27B at 52 and gpt-oss-120b at 12 on the current index, and the 27B needs less memory. gpt-oss wins on license origin, tooling maturity, and the 20B’s 16 GB footprint.

**Is gpt-oss on Yotta?**

You can run either size on a Yotta GPU Pod today with vLLM or Ollama. Gateway availability for hosted gpt-oss is listed in the console model catalog.

## Bottom line

gpt-oss-120b is the easiest 100B-class model to put on one GPU, because OpenAI did the quantization for you. An H100 holds it, an H200 holds it comfortably, and the 20B runs on almost anything. The reason to run it in 2026 is the license and the tooling, not the benchmark score, so test the 20B first, benchmark the 120B against a current open model on your own eval, and let that decide.
