Guide
How much VRAM do you need to run Llama, Qwen or gpt-oss locally?
Memory needed by the open models people run at home, at every common GGUF quantization and at 8k, 32k and 128k tokens of context, worked out from each model's config.json and llama.cpp's bits-per-weight figures. Includes the smallest GPU or Mac memory size that holds each one.
The short answer: at the usual Q4_K_M quantization, a model takes a little over half a gigabyte per billion parameters, plus a cache that grows with every token of context, plus about a gigabyte for the software. The table below works it out for the models people ask about, with the Can I run it? calculator's own code, so you can check any other combination there.
The rule, and why "parameters x 4 bits" is wrong
memory = parameters x bits per weight / 8 the weights
+ 2 x layers x KV heads x head dim x 2 bytes x context tokens the KV cache (F16)
+ about 1 GiB llama.cpp's buffers and the driver (an assumption)
A "4-bit" GGUF is not 4 bits per weight. llama.cpp's quantization README measures Q4_K_M at 4.89 bits per weight and Q8_0 at 8.50 on Llama 3.1 8B, because each block of weights stores its scales and the mixes keep some tensors at higher precision. So Llama 3.1 8B at Q4_K_M is 4.58 GiB, not 4 GB, and Llama 3.3 70B is 40.2 GiB.
Memory per model and quantization
Weights in GiB. "Total at 8k" adds an F16 KV cache for 8,192 tokens and the 1 GiB allowance; "fits in" is the smallest memory size GPUs and Macs ship with (8, 10, 12, 16, 20, 24, 32, 48, 64, 96, 128, 192, 256 or 512 GB) that holds that total.
| Model | Parameters | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | F16 | Total at 8k, Q4_K_M | Fits in |
|---|---|---|---|---|---|---|---|---|
| Llama 3.1 8B Instruct | 8B | 4.58 | 5.33 | 6.14 | 7.95 | 14.96 | 6.58 | 8 GB |
| Qwen3 8B | 8.2B | 4.67 | 5.44 | 6.26 | 8.11 | 15.26 | 6.79 | 8 GB |
| Qwen3 14B | 14.8B | 8.41 | 9.81 | 11.28 | 14.62 | 27.51 | 10.66 | 12 GB |
| Mistral Small 3.2 24B Instruct (2506) | 24B | 13.68 | 15.94 | 18.35 | 23.76 | 44.72 | 15.93 | 16 GB |
| Gemma 4 31B IT | 31.3B | 17.82 | 20.76 | 23.89 | 30.95 | 58.25 | 20.23 | 24 GB |
| Qwen3 32B | 32.8B | 18.67 | 21.75 | 25.03 | 32.42 | 61.02 | 21.67 | 24 GB |
| Llama 3.3 70B Instruct | 70.6B | 40.2 | 46.85 | 53.91 | 69.82 | 131.42 | 43.7 | 48 GB |
| Qwen2.5 72B Instruct | 72.7B | 41.43 | 48.28 | 55.55 | 71.95 | 135.43 | 44.93 | 48 GB |
| Qwen3 30B-A3B Instruct (2507) | 30.5B (3.3B active) | 17.4 | 20.27 | 23.33 | 30.22 | 56.87 | 19.15 | 20 GB |
| gpt-oss-20b | 20.9B (3.6B active) | 11.92 | 13.89 | 15.98 | 20.7 | 38.96 | 13.11 | 16 GB |
| Gemma 4 26B-A4B IT | 25.8B (3.8B active) | 14.7 | 17.13 | 19.72 | 25.54 | 48.07 | 16.06 | 20 GB |
| GLM-4.5-Air | 110.5B (12B active) | 62.94 | 73.35 | 84.41 | 109.32 | 205.76 | 65.38 | 96 GB |
| gpt-oss-120b | 116.8B (5.1B active) | 66.57 | 77.57 | 89.27 | 115.62 | 217.61 | 67.85 | 96 GB |
| Qwen3 235B-A22B Instruct (2507) | 235.1B (22B active) | 133.95 | 156.1 | 179.63 | 232.65 | 437.9 | 136.42 | 192 GB |
GiB is 1,073,741,824 bytes. Graphics card memory is binary, so a "24 GB" card holds 24 GiB. On a Mac the GPU may use only part of unified memory (see below), so a Mac needs more memory than the "fits in" column says.
gpt-oss is published in MXFP4, OpenAI's 4-bit format, and that is the file to run: gpt-oss-20b is 12.82 GiB (14.01 GiB at 8k, so a 16 GB card) and gpt-oss-120b is 60.77 GiB (62.05 GiB at 8k: more than any GeForce card holds, so a 72 or 96 GB workstation card, or a Mac or unified-memory PC with 96 GB or more, once the GPU's share is counted). The GGUF columns for gpt-oss show what a re-quantized file would take.
The KV cache: where long context goes
The cache stores a key and a value for every token in the context, for every layer, and llama.cpp allocates it for the whole context length when the model loads. It depends on the architecture, not just the size:
| Model | Cache per token (F16) | 8k tokens | 32k tokens | 128k tokens |
|---|---|---|---|---|
| Llama 3.1 8B Instruct | 128 KiB | 1 GiB | 4 GiB | 16 GiB |
| Qwen3 8B | 144 KiB | 1.13 GiB | 4.5 GiB | beyond its maximum |
| Qwen3 14B | 160 KiB | 1.25 GiB | 5 GiB | beyond its maximum |
| Mistral Small 3.2 24B Instruct (2506) | 160 KiB | 1.25 GiB | 5 GiB | 20 GiB |
| Gemma 4 31B IT | 80 KiB | 1.41 GiB | 3.28 GiB | 10.78 GiB |
| Qwen3 32B | 256 KiB | 2 GiB | 8 GiB | beyond its maximum |
| Llama 3.3 70B Instruct | 320 KiB | 2.5 GiB | 10 GiB | 40 GiB |
| Qwen2.5 72B Instruct | 320 KiB | 2.5 GiB | 10 GiB | beyond its maximum |
| Qwen3 30B-A3B Instruct (2507) | 96 KiB | 0.75 GiB | 3 GiB | 12 GiB |
| gpt-oss-20b | 24 KiB | 0.19 GiB | 0.75 GiB | 3 GiB |
| Gemma 4 26B-A4B IT | 20 KiB | 0.35 GiB | 0.82 GiB | 2.7 GiB |
| GLM-4.5-Air | 184 KiB | 1.44 GiB | 5.75 GiB | 23 GiB |
| gpt-oss-120b | 36 KiB | 0.29 GiB | 1.13 GiB | 4.5 GiB |
| Qwen3 235B-A22B Instruct (2507) | 188 KiB | 1.47 GiB | 5.88 GiB | 23.5 GiB |
Three things make the differences:
- Grouped-query attention. Llama 3.1 8B has 32 attention heads but caches only 8 KV heads, so its cache is 128 KiB per token. Calculators that use the attention-head count overstate it four times.
- Sliding windows. Gemma 4 and gpt-oss cache only a short window on most layers, so Gemma 4 31B needs 3.28 GiB at 32k tokens where Qwen3 32B needs 8. llama.cpp caches only the window by default (
--swa-fullis off). - Layers and KV heads. Llama 3.3 70B's 80 layers make its cache 320 KiB per token: 10 GiB at 32k tokens, more than an 8 GB card holds in total.
A q8_0 cache (--cache-type-k q8_0 --cache-type-v q8_0) takes 34 bytes per 32 values instead of 64, about half. Whether that changes the answers depends on the model; try it on your own prompts.
Mixture of experts: big in memory, fast to run
A mixture-of-experts model has to hold every expert in memory, but each token uses only a few of them. Qwen3 30B-A3B needs memory for 30.5 billion parameters (17.4 GiB at Q4_K_M) and reads only about 3.3 billion per token, so its speed ceiling is close to a 3B dense model's while it needs the memory of a 30B one. That is also why MoE models tolerate partial offload to system RAM better than dense ones.
When it does not fit
- Use a smaller quantization. Each step down saves memory and costs some accuracy. llama.cpp's own figures for Llama 3 8B show perplexity rising by 0.18 at Q4_K_M, 0.66 at Q3_K_M and 3.52 at Q2_K over the 16-bit model (quantize.cpp), which is why Q4_K_M is the usual compromise. Other models lose different amounts.
- Shorten the context. The cache is the part you control at run time.
- Offload to system RAM. llama.cpp runs the layers that do not fit on the CPU. It works, but the speed then depends on system RAM bandwidth: Llama 3.3 70B at Q4_K_M on a 24 GB card keeps 37 of its 80 layers in DDR5-5600, and its ceiling drops to 3.9 tokens per second, as the calculator shows layer by layer.
- Use unified memory. A Mac, a Ryzen AI Max mini-PC or a DGX Spark has far more memory than a graphics card, with less bandwidth. The Apple Silicon vs NVIDIA guide compares them.
Macs: the GPU gets only part of the memory
macOS lets the GPU use part of unified memory by default. Apple documents no fixed rule, but its 2021 tech talk on Metal compute gives two examples, 21 GB of 32 GB and 48 GB of 64 GB, and the calculator follows them: two thirds up to 32 GB, three quarters above. A 64 GB Mac therefore has about 48 GB for a model, enough for Llama 3.3 70B at Q4_K_M (43.7 GiB at 8k) but not at Q5_K_M. The limit can be raised with sudo sysctl iogpu.wired_limit_mb=<megabytes>, which Apple's MLX project documents, as long as macOS keeps enough for itself.
How these numbers were made
Every figure on this page is computed by the same code as the Can I run it? calculator, from the models' config.json files (models.json) and llama.cpp's bits-per-weight table (local-ai-assumptions.json). A test fails the site build if this page's table and the calculator ever disagree. Nothing here is a measurement, and actual memory use differs by engine and version by a few hundred megabytes; the methodology lists the assumptions.
Also available as Markdown.