The short answer: at the usual Q4_K_M quantization, a model takes a little over half a gigabyte per billion parameters, plus a cache that grows with every token of context, plus about a gigabyte for the software. The table below works it out for the models people ask about, with the Can I run it? calculator's own code, so you can check any other combination there.

The rule, and why "parameters x 4 bits" is wrong

memory = parameters x bits per weight / 8      the weights
       + 2 x layers x KV heads x head dim x 2 bytes x context tokens      the KV cache (F16)
       + about 1 GiB                           llama.cpp's buffers and the driver (an assumption)

A "4-bit" GGUF is not 4 bits per weight. llama.cpp's quantization README measures Q4_K_M at 4.89 bits per weight and Q8_0 at 8.50 on Llama 3.1 8B, because each block of weights stores its scales and the mixes keep some tensors at higher precision. So Llama 3.1 8B at Q4_K_M is 4.58 GiB, not 4 GB, and Llama 3.3 70B is 40.2 GiB.

Memory per model and quantization

Weights in GiB. "Total at 8k" adds an F16 KV cache for 8,192 tokens and the 1 GiB allowance; "fits in" is the smallest memory size GPUs and Macs ship with (8, 10, 12, 16, 20, 24, 32, 48, 64, 96, 128, 192, 256 or 512 GB) that holds that total.

Model Parameters Q4_K_M Q5_K_M Q6_K Q8_0 F16 Total at 8k, Q4_K_M Fits in
Llama 3.1 8B Instruct 8B 4.58 5.33 6.14 7.95 14.96 6.58 8 GB
Qwen3 8B 8.2B 4.67 5.44 6.26 8.11 15.26 6.79 8 GB
Qwen3 14B 14.8B 8.41 9.81 11.28 14.62 27.51 10.66 12 GB
Mistral Small 3.2 24B Instruct (2506) 24B 13.68 15.94 18.35 23.76 44.72 15.93 16 GB
Gemma 4 31B IT 31.3B 17.82 20.76 23.89 30.95 58.25 20.23 24 GB
Qwen3 32B 32.8B 18.67 21.75 25.03 32.42 61.02 21.67 24 GB
Llama 3.3 70B Instruct 70.6B 40.2 46.85 53.91 69.82 131.42 43.7 48 GB
Qwen2.5 72B Instruct 72.7B 41.43 48.28 55.55 71.95 135.43 44.93 48 GB
Qwen3 30B-A3B Instruct (2507) 30.5B (3.3B active) 17.4 20.27 23.33 30.22 56.87 19.15 20 GB
gpt-oss-20b 20.9B (3.6B active) 11.92 13.89 15.98 20.7 38.96 13.11 16 GB
Gemma 4 26B-A4B IT 25.8B (3.8B active) 14.7 17.13 19.72 25.54 48.07 16.06 20 GB
GLM-4.5-Air 110.5B (12B active) 62.94 73.35 84.41 109.32 205.76 65.38 96 GB
gpt-oss-120b 116.8B (5.1B active) 66.57 77.57 89.27 115.62 217.61 67.85 96 GB
Qwen3 235B-A22B Instruct (2507) 235.1B (22B active) 133.95 156.1 179.63 232.65 437.9 136.42 192 GB

GiB is 1,073,741,824 bytes. Graphics card memory is binary, so a "24 GB" card holds 24 GiB. On a Mac the GPU may use only part of unified memory (see below), so a Mac needs more memory than the "fits in" column says.

gpt-oss is published in MXFP4, OpenAI's 4-bit format, and that is the file to run: gpt-oss-20b is 12.82 GiB (14.01 GiB at 8k, so a 16 GB card) and gpt-oss-120b is 60.77 GiB (62.05 GiB at 8k: more than any GeForce card holds, so a 72 or 96 GB workstation card, or a Mac or unified-memory PC with 96 GB or more, once the GPU's share is counted). The GGUF columns for gpt-oss show what a re-quantized file would take.

The KV cache: where long context goes

The cache stores a key and a value for every token in the context, for every layer, and llama.cpp allocates it for the whole context length when the model loads. It depends on the architecture, not just the size:

Model Cache per token (F16) 8k tokens 32k tokens 128k tokens
Llama 3.1 8B Instruct 128 KiB 1 GiB 4 GiB 16 GiB
Qwen3 8B 144 KiB 1.13 GiB 4.5 GiB beyond its maximum
Qwen3 14B 160 KiB 1.25 GiB 5 GiB beyond its maximum
Mistral Small 3.2 24B Instruct (2506) 160 KiB 1.25 GiB 5 GiB 20 GiB
Gemma 4 31B IT 80 KiB 1.41 GiB 3.28 GiB 10.78 GiB
Qwen3 32B 256 KiB 2 GiB 8 GiB beyond its maximum
Llama 3.3 70B Instruct 320 KiB 2.5 GiB 10 GiB 40 GiB
Qwen2.5 72B Instruct 320 KiB 2.5 GiB 10 GiB beyond its maximum
Qwen3 30B-A3B Instruct (2507) 96 KiB 0.75 GiB 3 GiB 12 GiB
gpt-oss-20b 24 KiB 0.19 GiB 0.75 GiB 3 GiB
Gemma 4 26B-A4B IT 20 KiB 0.35 GiB 0.82 GiB 2.7 GiB
GLM-4.5-Air 184 KiB 1.44 GiB 5.75 GiB 23 GiB
gpt-oss-120b 36 KiB 0.29 GiB 1.13 GiB 4.5 GiB
Qwen3 235B-A22B Instruct (2507) 188 KiB 1.47 GiB 5.88 GiB 23.5 GiB

Three things make the differences:

A q8_0 cache (--cache-type-k q8_0 --cache-type-v q8_0) takes 34 bytes per 32 values instead of 64, about half. Whether that changes the answers depends on the model; try it on your own prompts.

Mixture of experts: big in memory, fast to run

A mixture-of-experts model has to hold every expert in memory, but each token uses only a few of them. Qwen3 30B-A3B needs memory for 30.5 billion parameters (17.4 GiB at Q4_K_M) and reads only about 3.3 billion per token, so its speed ceiling is close to a 3B dense model's while it needs the memory of a 30B one. That is also why MoE models tolerate partial offload to system RAM better than dense ones.

When it does not fit

Macs: the GPU gets only part of the memory

macOS lets the GPU use part of unified memory by default. Apple documents no fixed rule, but its 2021 tech talk on Metal compute gives two examples, 21 GB of 32 GB and 48 GB of 64 GB, and the calculator follows them: two thirds up to 32 GB, three quarters above. A 64 GB Mac therefore has about 48 GB for a model, enough for Llama 3.3 70B at Q4_K_M (43.7 GiB at 8k) but not at Q5_K_M. The limit can be raised with sudo sysctl iogpu.wired_limit_mb=<megabytes>, which Apple's MLX project documents, as long as macOS keeps enough for itself.

How these numbers were made

Every figure on this page is computed by the same code as the Can I run it? calculator, from the models' config.json files (models.json) and llama.cpp's bits-per-weight table (local-ai-assumptions.json). A test fails the site build if this page's table and the calculator ever disagree. Nothing here is a measurement, and actual memory use differs by engine and version by a few hundred megabytes; the methodology lists the assumptions.

Also available as Markdown.