Three questions decide whether a model runs well at home

  1. Does it fit? The weights at your quantization, plus a KV cache that grows with your context length, plus the software's buffers, have to fit in GPU memory. If they do not, llama.cpp can run the rest of the layers from system RAM, at a large cost in speed. On a Mac or a unified-memory mini-PC, the GPU may use only part of the memory by default.
  2. How fast can it go? Every generated token reads all the active weights and the whole cache once, so memory bandwidth divided by those bytes is the most tokens per second any software can reach. That ceiling is arithmetic, and real speed is lower.
  3. What does it cost? Electricity at your state's price, the hardware spread over the months you will use it, against an API serving the same open model.

The calculators

The guides

What fits in common memory sizes

The largest models from the VRAM guide that fit at Q4_K_M with an 8k context and a 1 GiB allowance (gpt-oss in its published MXFP4). On a Mac, use the memory the GPU may use, about three quarters of the total, not the total itself.

Memory Fits at Q4_K_M, 8k context
8 GB Qwen3 8B, Llama 3.1 8B Instruct
12 GB Qwen3 14B, Qwen3 8B, Llama 3.1 8B Instruct
16 GB Mistral Small 3.2 24B Instruct (2506), gpt-oss-20b, Qwen3 14B, and 2 smaller
24 GB Qwen3 32B, Gemma 4 31B IT, Qwen3 30B-A3B Instruct (2507), and 6 smaller
32 GB Qwen3 32B, Gemma 4 31B IT, Qwen3 30B-A3B Instruct (2507), and 6 smaller
48 GB Qwen2.5 72B Instruct, Llama 3.3 70B Instruct, Qwen3 32B, and 8 smaller
64 GB gpt-oss-120b, Qwen2.5 72B Instruct, Llama 3.3 70B Instruct, and 9 smaller
96 GB GLM-4.5-Air, gpt-oss-120b, Qwen2.5 72B Instruct, and 10 smaller
128 GB GLM-4.5-Air, gpt-oss-120b, Qwen2.5 72B Instruct, and 10 smaller
192 GB Qwen3 235B-A22B Instruct (2507), GLM-4.5-Air, gpt-oss-120b, and 11 smaller
256 GB Qwen3 235B-A22B Instruct (2507), GLM-4.5-Air, gpt-oss-120b, and 11 smaller
512 GB Qwen3 235B-A22B Instruct (2507), GLM-4.5-Air, gpt-oss-120b, and 11 smaller

How this section works

Specifications from the vendors. Memory sizes, memory bandwidth and rated power come from NVIDIA's, AMD's, Intel's and Apple's own spec pages, datasheets and whitepapers, with the link on every row of the hardware table. When a vendor does not publish a figure, the table says so and the calculator does not guess.

The models' own shapes. Layer counts, KV heads and head sizes come from each model's config.json, the same data as the VRAM calculator, and quantization sizes from llama.cpp's own tables.

Ceilings, not benchmarks. The speeds on these pages are upper bounds from memory bandwidth. Measured tokens per second on real home hardware are planned, and will be published with their scripts and raw results.

No prices from retailers. Hardware prices change too often to print. The electricity calculator takes what you paid, and the guides group hardware by what it can do. Some pages link to retailers; any paid link is labeled "(paid link)", and the affiliate disclosure lists every relationship.

Also available as Markdown.