How this list is made

Two numbers decide what a GPU can do with a language model at home:

  1. Memory decides what runs. The model's weights, its KV cache and the software's buffers have to fit. If they do not, the model either does not load or runs partly from system RAM, which is much slower. The VRAM guide has the figure for each model.
  2. Memory bandwidth decides how fast it can generate. Every generated token reads all the active weights once, so bandwidth divided by bytes per token is the most tokens per second any software can reach. Real engines stay below this ceiling.

So the tiers below are memory sizes, and within each tier the cards are ranked by bandwidth. Every speed is the ceiling from the Can I run it? calculator for one conversation at an 8,192-token context, in tokens per second: an upper bound, not a benchmark. "Offload, n" means the model only runs with part of it in system RAM (assumed 32 GB of DDR5-5600), and n is the ceiling then. "No" means it does not fit even with RAM.

There are no prices here, because they change weekly and differ between sellers. Within a generation, more memory and more bandwidth cost more, so the tiers run from the least to the most expensive kind of card; check current prices yourself before choosing between tiers.

Software matters as much as the chip. NVIDIA cards use CUDA, which llama.cpp, Ollama, LM Studio, vLLM and nearly every other engine support. AMD cards run llama.cpp through ROCm on the cards AMD lists for it, or through Vulkan, which works on any card with a Vulkan driver; AMD's ROCm 10.1 compatibility list covers the current Radeon RX 7000 and 9000 cards except the RX 7600 XT. Intel Arc cards use llama.cpp's SYCL backend (the Arc B580 is on its verified list) or Vulkan. Engines beyond llama.cpp, such as vLLM, support NVIDIA best. The software notes for every card are in the calculator's hardware table.

Entry tier: 8 to 12 GB

An 8 GB card holds 8-billion-parameter models at Q4_K_M with room for a long context (6.58 GiB for Llama 3.1 8B at 8k), and little more. A 12 GB card also holds Qwen3 14B at Q4_K_M (10.66 GiB at 8k). gpt-oss-20b needs 14.01 GiB, so on these cards part of it runs from system RAM; as a mixture-of-experts model it stays usable that way.

Card Memory Bandwidth Llama 3.1 8B Q4_K_M gpt-oss-20b MXFP4 Qwen3 32B Q4_K_M Llama 3.3 70B Q4_K_M
NVIDIA GeForce RTX 3080 Ti 12 GB 912 GB/s 161 offload, 187.3 offload, 8.1 no
NVIDIA GeForce RTX 3080 (10 GB) 10 GB 760 GB/s 134 offload, 106.8 offload, 6.7 no
NVIDIA GeForce RTX 5070 12 GB 672 GB/s 119 offload, 162.3 offload, 7.8 no
NVIDIA GeForce RTX 4070 12 GB 504 GB/s 89 offload, 138.8 offload, 7.4 no
NVIDIA GeForce RTX 4070 SUPER 12 GB 504 GB/s 89 offload, 138.8 offload, 7.4 no
NVIDIA GeForce RTX 4070 Ti 12 GB 504 GB/s 89 offload, 138.8 offload, 7.4 no
Intel Arc B580 Graphics 12 GB 456 GB/s 80 offload, 130.8 offload, 7.3 no
NVIDIA GeForce RTX 3060 Ti 8 GB 448 GB/s 79 offload, 70.4 offload, 5.7 no
NVIDIA GeForce RTX 3070 8 GB 448 GB/s 79 offload, 70.4 offload, 5.7 no
NVIDIA GeForce RTX 5060 8 GB 448 GB/s 79 offload, 70.4 offload, 5.7 no
NVIDIA GeForce RTX 5060 Ti (8 GB) 8 GB 448 GB/s 79 offload, 70.4 offload, 5.7 no
AMD Radeon RX 7700 XT 12 GB 432 GB/s 76 offload, 126.6 offload, 7.2 no
AMD Radeon RX 9070 GRE 12 GB 432 GB/s 76 offload, 126.6 offload, 7.2 no
Intel Arc B570 Graphics 10 GB 380 GB/s 67 offload, 85.8 offload, 6.2 no
NVIDIA GeForce RTX 3060 (12 GB) 12 GB 360 GB/s 64 offload, 112.7 offload, 7 no
NVIDIA GeForce RTX 5050 8 GB 320 GB/s 56 offload, 64.8 offload, 5.5 no
AMD Radeon RX 9060 XT (8GB) 8 GB 320 GB/s 56 offload, 64.8 offload, 5.5 no
NVIDIA GeForce RTX 4060 Ti (8 GB) 8 GB 288 GB/s 51 offload, 62.9 offload, 5.4 no
AMD Radeon RX 7600 8 GB 288 GB/s 51 offload, 62.9 offload, 5.4 no
AMD Radeon RX 9050 8 GB 288 GB/s 51 offload, 62.9 offload, 5.4 no
AMD Radeon RX 9060 8 GB 288 GB/s 51 offload, 62.9 offload, 5.4 no
NVIDIA GeForce RTX 4060 8 GB 272 GB/s 48 offload, 61.8 offload, 5.4 no

What the arithmetic says: if you can, choose 12 GB over 8 GB, because it moves you from 8B to 14B models. Among the 12 GB cards, the NVIDIA GeForce RTX 3080 Ti has the highest bandwidth in this list.

Mainstream tier: 16 GB

16 GB holds gpt-oss-20b in its published MXFP4 form entirely in VRAM, Qwen3 14B at up to Q6_K, and Mistral Small 3.2 24B Instruct (2506) at Q4_K_M with an 8k context (15.93 GiB, with very little to spare). 30B-class dense models do not fit at Q4_K_M.

Card Memory Bandwidth Llama 3.1 8B Q4_K_M gpt-oss-20b MXFP4 Qwen3 32B Q4_K_M Llama 3.3 70B Q4_K_M
NVIDIA GeForce RTX 5080 16 GB 960 GB/s 170 373 offload, 12.5 no
NVIDIA GeForce RTX 5070 Ti 16 GB 896 GB/s 158 348 offload, 12.4 no
NVIDIA GeForce RTX 4080 SUPER 16 GB 736 GB/s 130 286 offload, 11.8 no
NVIDIA GeForce RTX 4080 16 GB 716.8 GB/s 126 279 offload, 11.7 no
NVIDIA GeForce RTX 4070 Ti SUPER 16 GB 672 GB/s 119 261 offload, 11.5 no
AMD Radeon RX 9070 16 GB 640 GB/s 113 249 offload, 11.4 no
AMD Radeon RX 9070 XT 16 GB 640 GB/s 113 249 offload, 11.4 no
AMD Radeon RX 7700 16 GB 624 GB/s 110 242 offload, 11.3 no
AMD Radeon RX 7800 XT 16 GB 624 GB/s 110 242 offload, 11.3 no
AMD Radeon RX 7900 GRE 16 GB 576 GB/s 102 224 offload, 11 no
NVIDIA GeForce RTX 5060 Ti (16 GB) 16 GB 448 GB/s 79 174 offload, 10.1 no
NVIDIA RTX A4000 16 GB 448 GB/s 79 174 offload, 10.1 no
AMD Radeon RX 9060 XT (16GB) 16 GB 320 GB/s 56 124 offload, 8.8 no
AMD Radeon RX 9060 XT LP 16 GB 320 GB/s 56 124 offload, 8.8 no
NVIDIA GeForce RTX 4060 Ti (16 GB) 16 GB 288 GB/s 51 112 offload, 8.4 no
AMD Radeon RX 7600 XT 16 GB 288 GB/s 51 112 offload, 8.4 no
Intel Arc Pro B50 Graphics 16 GB 224 GB/s 40 87 offload, 7.4 no

What the arithmetic says: every card here runs the same models; bandwidth separates them by 4.3 times, from the NVIDIA GeForce RTX 5080 (960 GB/s) down to the Intel Arc Pro B50 Graphics (224 GB/s). Several 16 GB cards also come in an 8 GB version (RTX 4060 Ti, RTX 5060 Ti, RX 9060 XT); check the memory size, not just the name.

High-end tier: 20 to 32 GB

24 GB is where 30B-class dense models fit: Qwen3 32B at Q4_K_M takes 21.67 GiB at 8k and Gemma 4 31B 20.23 GiB. 32 GB adds headroom for longer context or Q6_K. A 70B model at Q4_K_M (43.7 GiB at 8k) does not fit on any single card in this tier.

Card Memory Bandwidth Llama 3.1 8B Q4_K_M gpt-oss-20b MXFP4 Qwen3 32B Q4_K_M Llama 3.3 70B Q4_K_M
NVIDIA GeForce RTX 5090 32 GB 1792 GB/s 316 696 82 offload, 6.4
NVIDIA GeForce RTX 3090 Ti 24 GB 1008 GB/s 178 392 46 offload, 3.9
NVIDIA GeForce RTX 4090 24 GB 1008 GB/s 178 392 46 offload, 3.9
AMD Radeon RX 7900 XTX 24 GB 960 GB/s 170 373 44 offload, 3.9
NVIDIA GeForce RTX 3090 24 GB 936 GB/s 165 364 43 offload, 3.9
NVIDIA RTX PRO 4500 Blackwell Workstation Edition 32 GB 896 GB/s 158 348 41 offload, 5.8
AMD Radeon RX 7900 XT 20 GB 800 GB/s 141 311 offload, 24.8 offload, 3.3
NVIDIA RTX A5000 24 GB 768 GB/s 136 298 35 offload, 3.8
NVIDIA RTX PRO 4000 Blackwell 24 GB 672 GB/s 119 261 31 offload, 3.8
AMD Radeon AI PRO R9600 32 GB 640 GB/s 113 249 30 offload, 5.3
AMD Radeon AI PRO R9600D 32 GB 640 GB/s 113 249 30 offload, 5.3
AMD Radeon AI PRO R9700 32 GB 640 GB/s 113 249 30 offload, 5.3
AMD Radeon AI PRO R9700S 32 GB 640 GB/s 113 249 30 offload, 5.3
Intel Arc Pro B65 Graphics 32 GB 608 GB/s 107 236 28 offload, 5.2
Intel Arc Pro B70 Graphics 32 GB 608 GB/s 107 236 28 offload, 5.2
NVIDIA RTX 5000 Ada Generation 32 GB 576 GB/s 102 224 26 offload, 5.2
AMD Radeon PRO W7800 32 GB 576 GB/s 102 224 26 offload, 5.2
Intel Arc Pro B60 Graphics 24 GB 456 GB/s 80 177 21 offload, 3.5
NVIDIA RTX 4500 Ada Generation 24 GB 432 GB/s 76 168 20 offload, 3.5
NVIDIA RTX 4000 Ada Generation 20 GB 360 GB/s 64 140 offload, 14 offload, 3

What the arithmetic says: the NVIDIA GeForce RTX 5090 has 32 GB and the most bandwidth in this tier (1792 GB/s), which puts its ceilings well above the rest. Among 24 GB cards, the RTX 3090, RTX 3090 Ti, RTX 4090 and RX 7900 XTX are within 8% of each other on bandwidth, so the ceilings are close; the differences between them are software, power draw and price. Several 32 GB workstation cards (AMD Radeon AI PRO, Intel Arc Pro B65 and B70) trade bandwidth for memory.

Two cards instead of one

Two 24 GB cards hold a 70B model at Q4_K_M entirely in VRAM: 44.7 GiB in 48 GiB. The ceiling is then 20.7 tokens per second on two RTX 3090s and 22.3 on two RTX 4090s, against 3.9 with one card and the rest in system RAM. A second card adds memory, not bandwidth: llama.cpp's default layer split passes each token through the cards one after the other. Two cards also need a motherboard with two suitable slots, a power supply for both, and room for the heat.

Workstation tier: 48 to 96 GB

These are professional workstation cards. 48 GB holds a 70B model at Q4_K_M on one card. 96 GB holds gpt-oss-120b in its published MXFP4 form (62.05 GiB at 8k) with room for context.

Card Memory Bandwidth Llama 3.1 8B Q4_K_M gpt-oss-20b MXFP4 Qwen3 32B Q4_K_M Llama 3.3 70B Q4_K_M
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 96 GB 1792 GB/s 316 696 82 40
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB 1792 GB/s 316 696 82 40
NVIDIA RTX PRO 5000 72GB Blackwell 72 GB 1344 GB/s 237 522 62 30
NVIDIA RTX PRO 5000 Blackwell (48 GB) 48 GB 1344 GB/s 237 522 62 30
NVIDIA RTX 6000 Ada Generation 48 GB 960 GB/s 170 373 44 21
AMD Radeon PRO W7800 48GB 48 GB 864 GB/s 152 336 40 19
AMD Radeon PRO W7900 48 GB 864 GB/s 152 336 40 19
NVIDIA RTX A6000 48 GB 768 GB/s 136 298 35 17

Unified memory instead of a graphics card

A Mac, an AMD Ryzen AI Max mini-PC or an NVIDIA DGX Spark shares one large pool of memory between CPU and GPU. They hold models no single consumer card can, at lower bandwidth than high-end cards. Each row is the largest memory configuration the vendor offers, with the GPU's default share of it:

Computer Memory (GPU share) Bandwidth Llama 3.1 8B Q4_K_M Llama 3.3 70B Q4_K_M gpt-oss-120b MXFP4 Qwen3 235B-A22B Q4_K_M
NVIDIA DGX Spark 128 GB (97%) 273 GB/s 48 6 86 no
AMD Ryzen AI Max+ 395 (Radeon 8060S) 128 GB (75%) 256 GB/s 45 6 81 no
AMD Ryzen AI Max+ PRO 495 (Radeon 8065S) 192 GB (83%) 273 GB/s 48 6 87 18
Apple M1 Max 64 GB (75%) 400 GB/s 71 9 no no
Apple M1 Ultra 128 GB (75%) 800 GB/s 141 18 254 no
Apple M3 Max (14-core CPU, 30-core GPU) 96 GB (75%) 300 GB/s 53 7 95 no
Apple M2 Max 96 GB (75%) 400 GB/s 71 9 127 no
Apple M3 Max (16-core CPU, 40-core GPU) 128 GB (75%) 400 GB/s 71 9 127 no
Apple M2 Ultra 192 GB (75%) 800 GB/s 141 18 254 53
Apple M4 Pro 64 GB (75%) 273 GB/s 48 6 no no
Apple M4 Max (16-core CPU, 40-core GPU) 128 GB (75%) 546 GB/s 96 12 173 no
Apple M3 Ultra 512 GB (75%) 819 GB/s 145 18 260 55
Apple M5 Pro 64 GB (75%) 307 GB/s 54 7 no no
Apple M5 Max (18-core CPU, 40-core GPU) 128 GB (75%) 614 GB/s 108 14 195 no
Apple M5 Ultra 512 GB (75%) 1200 GB/s 212 26 380 80

The Apple Silicon vs NVIDIA guide goes through the trade-off: capacity against bandwidth, and what each costs to run.

Where to buy

GPUCostLab does not list prices. The manufacturer pages linked from the hardware table give each card's full specifications; for current prices, check retailers such as Amazon and confirm the memory size on the listing. Cards from board partners use the same GPU and memory size; their clocks, power limits and coolers can differ.

What this guide does not tell you

Also available as Markdown.