# Best GPU for running LLMs at home, at each budget tier

> Graphics cards and unified-memory computers for running language models at home, grouped into budget tiers by memory size and ranked by memory bandwidth, with what each tier can run and its speed ceiling. Derived from vendor specifications and the calculator's arithmetic, not from benchmarks we have not run, and with no retailer prices.

By GPUCostLab. Updated 7 October 2026. Canonical URL: https://gpucostlab.com/best-gpu-for-local-llms/

No paid links on this page: vendor links go straight to the vendor.

## How this list is made

Two numbers decide what a GPU can do with a language model at home:

1. **Memory decides what runs.** The model's weights, its KV cache and the software's buffers have to fit. If they do not, the model either does not load or runs partly from system RAM, which is much slower. The [VRAM guide](/how-much-vram-to-run-llms-locally/) has the figure for each model.
2. **Memory bandwidth decides how fast it can generate.** Every generated token reads all the active weights once, so bandwidth divided by bytes per token is the most tokens per second any software can reach. Real engines stay below this ceiling.

So the tiers below are memory sizes, and within each tier the cards are ranked by bandwidth. Every speed is the ceiling from the [Can I run it?](/can-i-run-this-llm/) calculator for one conversation at an 8,192-token context, in tokens per second: an upper bound, not a benchmark. "Offload, n" means the model only runs with part of it in system RAM (assumed 32 GB of DDR5-5600), and n is the ceiling then. "No" means it does not fit even with RAM.

**There are no prices here**, because they change weekly and differ between sellers. Within a generation, more memory and more bandwidth cost more, so the tiers run from the least to the most expensive kind of card; check current prices yourself before choosing between tiers.

**Software matters as much as the chip.** NVIDIA cards use CUDA, which llama.cpp, Ollama, LM Studio, vLLM and nearly every other engine support. AMD cards run llama.cpp through ROCm on the cards AMD lists for it, or through Vulkan, which works on any card with a Vulkan driver; AMD's ROCm 10.1 compatibility list covers the current Radeon RX 7000 and 9000 cards except the RX 7600 XT. Intel Arc cards use llama.cpp's SYCL backend (the Arc B580 is on its verified list) or Vulkan. Engines beyond llama.cpp, such as vLLM, support NVIDIA best. The software notes for every card are in the [calculator's hardware table](/can-i-run-this-llm/#cir-hw-title).

## Entry tier: 8 to 12 GB

An 8 GB card holds 8-billion-parameter models at Q4_K_M with room for a long context (6.58 GiB for Llama 3.1 8B at 8k), and little more. A 12 GB card also holds Qwen3 14B at Q4_K_M (10.66 GiB at 8k). gpt-oss-20b needs 14.01 GiB, so on these cards part of it runs from system RAM; as a mixture-of-experts model it stays usable that way.

| Card | Memory | Bandwidth | Llama 3.1 8B Q4_K_M | gpt-oss-20b MXFP4 | Qwen3 32B Q4_K_M | Llama 3.3 70B Q4_K_M |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 3080 Ti | 12 GB | 912 GB/s | 161 | offload, 187.3 | offload, 8.1 | no |
| NVIDIA GeForce RTX 3080 (10 GB) | 10 GB | 760 GB/s | 134 | offload, 106.8 | offload, 6.7 | no |
| NVIDIA GeForce RTX 5070 | 12 GB | 672 GB/s | 119 | offload, 162.3 | offload, 7.8 | no |
| NVIDIA GeForce RTX 4070 | 12 GB | 504 GB/s | 89 | offload, 138.8 | offload, 7.4 | no |
| NVIDIA GeForce RTX 4070 SUPER | 12 GB | 504 GB/s | 89 | offload, 138.8 | offload, 7.4 | no |
| NVIDIA GeForce RTX 4070 Ti | 12 GB | 504 GB/s | 89 | offload, 138.8 | offload, 7.4 | no |
| Intel Arc B580 Graphics | 12 GB | 456 GB/s | 80 | offload, 130.8 | offload, 7.3 | no |
| NVIDIA GeForce RTX 3060 Ti | 8 GB | 448 GB/s | 79 | offload, 70.4 | offload, 5.7 | no |
| NVIDIA GeForce RTX 3070 | 8 GB | 448 GB/s | 79 | offload, 70.4 | offload, 5.7 | no |
| NVIDIA GeForce RTX 5060 | 8 GB | 448 GB/s | 79 | offload, 70.4 | offload, 5.7 | no |
| NVIDIA GeForce RTX 5060 Ti (8 GB) | 8 GB | 448 GB/s | 79 | offload, 70.4 | offload, 5.7 | no |
| AMD Radeon RX 7700 XT | 12 GB | 432 GB/s | 76 | offload, 126.6 | offload, 7.2 | no |
| AMD Radeon RX 9070 GRE | 12 GB | 432 GB/s | 76 | offload, 126.6 | offload, 7.2 | no |
| Intel Arc B570 Graphics | 10 GB | 380 GB/s | 67 | offload, 85.8 | offload, 6.2 | no |
| NVIDIA GeForce RTX 3060 (12 GB) | 12 GB | 360 GB/s | 64 | offload, 112.7 | offload, 7 | no |
| NVIDIA GeForce RTX 5050 | 8 GB | 320 GB/s | 56 | offload, 64.8 | offload, 5.5 | no |
| AMD Radeon RX 9060 XT (8GB) | 8 GB | 320 GB/s | 56 | offload, 64.8 | offload, 5.5 | no |
| NVIDIA GeForce RTX 4060 Ti (8 GB) | 8 GB | 288 GB/s | 51 | offload, 62.9 | offload, 5.4 | no |
| AMD Radeon RX 7600 | 8 GB | 288 GB/s | 51 | offload, 62.9 | offload, 5.4 | no |
| AMD Radeon RX 9050 | 8 GB | 288 GB/s | 51 | offload, 62.9 | offload, 5.4 | no |
| AMD Radeon RX 9060 | 8 GB | 288 GB/s | 51 | offload, 62.9 | offload, 5.4 | no |
| NVIDIA GeForce RTX 4060 | 8 GB | 272 GB/s | 48 | offload, 61.8 | offload, 5.4 | no |

**What the arithmetic says:** if you can, choose 12 GB over 8 GB, because it moves you from 8B to 14B models. Among the 12 GB cards, the NVIDIA GeForce RTX 3080 Ti has the highest bandwidth in this list.

## Mainstream tier: 16 GB

16 GB holds gpt-oss-20b in its published MXFP4 form entirely in VRAM, Qwen3 14B at up to Q6_K, and Mistral Small 3.2 24B Instruct (2506) at Q4_K_M with an 8k context (15.93 GiB, with very little to spare). 30B-class dense models do not fit at Q4_K_M.

| Card | Memory | Bandwidth | Llama 3.1 8B Q4_K_M | gpt-oss-20b MXFP4 | Qwen3 32B Q4_K_M | Llama 3.3 70B Q4_K_M |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5080 | 16 GB | 960 GB/s | 170 | 373 | offload, 12.5 | no |
| NVIDIA GeForce RTX 5070 Ti | 16 GB | 896 GB/s | 158 | 348 | offload, 12.4 | no |
| NVIDIA GeForce RTX 4080 SUPER | 16 GB | 736 GB/s | 130 | 286 | offload, 11.8 | no |
| NVIDIA GeForce RTX 4080 | 16 GB | 716.8 GB/s | 126 | 279 | offload, 11.7 | no |
| NVIDIA GeForce RTX 4070 Ti SUPER | 16 GB | 672 GB/s | 119 | 261 | offload, 11.5 | no |
| AMD Radeon RX 9070 | 16 GB | 640 GB/s | 113 | 249 | offload, 11.4 | no |
| AMD Radeon RX 9070 XT | 16 GB | 640 GB/s | 113 | 249 | offload, 11.4 | no |
| AMD Radeon RX 7700 | 16 GB | 624 GB/s | 110 | 242 | offload, 11.3 | no |
| AMD Radeon RX 7800 XT | 16 GB | 624 GB/s | 110 | 242 | offload, 11.3 | no |
| AMD Radeon RX 7900 GRE | 16 GB | 576 GB/s | 102 | 224 | offload, 11 | no |
| NVIDIA GeForce RTX 5060 Ti (16 GB) | 16 GB | 448 GB/s | 79 | 174 | offload, 10.1 | no |
| NVIDIA RTX A4000 | 16 GB | 448 GB/s | 79 | 174 | offload, 10.1 | no |
| AMD Radeon RX 9060 XT (16GB) | 16 GB | 320 GB/s | 56 | 124 | offload, 8.8 | no |
| AMD Radeon RX 9060 XT LP | 16 GB | 320 GB/s | 56 | 124 | offload, 8.8 | no |
| NVIDIA GeForce RTX 4060 Ti (16 GB) | 16 GB | 288 GB/s | 51 | 112 | offload, 8.4 | no |
| AMD Radeon RX 7600 XT | 16 GB | 288 GB/s | 51 | 112 | offload, 8.4 | no |
| Intel Arc Pro B50 Graphics | 16 GB | 224 GB/s | 40 | 87 | offload, 7.4 | no |

**What the arithmetic says:** every card here runs the same models; bandwidth separates them by 4.3 times, from the NVIDIA GeForce RTX 5080 (960 GB/s) down to the Intel Arc Pro B50 Graphics (224 GB/s). Several 16 GB cards also come in an 8 GB version (RTX 4060 Ti, RTX 5060 Ti, RX 9060 XT); check the memory size, not just the name.

## High-end tier: 20 to 32 GB

24 GB is where 30B-class dense models fit: Qwen3 32B at Q4_K_M takes 21.67 GiB at 8k and Gemma 4 31B 20.23 GiB. 32 GB adds headroom for longer context or Q6_K. A 70B model at Q4_K_M (43.7 GiB at 8k) does not fit on any single card in this tier.

| Card | Memory | Bandwidth | Llama 3.1 8B Q4_K_M | gpt-oss-20b MXFP4 | Qwen3 32B Q4_K_M | Llama 3.3 70B Q4_K_M |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 | 32 GB | 1792 GB/s | 316 | 696 | 82 | offload, 6.4 |
| NVIDIA GeForce RTX 3090 Ti | 24 GB | 1008 GB/s | 178 | 392 | 46 | offload, 3.9 |
| NVIDIA GeForce RTX 4090 | 24 GB | 1008 GB/s | 178 | 392 | 46 | offload, 3.9 |
| AMD Radeon RX 7900 XTX | 24 GB | 960 GB/s | 170 | 373 | 44 | offload, 3.9 |
| NVIDIA GeForce RTX 3090 | 24 GB | 936 GB/s | 165 | 364 | 43 | offload, 3.9 |
| NVIDIA RTX PRO 4500 Blackwell Workstation Edition | 32 GB | 896 GB/s | 158 | 348 | 41 | offload, 5.8 |
| AMD Radeon RX 7900 XT | 20 GB | 800 GB/s | 141 | 311 | offload, 24.8 | offload, 3.3 |
| NVIDIA RTX A5000 | 24 GB | 768 GB/s | 136 | 298 | 35 | offload, 3.8 |
| NVIDIA RTX PRO 4000 Blackwell | 24 GB | 672 GB/s | 119 | 261 | 31 | offload, 3.8 |
| AMD Radeon AI PRO R9600 | 32 GB | 640 GB/s | 113 | 249 | 30 | offload, 5.3 |
| AMD Radeon AI PRO R9600D | 32 GB | 640 GB/s | 113 | 249 | 30 | offload, 5.3 |
| AMD Radeon AI PRO R9700 | 32 GB | 640 GB/s | 113 | 249 | 30 | offload, 5.3 |
| AMD Radeon AI PRO R9700S | 32 GB | 640 GB/s | 113 | 249 | 30 | offload, 5.3 |
| Intel Arc Pro B65 Graphics | 32 GB | 608 GB/s | 107 | 236 | 28 | offload, 5.2 |
| Intel Arc Pro B70 Graphics | 32 GB | 608 GB/s | 107 | 236 | 28 | offload, 5.2 |
| NVIDIA RTX 5000 Ada Generation | 32 GB | 576 GB/s | 102 | 224 | 26 | offload, 5.2 |
| AMD Radeon PRO W7800 | 32 GB | 576 GB/s | 102 | 224 | 26 | offload, 5.2 |
| Intel Arc Pro B60 Graphics | 24 GB | 456 GB/s | 80 | 177 | 21 | offload, 3.5 |
| NVIDIA RTX 4500 Ada Generation | 24 GB | 432 GB/s | 76 | 168 | 20 | offload, 3.5 |
| NVIDIA RTX 4000 Ada Generation | 20 GB | 360 GB/s | 64 | 140 | offload, 14 | offload, 3 |

**What the arithmetic says:** the NVIDIA GeForce RTX 5090 has 32 GB and the most bandwidth in this tier (1792 GB/s), which puts its ceilings well above the rest. Among 24 GB cards, the RTX 3090, RTX 3090 Ti, RTX 4090 and RX 7900 XTX are within 8% of each other on bandwidth, so the ceilings are close; the differences between them are software, power draw and price. Several 32 GB workstation cards (AMD Radeon AI PRO, Intel Arc Pro B65 and B70) trade bandwidth for memory.

## Two cards instead of one

Two 24 GB cards hold a 70B model at Q4_K_M entirely in VRAM: 44.7 GiB in 48 GiB. The ceiling is then 20.7 tokens per second on two RTX 3090s and 22.3 on two RTX 4090s, against 3.9 with one card and the rest in system RAM. A second card adds memory, not bandwidth: llama.cpp's default layer split passes each token through the cards one after the other. Two cards also need a motherboard with two suitable slots, a power supply for both, and room for the heat.

## Workstation tier: 48 to 96 GB

These are professional workstation cards. 48 GB holds a 70B model at Q4_K_M on one card. 96 GB holds gpt-oss-120b in its published MXFP4 form (62.05 GiB at 8k) with room for context.

| Card | Memory | Bandwidth | Llama 3.1 8B Q4_K_M | gpt-oss-20b MXFP4 | Qwen3 32B Q4_K_M | Llama 3.3 70B Q4_K_M |
|---|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 96 GB | 1792 GB/s | 316 | 696 | 82 | 40 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 96 GB | 1792 GB/s | 316 | 696 | 82 | 40 |
| NVIDIA RTX PRO 5000 72GB Blackwell | 72 GB | 1344 GB/s | 237 | 522 | 62 | 30 |
| NVIDIA RTX PRO 5000 Blackwell (48 GB) | 48 GB | 1344 GB/s | 237 | 522 | 62 | 30 |
| NVIDIA RTX 6000 Ada Generation | 48 GB | 960 GB/s | 170 | 373 | 44 | 21 |
| AMD Radeon PRO W7800 48GB | 48 GB | 864 GB/s | 152 | 336 | 40 | 19 |
| AMD Radeon PRO W7900 | 48 GB | 864 GB/s | 152 | 336 | 40 | 19 |
| NVIDIA RTX A6000 | 48 GB | 768 GB/s | 136 | 298 | 35 | 17 |

## Unified memory instead of a graphics card

A Mac, an AMD Ryzen AI Max mini-PC or an NVIDIA DGX Spark shares one large pool of memory between CPU and GPU. They hold models no single consumer card can, at lower bandwidth than high-end cards. Each row is the largest memory configuration the vendor offers, with the GPU's default share of it:

| Computer | Memory (GPU share) | Bandwidth | Llama 3.1 8B Q4_K_M | Llama 3.3 70B Q4_K_M | gpt-oss-120b MXFP4 | Qwen3 235B-A22B Q4_K_M |
|---|---|---|---|---|---|---|
| NVIDIA DGX Spark | 128 GB (97%) | 273 GB/s | 48 | 6 | 86 | no |
| AMD Ryzen AI Max+ 395 (Radeon 8060S) | 128 GB (75%) | 256 GB/s | 45 | 6 | 81 | no |
| AMD Ryzen AI Max+ PRO 495 (Radeon 8065S) | 192 GB (83%) | 273 GB/s | 48 | 6 | 87 | 18 |
| Apple M1 Max | 64 GB (75%) | 400 GB/s | 71 | 9 | no | no |
| Apple M1 Ultra | 128 GB (75%) | 800 GB/s | 141 | 18 | 254 | no |
| Apple M3 Max (14-core CPU, 30-core GPU) | 96 GB (75%) | 300 GB/s | 53 | 7 | 95 | no |
| Apple M2 Max | 96 GB (75%) | 400 GB/s | 71 | 9 | 127 | no |
| Apple M3 Max (16-core CPU, 40-core GPU) | 128 GB (75%) | 400 GB/s | 71 | 9 | 127 | no |
| Apple M2 Ultra | 192 GB (75%) | 800 GB/s | 141 | 18 | 254 | 53 |
| Apple M4 Pro | 64 GB (75%) | 273 GB/s | 48 | 6 | no | no |
| Apple M4 Max (16-core CPU, 40-core GPU) | 128 GB (75%) | 546 GB/s | 96 | 12 | 173 | no |
| Apple M3 Ultra | 512 GB (75%) | 819 GB/s | 145 | 18 | 260 | 55 |
| Apple M5 Pro | 64 GB (75%) | 307 GB/s | 54 | 7 | no | no |
| Apple M5 Max (18-core CPU, 40-core GPU) | 128 GB (75%) | 614 GB/s | 108 | 14 | 195 | no |
| Apple M5 Ultra | 512 GB (75%) | 1200 GB/s | 212 | 26 | 380 | 80 |

The [Apple Silicon vs NVIDIA guide](/apple-silicon-vs-nvidia-for-local-ai/) goes through the trade-off: capacity against bandwidth, and what each costs to run.

## Where to buy

GPUCostLab does not list prices. The manufacturer pages linked from the [hardware table](/can-i-run-this-llm/#cir-hw-title) give each card's full specifications; for current prices, check retailers such as [Amazon](https://www.amazon.com/) and confirm the memory size on the listing. Cards from board partners use the same GPU and memory size; their clocks, power limits and coolers can differ.

## What this guide does not tell you

- **Measured speed.** Every speed here is a bandwidth ceiling. Real speed depends on the engine, its version, the quantization kernels and the driver, and is lower. Measured results are [planned](/methodology/#benchmarks-planned-no-results-yet) and will be published with their scripts.
- **Prompt processing.** Reading a long prompt is limited by compute, not bandwidth, and is not ranked here.
- **Quality.** A smaller quantization fits more but loses some accuracy, by an amount that depends on the model.
- **Noise, heat and power.** The [electricity calculator](/local-llm-electricity-cost/) turns each card's rated power into a monthly cost.
