Can my computer run this model?

Runs in your browser; nothing you enter is sent anywhere. Hardware specs from vendor pages, read 7 October 2026; model shapes from each model's config.json. Speeds are upper bounds from memory bandwidth, not measurements.

Your hardware
Model
Assumptions you can change

The runtime allowance (1 GiB per GPU by default) stands for llama.cpp's compute buffers and the GPU driver's own context. It is a round assumption, not a measurement: llama.cpp prints its buffers when it loads a model ("compute buffer size = ... MiB"), and a GPU that also drives your monitors uses some VRAM for the desktop, so raise it to match. The RAM kept for the operating system and other programs (4 GiB by default) is also an assumption; a desktop with a browser open often uses more, so check your task manager or Activity Monitor.

Hardware in the calculator

Memory, memory bandwidth and rated power from each vendor's own spec page, datasheet or whitepaper, read 7 October 2026. "Derived" bandwidth is the vendor's stated memory speed times bus width; "n/a" means the vendor does not publish it. Raw data: consumer-hardware.json.

NVIDIA GeForce

DeviceMemoryBandwidthRated powerSoftwareSource
NVIDIA GeForce RTX 4060 8 GB GDDR6 272 GB/s 115 W CUDA NVIDIA
NVIDIA GeForce RTX 4060 Ti (8 GB) 8 GB GDDR6 288 GB/s 160 W CUDA NVIDIA
NVIDIA GeForce RTX 5050 8 GB GDDR6 320 GB/s 130 W CUDA NVIDIA
NVIDIA GeForce RTX 3060 Ti 8 GB GDDR6 448 GB/s 200 W CUDA NVIDIA
NVIDIA GeForce RTX 3070 8 GB GDDR6 448 GB/s 220 W CUDA NVIDIA
NVIDIA GeForce RTX 5060 8 GB GDDR7 448 GB/s 145 W CUDA NVIDIA
NVIDIA GeForce RTX 5060 Ti (8 GB) 8 GB GDDR7 448 GB/s 180 W CUDA NVIDIA
NVIDIA GeForce RTX 3080 (10 GB) 10 GB GDDR6X 760 GB/s 320 W CUDA NVIDIA
NVIDIA GeForce RTX 3080 (12 GB) 12 GB GDDR6X n/a 350 W CUDA NVIDIA
NVIDIA GeForce RTX 3060 (12 GB) 12 GB GDDR6 360 GB/s 170 W CUDA NVIDIA
NVIDIA GeForce RTX 4070 12 GB GDDR6X 504 GB/s 200 W CUDA NVIDIA
NVIDIA GeForce RTX 4070 SUPER 12 GB GDDR6X 504 GB/s 220 W CUDA NVIDIA
NVIDIA GeForce RTX 4070 Ti 12 GB GDDR6X 504 GB/s 285 W CUDA NVIDIA
NVIDIA GeForce RTX 5070 12 GB GDDR7 672 GB/s 250 W CUDA NVIDIA
NVIDIA GeForce RTX 3080 Ti 12 GB GDDR6X 912 GB/s 350 W CUDA NVIDIA
NVIDIA GeForce RTX 4060 Ti (16 GB) 16 GB GDDR6 288 GB/s 165 W CUDA NVIDIA
NVIDIA GeForce RTX 5060 Ti (16 GB) 16 GB GDDR7 448 GB/s 180 W CUDA NVIDIA
NVIDIA GeForce RTX 4070 Ti SUPER 16 GB GDDR6X 672 GB/s 285 W CUDA NVIDIA
NVIDIA GeForce RTX 4080 16 GB GDDR6X 716.8 GB/s 320 W CUDA NVIDIA
NVIDIA GeForce RTX 4080 SUPER 16 GB GDDR6X 736 GB/s 320 W CUDA NVIDIA
NVIDIA GeForce RTX 5070 Ti 16 GB GDDR7 896 GB/s 300 W CUDA NVIDIA
NVIDIA GeForce RTX 5080 16 GB GDDR7 960 GB/s 360 W CUDA NVIDIA
NVIDIA GeForce RTX 3090 24 GB GDDR6X 936 GB/s 350 W CUDA NVIDIA
NVIDIA GeForce RTX 3090 Ti 24 GB GDDR6X 1,008 GB/s 450 W CUDA NVIDIA
NVIDIA GeForce RTX 4090 24 GB GDDR6X 1,008 GB/s 450 W CUDA NVIDIA
NVIDIA GeForce RTX 5090 32 GB GDDR7 1,792 GB/s 575 W CUDA NVIDIA

NVIDIA RTX workstation cards

DeviceMemoryBandwidthRated powerSoftwareSource
NVIDIA RTX A4000 16 GB GDDR6 (ECC) 448 GB/s 140 W CUDA NVIDIA
NVIDIA RTX 4000 Ada Generation 20 GB GDDR6 (ECC) 360 GB/s 130 W CUDA NVIDIA
NVIDIA RTX 4500 Ada Generation 24 GB GDDR6 (ECC) 432 GB/s 210 W CUDA NVIDIA
NVIDIA RTX PRO 4000 Blackwell 24 GB GDDR7 (ECC) 672 GB/s 145 W CUDA NVIDIA
NVIDIA RTX A5000 24 GB GDDR6 (ECC) 768 GB/s 230 W CUDA NVIDIA
NVIDIA RTX 5000 Ada Generation 32 GB GDDR6 (ECC) 576 GB/s 250 W CUDA NVIDIA
NVIDIA RTX PRO 4500 Blackwell Workstation Edition 32 GB GDDR7 (ECC) 896 GB/s 200 W CUDA NVIDIA
NVIDIA RTX A6000 48 GB GDDR6 (ECC) 768 GB/s 300 W CUDA NVIDIA
NVIDIA RTX 6000 Ada Generation 48 GB GDDR6 (ECC) 960 GB/s 300 W CUDA NVIDIA
NVIDIA RTX PRO 5000 Blackwell (48 GB) 48 GB GDDR7 (ECC) 1,344 GB/s 300 W CUDA NVIDIA
NVIDIA RTX PRO 5000 72GB Blackwell 72 GB GDDR7 (ECC) 1,344 GB/s 300 W CUDA NVIDIA
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 96 GB GDDR7 (ECC) 1,792 GB/s 300 W CUDA NVIDIA
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB GDDR7 (ECC) 1,792 GB/s 600 W CUDA NVIDIA

AMD Radeon

DeviceMemoryBandwidthRated powerSoftwareSource
AMD Radeon RX 7600 8 GB GDDR6 288 GB/s 165 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 9050 8 GB GDDR6 288 GB/s 92 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 9060 8 GB GDDR6 288 GB/s 132 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 9060 XT (8GB) 8 GB GDDR6 320 GB/s 150 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 7700 XT 12 GB GDDR6 432 GB/s 245 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 9070 GRE 12 GB GDDR6 432 GB/s 220 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 7600 XT 16 GB GDDR6 288 GB/s 190 W Vulkan; not on AMD's ROCm list AMD
AMD Radeon RX 9060 XT (16GB) 16 GB GDDR6 320 GB/s 160 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 9060 XT LP 16 GB GDDR6 320 GB/s 140 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 7900 GRE 16 GB GDDR6 576 GB/s 260 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 7700 16 GB GDDR6 624 GB/s 263 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 7800 XT 16 GB GDDR6 624 GB/s 263 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 9070 16 GB GDDR6 640 GB/s 220 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 9070 XT 16 GB GDDR6 640 GB/s 304 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 7900 XT 20 GB GDDR6 800 GB/s 315 W ROCm (Linux) and Vulkan AMD
AMD Radeon RX 7900 XTX 24 GB GDDR6 960 GB/s 355 W ROCm (Linux) and Vulkan AMD

AMD Radeon PRO and AI PRO

DeviceMemoryBandwidthRated powerSoftwareSource
AMD Radeon PRO W7800 32 GB GDDR6 576 GB/s 260 W ROCm (Linux) and Vulkan AMD
AMD Radeon AI PRO R9600 32 GB GDDR6 640 GB/s 150 W ROCm (Linux) and Vulkan AMD
AMD Radeon AI PRO R9600D 32 GB GDDR6 640 GB/s 150 W ROCm (Linux) and Vulkan AMD
AMD Radeon AI PRO R9700 32 GB GDDR6 640 GB/s 300 W ROCm (Linux) and Vulkan AMD
AMD Radeon AI PRO R9700S 32 GB GDDR6 640 GB/s 300 W ROCm (Linux) and Vulkan AMD
AMD Radeon PRO W7800 48GB 48 GB GDDR6 864 GB/s 260 W ROCm (Linux) and Vulkan AMD
AMD Radeon PRO W7900 48 GB GDDR6 864 GB/s 295 W ROCm (Linux) and Vulkan AMD

Intel Arc

DeviceMemoryBandwidthRated powerSoftwareSource
Intel Arc B570 Graphics 10 GB GDDR6 380 GB/s 150 W SYCL and Vulkan Intel
Intel Arc B580 Graphics 12 GB GDDR6 456 GB/s 190 W SYCL and Vulkan Intel
Intel Arc Pro B50 Graphics 16 GB GDDR6 224 GB/s 70 W SYCL and Vulkan Intel
Intel Arc Pro B60 Graphics 24 GB GDDR6 456 GB/s 200 W SYCL and Vulkan Intel
Intel Arc Pro B65 Graphics 32 GB GDDR6 608 GB/s 200 W SYCL and Vulkan Intel
Intel Arc Pro B70 Graphics 32 GB GDDR6 608 GB/s 230 W SYCL and Vulkan Intel

Apple Silicon (unified memory)

DeviceMemoryBandwidthRated powerSoftwareSource
Apple M1 Pro 16 / 32 GB unified 200 GB/s n/a Metal Apple
Apple M1 Max 32 / 64 GB unified 400 GB/s 115 W Metal Apple
Apple M1 Ultra 64 / 128 GB unified 800 GB/s 215 W Metal Apple
Apple M2 Pro 16 / 32 GB unified 200 GB/s 100 W Metal Apple
Apple M3 Pro 18 / 36 GB unified 150 GB/s n/a Metal Apple
Apple M3 Max (14-core CPU, 30-core GPU) 36 / 96 GB unified 300 GB/s n/a Metal Apple
Apple M2 Max 32 / 64 / 96 GB unified 400 GB/s 145 W Metal Apple
Apple M3 Max (16-core CPU, 40-core GPU) 48 / 64 / 128 GB unified 400 GB/s n/a Metal Apple
Apple M2 Ultra 64 / 128 / 192 GB unified 800 GB/s 295 W Metal Apple
Apple M4 16 / 24 / 32 GB unified 120 GB/s 65 W Metal Apple
Apple M4 Max (14-core CPU, 32-core GPU) 36 GB unified 410 GB/s 145 W Metal Apple
Apple M4 Pro 24 / 48 / 64 GB unified 273 GB/s 140 W Metal Apple
Apple M4 Max (16-core CPU, 40-core GPU) 48 / 64 / 128 GB unified 546 GB/s n/a Metal Apple
Apple M5 16 / 24 / 32 GB unified 153 GB/s n/a Metal Apple
Apple M3 Ultra 96 / 256 / 512 GB unified 819 GB/s 270 W Metal Apple
Apple M6 16 / 24 / 32 GB unified 170 GB/s 70 W Metal Apple
Apple M5 Max (18-core CPU, 32-core GPU) 36 GB unified 460 GB/s 200 W Metal Apple
Apple M5 Pro 24 / 48 / 64 GB unified 307 GB/s 145 W Metal Apple
Apple M5 Max (18-core CPU, 40-core GPU) 48 / 64 / 128 GB unified 614 GB/s n/a Metal Apple
Apple M5 Ultra 96 / 256 / 512 GB unified 1,200 GB/s 385 W Metal Apple

Unified-memory PCs (AMD Ryzen AI Max, NVIDIA DGX Spark)

DeviceMemoryBandwidthRated powerSoftwareSource
NVIDIA RTX Spark (N1X) desktop PCs 128 GB unified n/a 140 W CUDA NVIDIA
AMD Ryzen AI Max 385 (Radeon 8050S) 128 GB unified 256 GB/s (derived) 120 W ROCm and Vulkan AMD
AMD Ryzen AI Max 390 (Radeon 8050S) 128 GB unified 256 GB/s (derived) 120 W ROCm and Vulkan AMD
AMD Ryzen AI Max+ 395 (Radeon 8060S) 128 GB unified 256 GB/s 120 W ROCm and Vulkan AMD
NVIDIA DGX Spark 64 / 128 GB unified 273 GB/s 240 W CUDA NVIDIA
AMD Ryzen AI Max+ PRO 495 (Radeon 8065S) 192 GB unified 273.1 GB/s (derived) 120 W ROCm and Vulkan AMD
How each device's figures were read (derived bandwidth, power terms, software support)
NVIDIA GeForce RTX 4060
NVIDIA RTX 4060 launch spec-comparison image (on https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4060-4060ti/): Memory Subsystem '24 MB L2 / 272 GB/s (453 GB/s effective)', TGP '115 W'. 272 GB/s is the physical DRAM bandwidth; 453 GB/s is NVIDIA's cache-adjusted 'effective' figure. Per-pin data rate not stated by NVIDIA. Power: Total Graphics Power. Compare page (40 Series): '8 GB GDDR6', 128-bit, Total Graphics Power (W) '115'. Product page: https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4060-4060ti/. GeForce news dated May 18, 2023 says RTX 4060 arrives in July (2023). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/geforc....
NVIDIA GeForce RTX 4060 Ti (8 GB)
NVIDIA RTX 4060 Ti launch spec-comparison image (on https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4060-4060ti/): VRAM '8GB / 16GB', Memory Subsystem '32 MB L2 / 288 GB/s (554 GB/s effective)', TGP '160 W'. VRAM explainer (https://www.nvidia.com/en-us/geforce/news/rtx-40-series-vram-video-memory-explained/) also: 'an Ada GPU with 288 GB/sec of peak memory bandwidth' in the RTX 4060 Ti test. Per-pin data rate not stated by NVIDIA. Power: Total Graphics Power. Compare page: '16 GB GDDR6 or 8 GB GDDR6', 128-bit, Total Graphics Power '165 or 160' (8 GB = 160 W by position). Available May 24, 2023 per GeForce news dated May 18, 2023. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/geforc....
NVIDIA GeForce RTX 5050
Compare page (50 Series) 'Memory Bandwidth' row: '320 GB/sec'. Per-pin data rate not stated by NVIDIA. Power: Total Graphics Power. Later 50-series desktop card (not in the original request list). Compare page and product page: '8 GB GDDR6', 128-bit, Total Graphics Power (W) '130'. GeForce news (https://www.nvidia.com/en-us/geforce/news/rtx-5050-desktop-gpu-and-laptops/, seen via search result) says desktop cards arrive in the second half of July (2025). Second source: https://www.nvidia.com/en-us/geforce/graphics-cards/50-se....
NVIDIA GeForce RTX 3060 Ti
NVIDIA RTX 4060 Ti launch spec-comparison image (on https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4060-4060ti/) lists RTX 3060 Ti: '8GB' VRAM, Memory Subsystem '4MB L2 / 448 GB/s', TGP '200W'. Applies to the original GDDR6 card; no official bandwidth found for the later GDDR6X variant. Per-pin data rate not stated by NVIDIA. Power: Graphics Card Power. Compare page: Standard Memory Config '8 GB GDDR6 / 8 GB GDDR6X', 256-bit, Graphics Card Power (W) '200'. Release year from GeForce news 'GeForce RTX 3060 Ti out December 2' dated December 01, 2020 (https://www.nvidia.com/en-us/geforce/news/geforce-rtx-3060-ti-out-december-2/). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/geforc....
NVIDIA GeForce RTX 3070
Ampere GA102 whitepaper Table 10 (p.47, 'RTX 3070 Founders Edition'): Memory Interface 256-bit, 'Memory Clock (Data Rate) 14 Gbps', 'Memory Bandwidth 448 GB/sec', TGP 220 Watts. Same values in RTX Blackwell whitepaper Table 6 (p.54-55). Power: Graphics Card Power. Compare page: 8 GB GDDR6, 256-bit, Graphics Card Power (W) '220'. Product page: https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3070-3070ti/. Release year: RTX 30 Series introduced in GeForce news dated September 01, 2020. Second source: https://www.nvidia.com/content/PDF/nvidia-ampere-ga-102-g....
NVIDIA GeForce RTX 5060
Compare page (50 Series) 'Memory Bandwidth' row: '448 GB/sec'. Per-pin data rate not stated on any NVIDIA page found. Power: Total Graphics Power. Compare page / RTX 5060 family page: '8 GB GDDR7', 128-bit, Total Graphics Power (W) '145'. NVIDIA newsroom (April 15, 2025): 'GeForce RTX 5060 graphics cards will be available starting in May at $299'. Second source: https://www.nvidia.com/en-us/geforce/graphics-cards/50-se....
NVIDIA GeForce RTX 5060 Ti (8 GB)
Compare page (50 Series): RTX 5060 Ti column 'Standard Memory Config 16 GB / 8 GB GDDR7', 'Memory Bandwidth 448 GB/sec' (one value for both memory sizes). Per-pin data rate not stated. Power: Total Graphics Power. Total Graphics Power '180' (single value listed for both memory sizes). NVIDIA newsroom (April 15, 2025, https://nvidianews.nvidia.com/news/nvidia-blackwell-geforce-rtx-arrives-for-every-gamer-starting-at-299): 'GeForce RTX 5060 Ti graphics cards, equipped with 16GB or 8GB graphics memory, will be available starting April 16'. Second source: https://www.nvidia.com/en-us/geforce/graphics-cards/50-se....
NVIDIA GeForce RTX 3080 (10 GB)
Ampere GA102 whitepaper Table 2 (p.14-15, 'GeForce RTX 3080 10 GB Founders Edition'): 320-bit, 'Memory Clock (Data Rate) 19 Gbps', 'Memory Bandwidth 760 GB/sec', 'TGP (Total Graphics Power) 320W'. Also RTX Blackwell whitepaper Table 4 (p.49-50). Power: Graphics Card Power. Compare page: '12 GB GDDR6X / 10 GB GDDR6X', '384-bit / 320-bit', Graphics Card Power '350 / 320' (10 GB = 320 W). Product page: https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3080-3080ti/. Release year: RTX 30 Series GeForce news dated September 01, 2020 ('on sale starting September 17th'). Second source: https://www.nvidia.com/content/PDF/nvidia-ampere-ga-102-g....
NVIDIA GeForce RTX 3080 (12 GB)
No NVIDIA page found that states either the memory bandwidth or the per-pin data rate for the 12 GB RTX 3080, so bandwidth cannot be stated or derived. Left null. Power: Graphics Card Power. Compare page and RTX 3080 family page: '12 GB GDDR6X / 10 GB GDDR6X', '384-bit / 320-bit', Graphics Card Power '350 / 320' (12 GB = 350 W). release_year 2022 not verified against a dated NVIDIA page. Second source: https://www.nvidia.com/en-us/geforce/graphics-cards/30-se....
NVIDIA GeForce RTX 3060 (12 GB)
NVIDIA RTX 4060 launch spec-comparison image (on https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4060-4060ti/) lists RTX 3060: '12GB' VRAM, Memory Subsystem '3MB L2 / 360 GB/s', TGP '170W'. NVIDIA does not state the per-pin data rate on any page found. Power: Graphics Card Power. Compare page (30 Series): Standard Memory Config '12 GB GDDR6 / 8 GB GDDR6', Memory Interface Width '192-bit / 128-bit' (12 GB = 192-bit), Graphics Card Power (W) '170'. Release year from GeForce news 'GeForce RTX 3060' dated January 12, 2021 (https://www.nvidia.com/en-us/geforce/news/geforce-rtx-3060/). The 8 GB/128-bit variant is not included. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/geforc....
NVIDIA GeForce RTX 4070
RTX Blackwell whitepaper Table 6 (p.54-55, 'RTX 4070'): '12 GB GDDR6X', 192-bit, 'Memory Clock (Data Rate) 21 Gbps', 'Memory Bandwidth 504 GB/sec', TGP 200 W. These apply to the original GDDR6X card; NVIDIA states no bandwidth or data rate for the later GDDR6 variant. Power: Total Graphics Power. Compare page / RTX 4070 family page: '12 GB GDDR6 / 12 GB GDDR6X', 192-bit, Total Graphics Power '200'. Release year from GeForce news 'GeForce RTX 4070' dated April 12, 2023 (https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4070/). Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/black....
NVIDIA GeForce RTX 4070 SUPER
NVIDIA CES 2024 spec image for RTX 4070 SUPER (on https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4080-4070-ti-4070-super-gpu/): Frame Buffer '12GB G6X', Memory Subsystem '48MB L2 / 504 GB/sec', TGP '220 W'. Per-pin data rate not stated. Power: Total Graphics Power. Compare page: 12 GB GDDR6X, 192-bit, Total Graphics Power '220'. Release year from GeForce news dated January 08, 2024 (40 SUPER series launching January 2024). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/geforc....
NVIDIA GeForce RTX 4070 Ti
RTX Blackwell whitepaper Table 5 (p.51-52, 'RTX 4070 Ti'): '12 GB GDDR6X', 192-bit, 'Memory Clock (Data Rate) 21 Gbps', 'Memory Bandwidth 504 GB/sec', TGP 285 W. Power: Total Graphics Power. Compare page / RTX 4070 family page (https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4070-family/): 12 GB GDDR6X, 192-bit, Total Graphics Power '285'. Available January 5, 2023 per GeForce CES 2023 news (seen via search result only). Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/black....
NVIDIA GeForce RTX 5070
Compare page 'Memory Bandwidth' row: '672 GB/sec'. RTX Blackwell whitepaper Table 6 (p.54-55): 'Memory Clock (Data Rate) 28 Gbps', 'Memory Bandwidth 672 GB/sec', TGP 250 W. Power: Total Graphics Power. Compare page / RTX 5070 family page (https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5070-family/): '12 GB GDDR7', 192-bit, Total Graphics Power '250'. GeForce news 'RTX 5070 out now' dated March 05, 2025. Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/black....
NVIDIA GeForce RTX 3080 Ti
Ada whitepaper Table 3 / Appendix B (p.32-35, 'RTX 3080 Ti' column): 384-bit, 'Memory Clock (Data Rate) 19 Gbps', 'Memory Bandwidth 912 GB/sec', TGP 350 W. RTX 4080 SUPER launch spec image (https://www.nvidia.com/content/dam/en-zz/Solutions/geforce/news/geforce-rtx-4080-4070-ti-4070-super-gpu/nvidia-geforce-rtx-ces-2024-4080-super-specifications.png) also lists RTX 3080 Ti '12GB G6X', '6MB L2 / 912 GB/sec', TGP '350W'. Power: Graphics Card Power. Compare page: 12 GB GDDR6X, 384-bit, Graphics Card Power '350'. Release year: unveiled at COMPUTEX 2021 per GeForce news (https://www.nvidia.com/en-us/geforce/news/rtx-3080-ti-3070-ti-graphics-cards/), seen via search result only. Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/ada/n....
NVIDIA GeForce RTX 4060 Ti (16 GB)
Same NVIDIA RTX 4060 Ti spec image (VRAM '8GB / 16GB', '288 GB/s'); GeForce news (https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4060-4060ti/): the 16GB model has 'additional graphics memory but otherwise identical specifications'. Per-pin data rate not stated. Power: Total Graphics Power. Compare page / RTX 4060 family page: Total Graphics Power '165 or 160' listed against '16 GB GDDR6 or 8 GB GDDR6', so 16 GB = 165 W by position; the launch spec image shows only 160 W for the 4060 Ti column. Arrived July 2023 per GeForce news dated May 18, 2023. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/geforc....
NVIDIA GeForce RTX 5060 Ti (16 GB)
Compare page (50 Series): RTX 5060 Ti column '16 GB / 8 GB GDDR7', 128-bit, 'Memory Bandwidth 448 GB/sec'. Per-pin data rate not stated. Power: Total Graphics Power. Total Graphics Power '180'. Available April 16, 2025 per NVIDIA newsroom (April 15, 2025). Second source: https://www.nvidia.com/en-us/geforce/graphics-cards/50-se....
NVIDIA GeForce RTX 4070 Ti SUPER
NVIDIA CES 2024 spec image for RTX 4070 Ti SUPER (on https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4080-4070-ti-4070-super-gpu/): Frame Buffer '16GB G6X', Memory Subsystem '48MB L2 / 672 GB/sec', TGP '285 W'. Per-pin data rate not stated. Power: Total Graphics Power. Compare page: 16 GB GDDR6X, 256-bit, Total Graphics Power '285'. Release year from GeForce news dated January 08, 2024. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/geforc....
NVIDIA GeForce RTX 4080
Ada whitepaper Table 3 / Appendix B (p.32-35, 'RTX 4080 16 GB'): 256-bit, 'Memory Clock (Data Rate) 22.4 Gbps', 'Memory Bandwidth 716.8 GB/sec', TGP 320 W. Same in RTX Blackwell whitepaper Table 4. Power: Total Graphics Power. Compare page / RTX 4080 family page (https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4080-family/): 16 GB GDDR6X, 256-bit, Total Graphics Power '320'. release_year 2022 not verified against a dated NVIDIA page. Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/ada/n....
NVIDIA GeForce RTX 4080 SUPER
NVIDIA CES 2024 spec image for RTX 4080 SUPER (on https://www.nvidia.com/en-us/geforce/news/geforce-rtx-4080-4070-ti-4070-super-gpu/): Frame Buffer '16GB G6X', Memory Subsystem '64MB L2 / 736 GB/sec', TGP '320 W'. Article text: 'GDDR6X video memory (VRAM) running at 23 Gbps' (23 x 256 / 8 = 736, consistent). Power: Total Graphics Power. Compare page: 16 GB GDDR6X, 256-bit, Total Graphics Power '320'. GeForce news dated January 08, 2024: 'arrives January 31st'. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/geforc....
NVIDIA GeForce RTX 5070 Ti
Compare page 'Memory Bandwidth' row: '896 GB/sec'. RTX Blackwell whitepaper Table 5 (p.51-52): 'Memory Clock (Data Rate) 28 Gbps', 'Memory Bandwidth 896 GB/sec', TGP 300 W. Power: Total Graphics Power. Compare page / RTX 5070 family page: '16 GB GDDR7', 256-bit, Total Graphics Power '300'. GeForce news 'RTX 5070 Ti out now' dated February 20, 2025. Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/black....
NVIDIA GeForce RTX 5080
Compare page 'Memory Bandwidth' row: '960 GB/sec'. RTX Blackwell whitepaper Table 4 (p.49-50): 'Memory Clock (Data Rate) 30 Gbps', 'Memory Bandwidth 960 GB/sec'; text: 'RTX 5080 ships with 30 Gbps GDDR7 memory, delivering 960 GB/sec'. Power: Total Graphics Power. Compare page / product page (https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5080/): '16 GB GDDR7', 256-bit, Total Graphics Power '360'. GeForce news 'RTX 5090 & 5080 out now' dated January 30, 2025. Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/black....
NVIDIA GeForce RTX 3090
Ampere GA102 whitepaper Table 9 (p.44-45, 'RTX 3090 FE'): 384-bit, 'Memory Clock (Data Rate) 19.5 Gbps', 'Memory Bandwidth (GB/sec) 936 GB/sec', 'Total Graphics Power (TGP) 350 Watts'. Also RTX Blackwell whitepaper Table 1/3. Power: Graphics Card Power. Compare page: 24 GB GDDR6X, 384-bit, Graphics Card Power '350'. Product page: https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090-3090ti/. Release year: RTX 30 Series GeForce news dated September 01, 2020. Second source: https://www.nvidia.com/content/PDF/nvidia-ampere-ga-102-g....
NVIDIA GeForce RTX 3090 Ti
Ada whitepaper Table 1 (p.13) and Table 2 / Appendix A (p.29-31, 'GeForce RTX 3090 Ti'): 384-bit, 'Memory Clock (Data Rate) 21 Gbps', 'Memory Bandwidth 1008 GB/sec', 'TGP (Total Graphics Power) 450 W'. GeForce news (March 29, 2022): '24GB of the fastest 21Gbps GDDR6X memory'. Power: Graphics Card Power. Compare page: 24 GB GDDR6X, 384-bit, Graphics Card Power '450'. Release year from GeForce news 'GeForce RTX 3090 Ti out now' dated March 29, 2022 (https://www.nvidia.com/en-us/geforce/news/geforce-rtx-3090-ti-out-now/). Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/ada/n....
NVIDIA GeForce RTX 4090
Ada whitepaper Table 1 (p.13) and Table 2 / Appendix A (p.29-31): 384-bit, 'Memory Clock (Data Rate) 21 Gbps', 'Memory Bandwidth 1008 GB/sec', 'TGP (Total Graphics Power) 450 W'. Same in RTX Blackwell whitepaper Table 1/3. Power: Total Graphics Power. Compare page / product page (https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/): 24 GB GDDR6X, 384-bit, Total Graphics Power '450'. Available October 12, 2022 per GeForce news 'GeForce RTX 4090 Out Now' (seen via search result only). Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/ada/n....
NVIDIA GeForce RTX 5090
Compare page 'Memory Bandwidth' row: '1792 GB/sec'. RTX Blackwell whitepaper Table 1 (p.14) and Table 3 (p.46-47): 512-bit, 'Memory Clock (Data Rate) 28 Gbps', 'Memory Bandwidth 1792 GB/sec', TGP 575 W. Power: Total Graphics Power. Compare page / product page (https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/): '32 GB GDDR7', 512-bit, Total Graphics Power '575'. GeForce news dated January 30, 2025. Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/black....
NVIDIA RTX A4000
RTX A4000 datasheet (scanned PDF, read visually): 'Memory interface 256-bit', 'Memory bandwidth 448 GB/s', 'Power consumption Total board power: 140 W'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: GPU Memory '16GB GDDR6 with error-correction code (ECC)', Max Power Consumption '140 W' (datasheet calls it 'Total board power'). Announced April 12, 2021 per NVIDIA newsroom (seen via search result). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX 4000 Ada Generation
RTX 4000 Ada datasheet: 'GPU Memory 20GB GDDR6', 'Memory Interface 160 bit', 'Memory Bandwidth 360GB/s', 'Total board power: 130W'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: '20GB GDDR6 with error-correction code (ECC)', Max Power Consumption '130W', single slot. Announced at SIGGRAPH (Aug) 2023, available fall 2023 per NVIDIA newsroom (seen via search result). The separate RTX 4000 SFF Ada is not included. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX 4500 Ada Generation
RTX 4500 Ada datasheet: 'GPU Memory 24GB GDDR6', 'Memory Interface 192 bit', 'Memory Bandwidth 432GB/s', 'Total board power: 210W'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: '24GB GDDR6 with error-correction code (ECC)', Max Power Consumption '210W'. Available fall 2023 per NVIDIA newsroom (seen via search result). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX PRO 4000 Blackwell
Datasheet: 'GPU Memory 24 GB GDDR7 with ECC', 'Memory Interface 192-bit', 'Memory Bandwidth 672 GB/s', 'Total board power: 145 W'. Product page also: 'Memory Bandwidth 672 GB/sec'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: Memory Configuration '24GB GDDR7 with error-correcting code (ECC)', Max Power Consumption '145 W'. NVIDIA newsroom (March 18, 2025): RTX PRO 5000/4500/4000 'available in the summer' (2025). The RTX PRO 4000 Blackwell SFF Edition is a separate product, not included. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX A5000
RTX A5000 datasheet: 'Memory interface 384-bit', 'Memory bandwidth 768 GB/s', 'Power consumption Total board power: 230 W'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: '24GB GDDR6 with error-correction code (ECC)', Max Power Consumption '230 W'; NVLink 2-way. Announced April 12, 2021 per NVIDIA newsroom (seen via search result). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX 5000 Ada Generation
RTX 5000 Ada datasheet: 'GPU Memory 32GB GDDR6', 'Memory Interface 256 bit', 'Memory Bandwidth 576GB/s', 'Total board power: 250W'. The datasheet is the one linked from the product page (resources.nvidia.com viewer); the PDF file itself is served from cdn.pathfactory.com/assets/10412/contents/1095460/b57d512c-c1b8-4fea-b9e5-cb95bcb33c44.pdf. Per-pin data rate not stated. Power: Max Power Consumption. Product page: '32GB GDDR6 with error-correction code (ECC)', Max Power Consumption '250W'. Shipping from August 2023 (SIGGRAPH) per NVIDIA newsroom (seen via search result). Second source: https://resources.nvidia.com/en-us-briefcase-for-datashee....
NVIDIA RTX PRO 4500 Blackwell Workstation Edition
Datasheet: 'GPU memory 32 GB GDDR7 with ECC', 'Memory interface 256-bit', 'Memory bandwidth 896 GB/s', 'Total board power: 200 W'. Product page: 'Memory Bandwidth 896 GB/s'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: '32GB GDDR7 with error-correcting code (ECC)', Max Power Consumption '200 W'. Summer 2025 availability per NVIDIA newsroom (March 18, 2025). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/data-c....
NVIDIA RTX A6000
RTX A6000 datasheet: 'Memory interface 384-bit', 'Memory bandwidth 768 GB/s', 'Total board power: 300 W'. Ampere GA102 whitepaper Table 3 (p.15-16): 'Memory Clock (Data Rate) 16 Gbps', 'Memory Bandwidth 768 GB/sec', 'TGP 300W'. Power: Max Power Consumption. Product page: '48 GB GDDR6 with error-correcting code (ECC)', Max Power Consumption '300 W'. Data rate from https://www.nvidia.com/content/PDF/nvidia-ampere-ga-102-gpu-architecture-whitepaper-v2.pdf. release_year 2020 not verified against a dated NVIDIA page. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX 6000 Ada Generation
RTX 6000 Ada datasheet (nvidia.com copy; identical values in dam-cdn.nvd.orangelogic.com/AssetLink/83uo5f1t66fkm0dh16mi5uq1412l7ux5.pdf): 'GPU memory 48GB GDDR6', 'Memory interface 384-bit', 'Memory bandwidth 960 GB/s', 'Total board power: 300 W'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: '48GB GDDR6 with error-correcting code (ECC)', Max Power Consumption '300 W'. Announced September 20, 2022, available from December 2022 per NVIDIA newsroom (seen via search result). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/design....
NVIDIA RTX PRO 5000 Blackwell (48 GB)
Datasheet: 'GPU Memory 48 GB GDDR7 with ECC | 72 GB GDDR7 with ECC', 'Memory Interface 384-bit', 'Memory Bandwidth 1,344 GB/s', 'Total board power: 300 W'. Product page: 'Memory Bandwidth 1,344 GB/sec'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: Max Power Consumption '300 W'. Summer 2025 availability per NVIDIA newsroom (March 18, 2025). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX PRO 5000 72GB Blackwell
Same datasheet and product-page spec table as the 48 GB model: one column lists '48 GB GDDR7 with ECC' and '72 GB GDDR7 with ECC', 'Memory Interface 384-bit', 'Memory Bandwidth 1,344 GB/s' (one value for both). Per-pin data rate not stated. Power: Max Power Consumption. Max Power Consumption '300 W' (same column). NVIDIA blog dated December 18, 2025 (https://blogs.nvidia.com/blog/rtx-pro-5000-72gb-blackwell-gpu/): 'now generally available'. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
Max-Q datasheet: 'GPU Memory 96 GB GDDR7 with ECC', 'Memory Interface 512-bit', 'Memory Bandwidth 1792 GB/s', 'Total board power: 300 W'. Product page: 'Memory Bandwidth 1792 GB/sec'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: Max Power Consumption '300 W'; RTX PRO 6000 series page lists Power Consumption '300 W' for Max-Q. Available from April 2025 per NVIDIA newsroom (March 18, 2025). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/produc....
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Datasheet: 'GPU Memory 96 GB GDDR7 with ECC', 'Memory Interface 512-bit', 'Memory Bandwidth 1792 GB/s', 'Total board power: 600 W'. Product page: 'Memory Bandwidth 1792 GB/sec'. Per-pin data rate not stated. Power: Max Power Consumption. Product page: Max Power Consumption '600 W', 5.4" x 12" dual slot, double flow-through. NVIDIA newsroom (March 18, 2025): available from distributors starting April (2025). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/data-c....
AMD Radeon RX 7600
AMD product page: 'Memory Bandwidth Up to 288 GB/s'; 'Effective Memory Bandwidth Up to 477 GB/s' (includes 32 MB Infinity Cache) - not used. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1102; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); NOT in the ROCm 7.14.0 Linux GPU table (newly listed in the 10.x stream). ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2; no WSL2 option for this GPU); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: AMD press release May 24, 2023 (available from May 25). Architecture per ROCm compatibility matrix / GPU specs page (RDNA 3). Second source: https://ir.amd.com/news-events/press-releases/detail/1133....
AMD Radeon RX 9050
AMD product page: 'Up to 288 GB/s'. 32 MB Infinity Cache. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1200; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); NOT in the ROCm 7.14.0 Linux GPU table (newly listed in the 10.x stream). ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2; no WSL2 option for this GPU); not listed in the HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Extra 9000-series desktop card found on AMD spec page. No launch date found on AMD pages; release_year null. Architecture per ROCm compatibility matrix / GPU specs page and AMD RDNA 4 press releases. Second source: https://www.amd.com/en/products/specifications/graphics.html.
AMD Radeon RX 9060
AMD product page: 'Up to 288 GB/s'. 32 MB Infinity Cache. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1200; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Extra 9000-series desktop card found on AMD spec page. No launch date found on AMD pages; release_year null. Architecture per ROCm compatibility matrix / GPU specs page and AMD RDNA 4 press releases. Second source: https://www.amd.com/en/products/specifications/graphics.html.
AMD Radeon RX 9060 XT (8GB)
AMD product page: 'Up to 320 GB/s'. 32 MB Infinity Cache. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1200; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: Computex press release May 20, 2025 (TBP 'Starting at 150W', availability 'later this year'). ROCm docs list 'AMD Radeon RX 9060 XT' without distinguishing 8 GB vs 16 GB. Architecture per ROCm compatibility matrix / GPU specs page and AMD RDNA 4 press releases. Second source: https://ir.amd.com/news-events/press-releases/detail/1253....
AMD Radeon RX 7700 XT
AMD product page: 'Up to 432 GB/s'; effective (with 48 MB Infinity Cache) 'Up to 1995 GB/s' - not used. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1101; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2; no WSL2 option for this GPU); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: AMD press release Aug 25, 2023 (board-partner availability from Sept 6, 2023). Architecture per ROCm compatibility matrix / GPU specs page (RDNA 3). Second source: https://ir.amd.com/news-events/press-releases/detail/1153....
AMD Radeon RX 9070 GRE
AMD product page: 'Up to 432 GB/s'. 48 MB Infinity Cache. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1201; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: AMD Computex 2026 blog (May 31, 2026) says the card launches 'globally' with availability 'starting June 2'; product page has no launch date and any earlier regional launch was not verified from an AMD source. CONFLICT: ROCm GPU specs page lists 'VRAM (GiB) 16' for RX 9070 GRE; AMD product page and the Computex 2026 blog say 12 GB - 12 used. Architecture per ROCm compatibility matrix / GPU specs page and AMD RDNA 4 press releases. Second source: https://www.amd.com/en/blogs/2026/amd-computex-2026-10-ye....
AMD Radeon RX 7600 XT
AMD product page: 'Up to 288 GB/s'; effective (with 32 MB Infinity Cache) 'Up to 477 GB/s' - not used. Power: Typical Board Power (Desktop). ROCm on Linux: not listed (absent from ROCm 10.1.0 compatibility matrix and ROCm 7.14.0 Linux GPU table; same gfx1102 target as the listed RX 7600). ROCm on Windows: not listed in ROCm 10.1.0 matrix; listed as supported (Runtime and HIP SDK) in HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: AMD press release Jan 8, 2024 (available from Jan 24, 2024); press release lists TBP 'Starting at 190W'. 16 GB on a 128-bit bus. ROCm caveat: not on the ROCm 10.1.0 matrix even though RX 7600 (same gfx1102) is. Architecture per ROCm compatibility matrix / GPU specs page (RDNA 3). Second source: https://ir.amd.com/news-events/press-releases/detail/1176....
AMD Radeon RX 9060 XT (16GB)
AMD product page: 'Up to 320 GB/s'. 32 MB Infinity Cache. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1200; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: Computex press release May 20, 2025 (TBP 'Starting at 160W'). ROCm docs list 'AMD Radeon RX 9060 XT' without memory variant. Architecture per ROCm compatibility matrix / GPU specs page and AMD RDNA 4 press releases. Second source: https://ir.amd.com/news-events/press-releases/detail/1253....
AMD Radeon RX 9060 XT LP
AMD product page: 'Up to 320 GB/s'. 32 MB Infinity Cache. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1200; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2; no WSL2 option for this GPU); not listed in the HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Extra 9000-series desktop card (low-power 16 GB variant). Product page power reads 'Up to 140W'. No launch date found; release_year null. Architecture per ROCm compatibility matrix / GPU specs page and AMD RDNA 4 press releases. Second source: https://www.amd.com/en/products/specifications/graphics.html.
AMD Radeon RX 7900 GRE
AMD product page: 'Up to 576 GB/s'; effective (with 64 MB Infinity Cache) 'Up to 2250 GB/s' - not used. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1100; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); not listed in the HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year from AMD product page 'Launch Date 7/27/2023'. ROCm caveat: in ROCm 10.1.0 matrix (incl. Windows) but not in the HIP SDK for Windows 7.2.0 GPU table. Architecture per ROCm compatibility matrix / GPU specs page (RDNA 3). Second source: https://rocm.docs.amd.com/en/latest/compatibility/compati....
AMD Radeon RX 7700
AMD product page: 'Up to 624 GB/s'; effective (with 40 MB Infinity Cache) 1927 GB/s - not used. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1101; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2; no WSL2 option for this GPU); not listed in the HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Extra card (not in the requested list) included because it is a 16 GB RX 7000 desktop card on AMD's spec page. No launch date on AMD page; release_year null. Values from AMD graphics specifications data (product page URL as listed there). Architecture per ROCm compatibility matrix / GPU specs page (RDNA 3). Second source: https://www.amd.com/en/products/specifications/graphics.html.
AMD Radeon RX 7800 XT
AMD product page: 'Up to 624 GB/s'; effective (with 64 MB Infinity Cache) 'Up to 2708 GB/s' - not used. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1101; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: AMD press release Aug 25, 2023 (available from Sept 6, 2023). Architecture per ROCm compatibility matrix / GPU specs page (RDNA 3). Second source: https://ir.amd.com/news-events/press-releases/detail/1153....
AMD Radeon RX 9070
AMD product page: 'Up to 640 GB/s'. 64 MB Infinity Cache. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1201; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: AMD press release Feb 28, 2025 (available from March 6, 2025). Architecture per ROCm compatibility matrix / GPU specs page and AMD RDNA 4 press releases. Second source: https://ir.amd.com/news-events/press-releases/detail/1238....
AMD Radeon RX 9070 XT
AMD product page: 'Up to 640 GB/s'. 64 MB Infinity Cache. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1201; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Release year: AMD press release Feb 28, 2025 (available from March 6, 2025). ROCm 10.1.0 release notes list as resolved a PyTorch training/fine-tuning GPU-reset issue that affected RX 9070 Series and Radeon AI PRO R9700. Architecture per ROCm compatibility matrix / GPU specs page and AMD RDNA 4 press releases. Second source: https://ir.amd.com/news-events/press-releases/detail/1238....
AMD Radeon RX 7900 XT
AMD product page: 'Up to 800 GB/s'; effective (with 80 MB Infinity Cache) 'Up to 2900 GB/s' - not used. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1100; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Product page 'Launch Date 12/13/2022'. Bus width 320-bit from AMD press release Nov 3, 2022 (product page omits Memory Interface; 320 x 20 / 8 = 800 matches stated bandwidth). CONFLICT: product page 'Typical Board Power (Desktop) 315 W' vs press release table 'TBP 300W'; product page value used. Architecture per ROCm compatibility matrix / GPU specs page (RDNA 3). Second source: https://ir.amd.com/news-events/press-releases/detail/1099....
AMD Radeon RX 7900 XTX
AMD product page: 'Memory Bandwidth Up to 960 GB/s'; 'Effective Memory Bandwidth Up to 3500 GB/s' (includes 96 MB Infinity Cache) - not used. Power: Typical Board Power (Desktop). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1100; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). Product page 'Launch Date 12/13/2022'. Bus width 384-bit from AMD press release Nov 3, 2022 ('24GB of high-speed GDDR6 memory running at 20Gbps over a 384-bit memory bus'); product page omits Memory Interface. Architecture per ROCm compatibility matrix / GPU specs page (RDNA 3). Second source: https://ir.amd.com/news-events/press-releases/detail/1099....
AMD Radeon PRO W7800
AMD product page: 'Peak Memory Bandwidth 576 GB/s'. 64 MB Infinity Cache. Power: Total Board Power (TBP). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1100; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). ECC memory supported per AMD spec. Memory speed: product page lists none; AMD press release Apr 13, 2023 table header reads 'GDDR6 ECC Memory (18GB/s)' (apparently 18 Gbps; 256 x 18 / 8 = 576 matches). Release year: press release Apr 13, 2023 (retail availability 'starting in Q2, 2023'). Architecture from product page 'GPU Architecture AMD RDNA 3'. Second source: https://ir.amd.com/news-events/press-releases/detail/1123....
AMD Radeon AI PRO R9600
AMD product page: 'Peak Memory Bandwidth 640 GB/s'. Memory speed not stated. Power: Total Board Power (TBP). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1201; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); NOT in the ROCm 7.14.0 Linux GPU table (newly listed in the 10.x stream). ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2; no WSL2 option for this GPU); not listed in the HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). ECC 'Yes (Linux Only)'. Newer Radeon AI PRO card: active cooling, single slot, 48 CUs (ROCm GPU specs). ROCm 10.1.0 release notes: 'ROCm 10.1.0 adds support for AMD Radeon AI PRO R9600 GPUs.' No launch date found; release_year null. Second source: https://rocm.docs.amd.com/en/latest/about/release-notes.html.
AMD Radeon AI PRO R9600D
AMD product page: 'Peak Memory Bandwidth 640 GB/s'. Memory speed not stated. Power: Total Board Power (TBP). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1201; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2; no WSL2 option for this GPU); not listed in the HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). ECC 'Yes (Linux Only)'. Newer Radeon AI PRO card: passive cooling, single slot (server/chassis-airflow design). No launch date found; release_year null. Second source: https://www.amd.com/en/products/specifications/profession....
AMD Radeon AI PRO R9700
AMD product page: 'Peak Memory Bandwidth 640 GB/s'. 64 MB Infinity Cache. Memory speed not stated (640 x 8 / 256 implies 20 Gbps). Power: Total Board Power (TBP). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1201; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). ECC 'Yes (Linux Only)' per AMD spec. Release year: Computex press release May 20, 2025 ('expected to be available from leading board partners starting in July 2025'). Active cooling, double slot, 12V-2x6 connector. ROCm 10.1.0 release notes list as resolved a PyTorch training GPU-reset issue affecting R9700. Second source: https://ir.amd.com/news-events/press-releases/detail/1253....
AMD Radeon AI PRO R9700S
AMD product page: 'Peak Memory Bandwidth 640 GB/s'. Memory speed not stated. Power: Total Board Power (TBP). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1201; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); NOT in the ROCm 7.14.0 Linux GPU table (newly listed in the 10.x stream). ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); not listed in the HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). ECC 'Yes (Linux Only)'. Newer Radeon AI PRO card: passive cooling, double slot (server/chassis-airflow design, not a typical home card). No launch date found; release_year null. Not in ROCm 7.14.0 Linux table or HIP SDK Windows 7.2.0 table; listed in ROCm 10.1.0 matrix. Second source: https://www.amd.com/en/products/specifications/profession....
AMD Radeon PRO W7800 48GB
AMD product page: 'Peak Memory Bandwidth 864 GB/s'. 96 MB Infinity Cache. Memory speed not stated (864 x 8 / 384 implies 18 Gbps). Power: Total Board Power (TBP). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1100; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); not listed in the HIP SDK for Windows 7.2.0 GPU table. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). ECC memory supported per AMD spec. Extra card (48 GB variant of W7800). No launch date found on AMD pages; release_year null. Architecture from product page 'AMD RDNA 3'. ROCm caveat: not in HIP SDK for Windows 7.2.0 GPU table. Second source: https://www.amd.com/en/products/specifications/profession....
AMD Radeon PRO W7900
AMD product page: 'Peak Memory Bandwidth 864 GB/s'. 96 MB Infinity Cache. Power: Total Board Power (TBP). ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1100; Ubuntu 26.04.1/24.04.5/22.04.5, RHEL 10.2/9.8); also listed in ROCm 7.14.0 Linux GPU table. ROCm on Windows: supported (ROCm 10.1.0 matrix: Windows 11 25H2 and WSL2); HIP SDK for Windows 7.2.0 table: Runtime and HIP SDK supported. Vulkan: llama.cpp Vulkan backend (README backend table: 'Vulkan | GPU'). llama.cpp HIP backend (needs ROCm; build.md). ECC memory supported per AMD spec. Memory speed: product page lists none; AMD press release Apr 13, 2023 header 'GDDR6 ECC Memory (18GB/s)' (apparently 18 Gbps; 384 x 18 / 8 = 864 matches). Release year: press release Apr 13, 2023. Triple-slot board; AMD also lists a 'Radeon PRO W7900 Dual Slot' with identical memory specs (48 GB, 384-bit, 864 GB/s, 295 W TBP) and the same ROCm status - not given a separate record. Second source: https://ir.amd.com/news-events/press-releases/detail/1123....
Intel Arc B570 Graphics
Intel ark: 'Graphics Memory Bandwidth 380 GB/s'. Power: Total Board Power (TBP). Vulkan: Intel spec: Vulkan 1.3; llama.cpp Vulkan backend. llama.cpp SYCL backend (docs/backend/SYCL.md 'Verified devices': 'Intel Arc B-Series | Support | Arc B580'; Linux and Windows 11 supported). SYCL.md notes oneAPI 2026.1 build speed-up 'measured ... on Arc B570'. Intel OpenVINO 2025.2 release notes cite optimizations for 'Intel Arc B Series Graphics'. Intel B-series marketing cites AI Playground. Intel ark 'Launch Date Q4'24' (release_year 2024 from ark); Intel press release Dec 3, 2024 gives retail availability 'starting Jan. 16, 2025'. PCIe 4.0 x8. Second source: https://www.intel.com/content/www/us/en/newsroom/news/int....
Intel Arc B580 Graphics
Intel ark: 'Graphics Memory Bandwidth 456 GB/s'. Power: Total Board Power (TBP). Vulkan: Intel spec: Vulkan 1.3; llama.cpp Vulkan backend. llama.cpp SYCL backend (docs/backend/SYCL.md 'Verified devices': 'Intel Arc B-Series | Support | Arc B580'; Linux and Windows 11 supported). Intel OpenVINO 2025.2 release notes cite optimizations for 'Intel Arc B Series Graphics'. Intel ark 'Launch Date Q4'24'; press release Dec 3, 2024: available 'starting Dec. 13'. PCIe 4.0 x8. Only B-series card named in llama.cpp SYCL verified-device table. Second source: https://www.intel.com/content/www/us/en/newsroom/news/int....
Intel Arc Pro B50 Graphics
Intel ark: 'Graphics Memory Bandwidth 224 GB/s'. Power: Total Board Power (TBP). Vulkan: Intel spec: Vulkan 1.3; llama.cpp Vulkan backend. llama.cpp SYCL backend (docs/backend/SYCL.md 'Verified devices': 'Intel Arc B-Series | Support | Arc B580'; Linux and Windows 11 supported). Intel datasheet 'API Support: DirectX 12 Ultimate, oneAPI, OpenCL 3.0, OpenGL 4.6, OpenVINO, Vulkan 1.3'; OS support Windows 11/10 and Linux Ubuntu. Intel ark 'Launch Date Q3'25'. Datasheet: '70W Total Board Power', 'No Connector Required', dual-slot low profile. PCIe 5.0 x8. Second source: https://download.intel.com/newsroom/2025/client-computing....
Intel Arc Pro B60 Graphics
Intel ark: 'Graphics Memory Bandwidth 456 GB/s'. Power: Total Board Power (TBP). Vulkan: Intel spec: Vulkan 1.3; llama.cpp Vulkan backend. llama.cpp SYCL backend (docs/backend/SYCL.md 'Verified devices': 'Intel Arc B-Series | Support | Arc B580'; Linux and Windows 11 supported). Intel datasheet 'API Support: DirectX 12 Ultimate, oneAPI, OpenCL 3.0, OpenGL 4.6, OpenVINO, Vulkan 1.3'; 'Multi-GPU Linux Ready'; OS support Windows 11/10 and Linux Ubuntu. Intel ark 'Launch Date Q2'25'; ark 'TBP 200 W' and 'TBP - LP 120 - 200 W'; 2025 datasheet 'Consumption 120-200W Total Board Power' (partner boards vary). Sold via board partners. PCIe 5.0 x8. Second source: https://download.intel.com/newsroom/2025/client-computing....
Intel Arc Pro B65 Graphics
Intel ark and datasheet: '608 GB/s'. Memory speed not stated (608 x 8 / 256 implies 19 Gbps). Power: Total Board Power (TBP). Vulkan: Intel spec: Vulkan 1.3; llama.cpp Vulkan backend. Not named in llama.cpp SYCL verified-device table (only B580 is). Intel datasheet: 'Leverage Linux multi-GPU compatibility and OneAPI integration'; OS support Windows 11/10 and Linux Ubuntu. Additional Intel B-series card. Intel ark 'Launch Date Q1'26'; datasheet '200W Total Board Power', 'Available as Partner Branded Card'. PCIe 5.0 x16. Second source: https://www.intel.com/content/dam/www/central-libraries/u....
Intel Arc Pro B70 Graphics
Intel ark and datasheet: '608 GB/s'. Memory speed not stated (608 x 8 / 256 implies 19 Gbps). Power: Total Board Power (TBP). Vulkan: Intel spec: Vulkan 1.3; llama.cpp Vulkan backend. Not named in llama.cpp SYCL verified-device table. Intel datasheet 'API Support: DirectX 12 Ultimate, oneAPI, OpenCL 3.0, OpenGL 4.6, OpenVINO, Vulkan 1.3'; 'Scalable multiple-GPU LLM Linux support'. Intel Arc Pro B-series page benchmarks use container 'intel/llm-scaler-vllm:1.3' (vLLM). Additional Intel B-series card (32 Xe-cores). Intel ark 'Launch Date Q1'26'; ark 'TBP 230 W' and 'TBP - LP 160-290'; datasheet 'Consumption 160W-290W, 230W for Intel Branded Card'. PCIe 5.0 x16. Second source: https://www.intel.com/content/dam/www/central-libraries/u....
Apple M1 Pro
Tech specs: "200GB/s memory bandwidth"; newsroom (Oct 18, 2021): "M1 Pro offers up to 200GB/s of memory bandwidth with support for up to 32GB of unified memory." MacBook Pro 14-inch (2021) and 16-inch (2021): 16GB, configurable to 32GB (support.apple.com/en-us/111902, 111901). Laptop-only chip. GPU variants: 14-core or 16-core (CPU 8- or 10-core); same 200GB/s. gpu_cores = top variant. Shipped only in MacBook Pro, so no desktop power figure. Second source: https://support.apple.com/en-us/111901.
Apple M1 Max
Mac Studio (2022) tech specs: "400GB/s memory bandwidth" (M1 Max 10-core CPU/24-core GPU base, configurable to 32-core GPU); newsroom: "M1 Max delivers up to 400GB/s of memory bandwidth". Power: Mac Studio (2022): Apple's maximum power for the whole computer, at the wall. MacBook Pro 14/16-inch (2021): 32GB or 64GB (support.apple.com/en-us/111902, 111901); Mac Studio (2022): 32GB, configurable to 64GB (support.apple.com/en-us/111900). GPU variants: 24-core or 32-core, both 400GB/s. gpu_cores = top variant. Apple's power figure was measured on: M1 Max 10-Core CPU & 32-Core GPU, 32GB, 2TB SSD (idle 11 W). Second source: https://support.apple.com/en-us/102027.
Apple M1 Ultra
Mac Studio (2022) tech specs: "800GB/s memory bandwidth"; newsroom (Mar 8, 2022): "Memory bandwidth is increased to 800GB/s". Power: Mac Studio (2022): Apple's maximum power for the whole computer, at the wall. Mac Studio (2022) only: 64GB, configurable to 128GB (support.apple.com/en-us/111900). GPU variants: 48-core or 64-core, both 800GB/s. gpu_cores = top variant; note Apple's power figure is for the 48-core GPU config. Apple's power figure was measured on: M1 Ultra 20-Core CPU & 48-Core GPU, 64GB, 1TB SSD (idle 13 W). Second source: https://support.apple.com/en-us/102027.
Apple M2 Pro
Tech specs: "200GB/s memory bandwidth"; newsroom (Jan 17, 2023): "M2 Pro features 40 billion transistors, 200GB/s of unified memory bandwidth, and up to 32GB of fast, low-latency unified memory." Power: Mac mini (2023): Apple's maximum power for the whole computer, at the wall. MacBook Pro 14/16-inch (2023): 16GB, configurable to 32GB (support.apple.com/en-us/111340, 111838); Mac mini (2023): 16GB, configurable to 32GB (support.apple.com/en-us/111837). GPU variants: 16-core (10-core CPU) or 19-core (12-core CPU), both 200GB/s. gpu_cores = top variant. Apple's power table does not say which GPU variant was measured. Apple's power figure was measured on: M2 Pro, 32GB, 8TB SSD (idle 7 W). Second source: https://support.apple.com/en-us/103253.
Apple M3 Pro
Tech specs: "150GB/s memory bandwidth" (the Oct 30, 2023 newsroom gives no bandwidth figure). MacBook Pro 14-inch and 16-inch (Nov 2023): 18GB, configurable to 36GB (support.apple.com/en-us/117736, 117737). Laptop-only chip. GPU variants: 14-core (11-core CPU) or 18-core (12-core CPU), both 150GB/s. gpu_cores = top variant. Laptop-only, so no desktop power figure. Second source: https://support.apple.com/en-us/117737.
Apple M3 Max (14-core CPU, 30-core GPU)
Tech specs: "M3 Max with 14-core CPU and 30-core GPU (300GB/s memory bandwidth)". MacBook Pro 14/16-inch (Nov 2023): 36GB, or 96GB ('96GB (M3 Max with 14-core CPU)') (support.apple.com/en-us/117736, 117737). Laptop-only. Laptop-only chip, so no desktop power figure. Second source: https://support.apple.com/en-us/117737.
Apple M2 Max
Mac Studio (2023) tech specs: "400GB/s memory bandwidth"; newsroom: "M2 Max features 67 billion transistors, 400GB/s of unified memory bandwidth, and up to 96GB of fast, low-latency unified memory." Power: Mac Studio (2023): Apple's maximum power for the whole computer, at the wall. MacBook Pro 14/16-inch (2023): 32GB, 64GB, or 96GB, where 96GB requires the 38-core GPU (support.apple.com/en-us/111340, 111838); Mac Studio (2023): 32GB, configurable to 64GB or 96GB, 96GB with the 38-core GPU (support.apple.com/en-us/111835). GPU variants: 30-core or 38-core, both 400GB/s. gpu_cores = top variant; Apple's power figure is for the 30-core config. Apple's power figure was measured on: M2 Max 12-Core CPU & 30-Core GPU, 32GB, 512GB SSD (idle 9 W). Second source: https://support.apple.com/en-us/102027.
Apple M3 Max (16-core CPU, 40-core GPU)
Tech specs: "M3 Max with 16-core CPU and 40-core GPU (400GB/s memory bandwidth)". MacBook Pro 14/16-inch (Nov 2023): 48GB, 64GB, or 128GB ('48GB, 64GB, or 128GB (M3 Max with 16-core CPU)') (support.apple.com/en-us/117736, 117737). Laptop-only. Laptop-only chip, so no desktop power figure. Second source: https://support.apple.com/en-us/117737.
Apple M2 Ultra
Tech specs: "800GB/s memory bandwidth"; newsroom (Jun 5, 2023): "M2 Ultra features 800GB/s of system memory bandwidth". Power: Mac Studio (2023): Apple's maximum power for the whole computer, at the wall. Mac Studio (2023): 64GB, configurable to 128GB or 192GB (support.apple.com/en-us/111835); Mac Pro (2023): 64GB, configurable to 128GB or 192GB (support.apple.com/en-us/111343). GPU variants: 60-core or 76-core, both 800GB/s. gpu_cores = top variant. Apple's power figure was measured on: M2 Ultra 24-Core CPU & 76-Core GPU, 192GB, 8TB SSD (idle 10 W). Mac Pro (2023) is listed at 330 W (76-core GPU, 192GB) and 300 W (60-core GPU, 64GB) at support.apple.com/en-us/102839. Second source: https://support.apple.com/en-us/102027.
Apple M4
Tech specs: "120GB/s memory bandwidth"; newsroom (Oct 30, 2024): "M4 supports up to 32GB of unified memory and has higher memory bandwidth of 120GB/s." Power: Mac mini (2024): Apple's maximum power for the whole computer, at the wall. MacBook Pro 14-inch (M4, 2024): 16/24/32GB (support.apple.com/en-us/121552); Mac mini (2024): 16GB, configurable to 24 or 32GB (121555); iMac (24-inch, 2024): 16/24/32GB on the 10-core-GPU model, 16/24GB on the 8-core-GPU model (121557, 121556); MacBook Air (M4, 2025): 16/24/32GB (122209, 122210). iPad Pro is not included. GPU variants: 8-core or 10-core (both 120GB/s). First shipped in iPad Pro (May 2024); Macs from Oct 2024. Apple's power figure was measured on: M4, 16GB, 256GB SSD (idle 4 W). Second source: https://support.apple.com/en-us/103253.
Apple M4 Max (14-core CPU, 32-core GPU)
Tech specs: "410GB/s memory bandwidth" (M4 Max 14-core CPU, 32-core GPU). Power: Mac Studio (2025): Apple's maximum power for the whole computer, at the wall. MacBook Pro 14/16-inch (2024): 36GB ('36GB (M4 Max with 14-core CPU)') (support.apple.com/en-us/121553, 121554); Mac Studio (2025): 36GB (support.apple.com/en-us/122211). Apple's power figure was measured on: M4 Max 14-Core CPU & 32-Core GPU, 36GB, 512GB SSD (idle 6 W). Second source: https://support.apple.com/en-us/102027.
Apple M4 Pro
Tech specs: "273GB/s memory bandwidth"; newsroom (Oct 30, 2024): "M4 Pro supports up to 64GB of fast unified memory and 273GB/s of memory bandwidth". Power: Mac mini (2024): Apple's maximum power for the whole computer, at the wall. MacBook Pro 14/16-inch (2024): 24GB or 48GB (support.apple.com/en-us/121553, 121554); Mac mini (2024): 24GB, configurable to 48GB or 64GB (support.apple.com/en-us/121555). GPU variants: 16-core (12-core CPU) or 20-core (14-core CPU), both 273GB/s. gpu_cores = top variant. Apple's power figure was measured on: M4 Pro, 64GB, 8TB SSD (idle 5 W); Apple does not say which CPU/GPU variant. Second source: https://support.apple.com/en-us/103253.
Apple M4 Max (16-core CPU, 40-core GPU)
Tech specs: "M4 Max with 16-core CPU and 40-core GPU (546GB/s memory bandwidth)"; newsroom: "up to 546GB/s of memory bandwidth". MacBook Pro 14/16-inch (2024): 48GB, 64GB, or 128GB (M4 Max with 16-core CPU) (support.apple.com/en-us/121553, 121554); Mac Studio (2025): 48GB, 64GB, or 128GB (M4 Max with 16-core CPU and 40-core GPU) (support.apple.com/en-us/122211). power_w is null because Apple's Mac Studio (2025) power table (support.apple.com/en-us/102027) lists only the 14-core CPU/32-core GPU, 36GB config (145 W max). It gives no figure for the 40-core GPU config. Second source: https://support.apple.com/en-us/121554.
Apple M5
Tech specs: "153GB/s memory bandwidth"; newsroom (Oct 15, 2025): "M5 offers unified memory bandwidth of 153GB/s". MacBook Pro 14-inch (M5): 16/24/32GB (support.apple.com/en-us/125405); MacBook Air 13/15-inch (M5): 16GB, configurable to 24 or 32GB (126320, 126321). iPad Pro and Apple Vision Pro are not included. No desktop Mac uses M5 (Mac mini moved to M6). GPU variants: 8-core (MacBook Air 13-inch base) or 10-core. Ships only in laptops among Macs, so no desktop power figure. Apple now calls the M5 CPU '4 super cores and 6 efficiency cores'. Second source: https://www.apple.com/newsroom/2025/10/apple-unleashes-m5....
Apple M3 Ultra
Mac Studio (2025) tech specs: "819GB/s memory bandwidth" (both 28-core CPU/60-core GPU and 32-core CPU/80-core GPU); newsroom: "over 800GB/s of memory bandwidth". Power: Mac Studio (2025): Apple's maximum power for the whole computer, at the wall. Mac Studio (2025) only. The current tech specs page (support.apple.com/en-us/122211, read 2026-10-07) lists '96GB unified memory, Configurable to: 256GB'. The launch newsroom (Mar 5, 2025) says it 'starts with 96GB of unified memory, which can be configured up to 512GB'. The 512GB option no longer appears on the specs page, but Apple's power table still lists a 512GB config. GPU variants: 60-core or 80-core, both 819GB/s. gpu_cores = top variant. Apple has not announced an M4 Ultra: Mac Studio (2025) pairs M4 Max with M3 Ultra, and the 2026 Mac Studio moved to M5 Ultra. Apple's power figure was measured on: M3 Ultra 32-Core CPU & 80-Core GPU, 512GB, 16TB SSD (idle 9 W). Second source: https://support.apple.com/en-us/102027.
Apple M6
Tech specs: "153GB/s memory bandwidth, 170GB/s memory bandwidth" and memory "Configurable to: 24GB or 32GB (170GB/s memory bandwidth)". The 16GB config is 153GB/s and the 24/32GB configs are 170GB/s. Newsroom (Aug 25, 2026): "up to 170GB/s of unified memory bandwidth". Power: Mac mini (M6, 2026): Apple's maximum power for the whole computer, at the wall. Mac mini (M6, 2026): 16GB, configurable to 24GB or 32GB (support.apple.com/en-us/128108). It is the only Mac with M6 as of 2026-10-07. Not in the requested list; added because Apple announced it on Aug 25, 2026 (Mac mini). Bandwidth depends on the memory config, so 170 is the 'up to' value and 16GB configs are 153GB/s. Apple's power figure was measured on: M6, 16GB, 256GB SSD (idle 4 W). Second source: https://support.apple.com/en-us/103253.
Apple M5 Max (18-core CPU, 32-core GPU)
Tech specs: "M5 Max with 18-core CPU and 32-core GPU (460GB/s memory bandwidth)". Power: Mac Studio (M5 Max, 2026): Apple's maximum power for the whole computer, at the wall. MacBook Pro 14/16-inch (M5 Max): 36GB ('36GB (M5 Max with 32-core GPU)') (support.apple.com/en-us/126318, 126319); Mac Studio (M5 Max, 2026): 36GB (support.apple.com/en-us/128107). Announced Mar 3, 2026 (MacBook Pro); Mac Studio with M5 Max announced Aug 25, 2026. Apple's power figure was measured on: M5 Max 18-Core CPU & 32-Core GPU, 36GB, 512GB SSD (idle 7 W). Second source: https://support.apple.com/en-us/102027.
Apple M5 Pro
Tech specs: "307GB/s memory bandwidth" (both 15-core CPU/16-core GPU and 18-core CPU/20-core GPU); newsroom (Mar 3, 2026): "M5 Pro supports up to 64GB of unified memory with higher unified memory bandwidth up to 307GB/s." Power: Mac mini (M5 Pro, 2026): Apple's maximum power for the whole computer, at the wall. MacBook Pro 14/16-inch (M5 Pro or M5 Max): 24GB, configurable to 48GB (any M5 Pro) or 64GB ('M5 Pro with 20-core GPU') (support.apple.com/en-us/126318, 126319); Mac mini (M5 Pro, 2026): 24GB, configurable to 48GB or 64GB (support.apple.com/en-us/128108). GPU variants: 16-core (15-core CPU) or 20-core (18-core CPU), both 307GB/s. gpu_cores = top variant. Announced Mar 3, 2026 (MacBook Pro); Mac mini with M5 Pro announced Aug 25, 2026 and available Sep 22, 2026. Apple's power figure was measured on: M5 Pro, 64GB, 8TB SSD (idle 6 W); Apple does not say which CPU/GPU variant. Second source: https://support.apple.com/en-us/103253.
Apple M5 Max (18-core CPU, 40-core GPU)
Tech specs: "M5 Max with 18-core CPU and 40-core GPU (614GB/s memory bandwidth)"; newsroom (Mar 3, 2026): "M5 Max supports up to 128GB of unified memory with higher unified memory bandwidth up to 614GB/s." MacBook Pro 14/16-inch (M5 Max): 48GB, 64GB, or 128GB (M5 Max with 40-core GPU) (support.apple.com/en-us/126318, 126319); Mac Studio (M5 Max, 2026): 48GB, 64GB, or 128GB (M5 Max with 18-core CPU and 40-core GPU) (support.apple.com/en-us/128107). power_w is null because Apple's Mac Studio power table (support.apple.com/en-us/102027) lists only the M5 Max 18-core CPU/32-core GPU, 36GB config (200 W max). Second source: https://support.apple.com/en-us/126319.
Apple M5 Ultra
Tech specs: "1.2TB/s memory bandwidth" (both 30-core CPU/64-core GPU and 36-core CPU/80-core GPU); newsroom (Aug 25, 2026): "1.2TB/s of unified memory bandwidth, 50 percent more than M3 Ultra". Apple states 1.2TB/s, a rounded figure, recorded here as 1200 GB/s. Power: Mac Studio (M5 Ultra, 2026): Apple's maximum power for the whole computer, at the wall. Mac Studio (M5 Ultra, 2026) only: 96GB, configurable to 256GB or 512GB, with 512GB requiring the 36-core CPU/80-core GPU (support.apple.com/en-us/128107; the configure-to-order list shows '256GB unified memory' and '512GB unified memory (M5 Ultra with 36-core CPU and 80-core GPU)'). GPU variants: 64-core (30-core CPU) or 80-core (36-core CPU). gpu_cores = top variant. Announced Aug 25, 2026; available Sep 22, 2026. Apple's power figure was measured on: M5 Ultra 36-Core CPU & 80-Core GPU, 512GB, 16TB SSD (idle 9 W). Second source: https://support.apple.com/en-us/102027.
NVIDIA RTX Spark (N1X) desktop PCs
NVIDIA states neither memory bandwidth nor bus width nor data rate for RTX Spark on the product page or the May 31, 2026 press release, so bandwidth is null. Power: Chip TDP as NVIDIA lists it (the chip only, not the whole PC). OEM Windows desktop PCs, not an NVIDIA-built box. Product page desktop spec column: 'RTX Spark N1X', GPU '6144-core NVIDIA Blackwell RTX GPU', CPU '20-core NVIDIA Grace CPU', Memory Type 'Unified', 'Power (TDP) 140 W' (superchip), Memory Configuration 'Up to 128 GB LPDDR5X'. Laptop versions are 45-80 W TDP, with a 5120-core / up to 64 GB variant. Press release (May 31, 2026): 'available this fall' from ASUS, Dell, HP, Lenovo, Microsoft Surface and MSI; product page shows 'Preorder Now'. Second source: https://nvidianews.nvidia.com/news/nvidia-microsoft-windo....
AMD Ryzen AI Max 385 (Radeon 8050S)
Derived: LPDDR5x-8000 (8000 MT/s) x 256-bit / 8 / 1000 = 256 GB/s, both inputs from the AMD product page. Power: Top of AMD's configurable TDP range (45-120 W; default 55 W): the chip only, and the computer maker sets it. ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1151; Ubuntu 26.04.1 and 24.04.5 (HWE 7.0) with inbox kernel driver; RHEL not offered for APUs). Vulkan: llama.cpp Vulkan backend. llama.cpp HIP backend; llama.cpp build.md: on Linux set GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 to use UMA with an integrated GPU. AMD blogs show LM Studio (llama.cpp runtime) on Windows with Variable Graphics Memory. Product page: 256-bit LPDDR5x, Max. Memory 128 GB, LPDDR5x-8000, Radeon 8050S (32 graphics cores), 8 CPU cores, Default TDP 55W, cTDP 45-120W, Strix Halo. gpu_memory_max_gb 96 from CES 2025 press release (Ryzen AI Max series, 'up to 96GB available for graphics'). Release year: CES press release Jan 6, 2025. Linux: ROCm RDNA3.5 optimization doc says GTT (GPU-mappable system RAM) defaults to 'approximately 50 percent of total system RAM' and can be raised via the TTM page limit (example sets 100 GB on a ~128 GB system); it recommends a small BIOS VRAM reservation (e.g. 0.5 GB) plus a larger GTT limit. AMD states a range of memory sizes up to 128 GB, not a list, so only the maximum is offered here; for a smaller system, enter it as an 'Other' unified-memory device with 256 GB/s. Second source: https://ir.amd.com/news-events/press-releases/detail/1232....
AMD Ryzen AI Max 390 (Radeon 8050S)
Derived: LPDDR5x-8000 (8000 MT/s, product page 'Max Memory Speed') x 256-bit (product page 'System Memory Type 256-bit LPDDR5x') / 8 / 1000 = 256 GB/s. AMD's VGM FAQ states 256 GB/s for the Strix Halo platform generally. Power: Top of AMD's configurable TDP range (45-120 W; default 55 W): the chip only, and the computer maker sets it. ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1151; Ubuntu 26.04.1 and 24.04.5 (HWE 7.0) with inbox kernel driver; RHEL not offered for APUs). Vulkan: llama.cpp Vulkan backend. llama.cpp HIP backend; llama.cpp build.md: on Linux set GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 to use UMA with an integrated GPU. AMD blogs show LM Studio (llama.cpp runtime) on Windows with Variable Graphics Memory. Product page: 256-bit LPDDR5x, Max. Memory 128 GB, LPDDR5x-8000, Radeon 8050S (32 graphics cores), 12 CPU cores, Default TDP 55W, cTDP 45-120W, Strix Halo. gpu_memory_max_gb 96 from CES 2025 press release for the Ryzen AI Max series: 'Featuring up to 128GB of unified memory with up to 96GB available for graphics'. Release year: CES press release Jan 6, 2025. Linux: ROCm RDNA3.5 optimization doc says GTT (GPU-mappable system RAM) defaults to 'approximately 50 percent of total system RAM' and can be raised via the TTM page limit (example sets 100 GB on a ~128 GB system); it recommends a small BIOS VRAM reservation (e.g. 0.5 GB) plus a larger GTT limit. AMD states a range of memory sizes up to 128 GB, not a list, so only the maximum is offered here; for a smaller system, enter it as an 'Other' unified-memory device with 256 GB/s. Second source: https://ir.amd.com/news-events/press-releases/detail/1232....
AMD Ryzen AI Max+ 395 (Radeon 8060S)
AMD blogs state 256 GB/s for Strix Halo ('This platform, codenamed "Strix Halo", doubles the bandwidth to 256 GB/s' - VGM FAQ; 'taking full advantage of the 256 GB/s bandwidth' - Max+ 395 blog). Cross-check from product page: LPDDR5x-8000 x 256-bit = 8000 x 256 / 8 / 1000 = 256 GB/s. Power: Top of AMD's configurable TDP range (45-120 W; default 55 W): the chip only, and the computer maker sets it. ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1151; Ubuntu 26.04.1 and 24.04.5 (HWE 7.0) with inbox kernel driver; RHEL not offered for APUs). Vulkan: llama.cpp Vulkan backend. llama.cpp HIP backend; llama.cpp build.md: on Linux set GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 to use UMA with an integrated GPU. AMD blogs show LM Studio (llama.cpp runtime) on Windows with Variable Graphics Memory. Product page: 'System Memory Type 256-bit LPDDR5x', 'Max. Memory 128 GB', 'Max Memory Speed LPDDR5x-8000', Radeon 8060S (40 graphics cores), Default TDP 55W, cTDP 45-120W, codename Strix Halo, form factor Laptops and Desktops. Memory options: AMD blog says 'system memory options ranging from 32GB all the way up to 128GB of unified memory - out of which up to 96GB can be converted to VRAM through AMD Variable Graphics Memory' (no explicit list, so memory_options_gb null). GPU memory: CES 2025 press release 'up to 128GB of unified memory with up to 96GB available for graphics'; VGM FAQ: 'the iGPU (set to 96GB VGM) can technically access a total graphics memory size of 112GB on AMD Ryzen AI MAX+ powered 128GB systems'; on 64 GB systems the FAQ cites 'up to 48GB dedicated graphics memory through VGM'. Release year: CES press release Jan 6, 2025 (systems 'starting in Q1 2025'). AMD also sells the Ryzen AI Halo developer mini-PC with this chip (128GB). AMD additionally lists Ryzen AI Max+ 392 and 388 with the same Radeon 8060S, 128 GB max and LPDDR5x-8000 (not given separate records). Linux: ROCm RDNA3.5 optimization doc says GTT (GPU-mappable system RAM) defaults to 'approximately 50 percent of total system RAM' and can be raised via the TTM page limit (example sets 100 GB on a ~128 GB system); it recommends a small BIOS VRAM reservation (e.g. 0.5 GB) plus a larger GTT limit. AMD states a range of memory sizes up to 128 GB, not a list, so only the maximum is offered here; for a smaller system, enter it as an 'Other' unified-memory device with 256 GB/s. Second source: https://www.amd.com/en/blogs/2025/amd-ryzen-ai-max-395-pr....
NVIDIA DGX Spark
Product page spec table: 'System Memory 64 GB LPDDR5X* or 128 GB LPDDR5x, coherent unified system memory', 'Memory Interface 256-bit', 'Memory Bandwidth 273 GB/s'. Datasheet: 'Memory Bandwidth Up to 273 GB/s'. Per-pin data rate not stated. Power: Power supply / power consumption of the whole DGX Spark, as NVIDIA lists it. Unified CPU+GPU memory. Product page: 'Power Supply 240 Watts', 'GB10 TDP 140 W' (TDP of the GB10 chip incl. CPU and GPU); datasheet: 'Power Consumption 240 W'. 128 GB model shipped Oct 15, 2025 (NVIDIA newsroom, seen via search result). '*64 GB memory configuration is available exclusively through participating OEM partners' - NVIDIA blog dated October 2, 2026 (https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/): 64GB available from Friday, Oct. 23 (2026). Second source: https://dam-cdn.nvd.orangelogic.com/AssetLink/es6d60li4v5....
AMD Ryzen AI Max+ PRO 495 (Radeon 8065S)
Derived: LPDDR5x-8533 (8533 MT/s, product page 'Max Memory Speed') x 256-bit (product page) / 8 / 1000 = 273.056 GB/s. No AMD-stated bandwidth found. Power: Top of AMD's configurable TDP range (45-120 W; default 55 W): the chip only, and the computer maker sets it. ROCm on Linux: supported (ROCm 10.1.0 compatibility matrix, gfx1151; Ubuntu 26.04.1 and 24.04.5 (HWE 7.0) with inbox kernel driver; RHEL not offered for APUs). Vulkan: llama.cpp Vulkan backend. llama.cpp HIP backend; llama.cpp build.md: on Linux set GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 to use UMA with an integrated GPU. AMD blogs show LM Studio (llama.cpp runtime) on Windows with Variable Graphics Memory. Newer large-memory APU (codename Gorgon Halo). Product page: 'System Memory Type 256-bit LPDDR5x', 'Max. Memory 192 GB', 'Max Memory Speed LPDDR5x-8533', Radeon 8065S (40 graphics cores), Default TDP 55W, cTDP 45-120W. AMD blog May 20, 2026: Ryzen AI Max PRO 400 Series feature 'up to 192GB of unified memory and 160GB of VRAM'; footnote GRHP-01: 'As of 5/11/2026, the AMD Ryzen AI Max+ 495 PRO processor supports up to 160 GB dedicated graphics memory'. release_year 2026 = unveiling date of that blog; product page shows no launch date, and the AMD Ryzen AI Halo page lists the 495 version as 'Available Soon'. Linux: ROCm RDNA3.5 optimization doc says GTT (GPU-mappable system RAM) defaults to 'approximately 50 percent of total system RAM' and can be raised via the TTM page limit (example sets 100 GB on a ~128 GB system); it recommends a small BIOS VRAM reservation (e.g. 0.5 GB) plus a larger GTT limit. AMD states a range of memory sizes up to 192 GB, not a list, so only the maximum is offered here; for a smaller system, enter it as an 'Other' unified-memory device with 273.1 GB/s. Second source: https://www.amd.com/en/blogs/2026/amd-powers-next-generat....
Software support: CUDA, ROCm, Vulkan, SYCL and Metal (sources)

NVIDIA cards run every major engine through CUDA. AMD cards run llama.cpp through ROCm (HIP) where AMD lists the card, or through Vulkan; Intel Arc cards through SYCL or Vulkan; Macs through Metal (llama.cpp) and MLX. What AMD's and llama.cpp's own documents say:

Vulkan
README 'Supported backends' table: 'Vulkan | GPU' (vendor-neutral). Description: 'Vulkan and SYCL backend support'. Build instructions in docs/build.md#vulkan for Windows and Linux. Source.
HIP (ROCm)
README table: 'HIP | AMD GPU'. build.md: 'This provides GPU acceleration on HIP-supported AMD GPUs. Make sure to have ROCm installed.' Linux and Windows CMake examples; 'gfx1100 ... corresponds to Radeon RX 7900XTX/XT/GRE'. On Linux, GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 enables UMA for integrated GPUs. Source.
SYCL
README table: 'SYCL | Intel GPU'. SYCL.md: supports Intel Data Center Max, Flex, Arc series, built-in Arc GPU and iGPUs; 'Verified devices' table: 'Intel Arc B-Series | Support | Arc B580'; OS: Linux (Ubuntu 22.04, Fedora Silverblue 39, Arch) and Windows 11 supported. Source.
OpenVINO
README table: 'OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs'. OPENVINO.md: work in progress; 'validated specifically on AI PCs such as the Intel Core Ultra Series 1 and Series 2' (Arc B-series not specifically validated). Source.
ROCm
ROCm Core SDK 10.1.0 (released October 5, 2026) is the current ROCm release; its compatibility matrix 'Applies to Linux and Windows' and lists Radeon GPUs: RX 9070 XT/GRE/9070, RX 9060 XT LP/9060 XT/9060, RX 9050 and RX 9050 (4GB), RX 7900 XTX/XT/GRE, RX 7800 XT, RX 7700 XT, RX 7700, RX 7600, PRO W7900 (and Dual Slot), W7800 48GB, W7800, W7700, AI PRO R9700S/R9700/R9600D/R9600, plus Ryzen AI Max/Max+ (gfx1151). RX 7600 XT is not listed. Per its OS selector, R9700S/R9700/RX 9070-series/9060 XT/9060/W7900/W7800/RX 7900-series/7800 XT offer Ubuntu, RHEL, Windows and WSL2; R9600D/R9600/9060 XT LP/9050/7700 XT/7700/7600 offer Ubuntu, RHEL and Windows (no WSL2); Ryzen AI Max offers Ubuntu, Windows, WSL2. Source.
ROCm
ROCm release history: 10.1.0 Oct 5, 2026; 10.0.0 Aug 26, 2026; 7.14.1 Sept 2, 2026; 7.14.0 July 15, 2026; 7.2.4 May 29, 2026. Two version lines are current, so a 'ROCm version' needs the stream stated. Source.
ROCm
ROCm 10.1.0 release notes: 'Since ROCm 7.14, ROCm uses TheRock as its build and release system.' 'ROCm 10.1.0 adds support for AMD Radeon AI PRO R9600 GPUs.' 'If your GPU is not listed, it might be community-enabled through TheRock nightly builds.' Resolved issue: PyTorch training/fine-tuning (Llama-Factory, Unsloth) could cause GPU resets or crashes on 'Radeon RX 9070 Series and Radeon AI PRO R9700'. Source.
ROCm
ROCm 7.14.0 Linux system requirements (page dated 2026-07-15, now marked 'This page has moved'): 'If a GPU is not listed on this table, it's not officially supported by AMD.' Radeon list excludes RX 7600, RX 7600 XT, RX 9050, AI PRO R9700S and R9600. Footnote: listed Radeon/Radeon PRO GPUs 'only support Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7.' Source.
ROCm
HIP SDK for Windows 7.2.0 system requirements (page dated 2026-07-31): 'If a GPU is not listed on this table, it is not officially supported by AMD.' Lists Runtime+HIP SDK support for RX 9070 XT, 9070, 9070 GRE, 9060 XT, 9060, 7900 XTX, 7900 XT, 7800 XT, 7700 XT, 7650 GRE, 7600 XT, 7600, PRO W7900 (and Dual Slot), W7800, W7700, AI PRO R9700, and Ryzen AI Max+/Max 395/390/385 (and PRO), plus 'AMD Ryzen AI MAX+ 495'. RX 7900 GRE, W7800 48GB, R9700S, R9600/R9600D, RX 9060 XT LP and RX 9050 are absent. ROCm Debugger is unsupported on the RDNA 3 entries. Source.
ROCm
ROCm GPU specifications page lists VRAM 'Dynamic + carveout' for Ryzen APUs and gives RX 9070 GRE VRAM as 16 GiB, which conflicts with AMD's product page (12 GB). Source.
ROCm
ROCm RDNA3.5 system optimization (Strix Halo): GTT defaults to approximately 50 percent of system RAM; the limit can be raised via the kernel TTM page limit (amd-ttm --set 100 example); a small BIOS VRAM reservation plus a larger GTT is recommended. Source.
ROCm
llama.cpp build.md (HIP): 'If your GPU is not officially supported you can use the environment variable HSA_OVERRIDE_GFX_VERSION set to a similar GPU ... 11.0.0 on RDNA3. Note that HSA_OVERRIDE_GFX_VERSION is not supported on Windows'. That is a community workaround, not AMD support. Source.

How much unified memory the GPU may use

Apple Silicon. macOS lets the GPU use only part of unified memory by default, and Apple documents no fixed rule. Metal reports the limit at run time as recommendedMaxWorkingSetSize, "an approximation of how much memory, in bytes, this GPU device can allocate without affecting its runtime performance", and llama.cpp and Ollama read it as the Mac's GPU memory. Apple's 2021 tech talk on Metal compute gives two examples: an M1 Pro or M1 Max with 32 GB lets the GPU access 21 GB, and an M1 Max with 64 GB, 48 GB. The calculator follows those examples: two thirds of memory up to 32 GB, three quarters above; enter your own share if your Mac reports a different limit. Apple's MLX project documents raising the limit with sudo sysctl iogpu.wired_limit_mb=<megabytes>, to a value larger than the model but smaller than the machine's memory. Community reports say the setting does not survive a reboot, and an MLX maintainer linked a kernel panic to a limit set too high, so leave macOS several gigabytes. Apple: recommendedMaxWorkingSetSize. Apple tech talk: Metal Compute on MacBook Pro. mlx-lm README: large models. MLX: set_wired_limit. mlx-lm issue 883 (kernel panic). llama.cpp discussion 2182 (community reports).

AMD Ryzen AI Max. AMD says up to 96 GB of a 128 GB Ryzen AI Max system can be assigned to graphics with Variable Graphics Memory, and up to 160 GB of 192 GB on the Ryzen AI Max+ PRO 495; the calculator uses those shares (75% and 83%). On Linux the GPU can also map system memory through GTT, which AMD's ROCm guide says defaults to about half of RAM and can be raised, so the usable share depends on how the machine is set up. Enter your own share if you have changed it. AMD CES 2025 press release. AMD blog: Ryzen AI Max+ 395. AMD blog: Ryzen AI Max PRO 400. ROCm: RDNA 3.5 system optimization.

NVIDIA DGX Spark and RTX Spark PCs. NVIDIA documents no separate GPU limit for the DGX Spark's unified LPDDR5X (128 GB, and 64 GB in versions sold through other makers), so the calculator lets the GPU use all of it except the RAM kept for the operating system. NVIDIA publishes no memory bandwidth for RTX Spark PCs, so the calculator shows no speed ceiling for them. NVIDIA DGX Spark. NVIDIA RTX Spark.

Weight formats and KV-cache types

Bits per weight decide the file size and the bytes read for every token. GGUF figures are llama.cpp's effective bits per weight for a whole model file, which mixes tensor types; other formats are sized like the VRAM calculator, with embeddings and the LM head kept at 16-bit. Raw data: local-ai-assumptions.json.

FormatUsed byBits per weightBasisSource
Q2_KGGUF (llama.cpp, Ollama, LM Studio)3.159Effective bits per weight of a whole Q2_K file of Llama 3.1 8B in llama.cpp's table (2.95 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent.llama.cpp quantize README
Q3_K_SGGUF (llama.cpp, Ollama, LM Studio)3.643Effective bits per weight of a whole Q3_K_S file of Llama 3.1 8B in llama.cpp's table (3.41 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent.llama.cpp quantize README
Q3_K_MGGUF (llama.cpp, Ollama, LM Studio)3.996Effective bits per weight of a whole Q3_K_M file of Llama 3.1 8B in llama.cpp's table (3.74 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent.llama.cpp quantize README
IQ4_XSGGUF (llama.cpp, Ollama, LM Studio)4.46Effective bits per weight of a whole IQ4_XS file of Llama 3.1 8B in llama.cpp's table (4.17 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent.llama.cpp quantize README
Q4_K_SGGUF (llama.cpp, Ollama, LM Studio)4.667Effective bits per weight of a whole Q4_K_S file of Llama 3.1 8B in llama.cpp's table (4.36 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent.llama.cpp quantize README
Q4_K_MGGUF (llama.cpp, Ollama, LM Studio)4.894Effective bits per weight of a whole Q4_K_M file of Llama 3.1 8B in llama.cpp's table (4.58 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. Q4_K itself is 4.5 bits (144 bytes per 256 weights); Q4_K_M stores half of the attention V and FFN down tensors, and the output tensor, in Q6_K.llama.cpp quantize README
Q5_K_SGGUF (llama.cpp, Ollama, LM Studio)5.57Effective bits per weight of a whole Q5_K_S file of Llama 3.1 8B in llama.cpp's table (5.21 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent.llama.cpp quantize README
Q5_K_MGGUF (llama.cpp, Ollama, LM Studio)5.704Effective bits per weight of a whole Q5_K_M file of Llama 3.1 8B in llama.cpp's table (5.33 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent.llama.cpp quantize README
Q6_KGGUF (llama.cpp, Ollama, LM Studio)6.563Effective bits per weight of a whole Q6_K file of Llama 3.1 8B in llama.cpp's table (6.14 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent.llama.cpp quantize README
Q8_0GGUF (llama.cpp, Ollama, LM Studio)8.501Effective bits per weight of a whole Q8_0 file of Llama 3.1 8B in llama.cpp's table (7.95 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. Q8_0 itself is 8.5 bits (34 bytes per 32 weights).llama.cpp quantize README
AWQ / GPTQ 4-bit (group size 128)GPU engines (vLLM, SGLang)4.1564-bit weights plus a 16-bit scale and a 4-bit zero point per group of 128 (4.156 bits); embeddings and LM head kept at 16-bit, as the VRAM calculator counts them. Group size 128 is AutoAWQ's and GPTQModel's default.AutoAWQ defaults
FP8 (8-bit float)GPU engines (vLLM, SGLang)81 byte per weight; embeddings and LM head kept at 16-bit. Fast only on GPUs with FP8 tensor cores (NVIDIA Ada and newer).vLLM FP8 docs
MXFP4 (how gpt-oss is published)As published4.254-bit values plus one 8-bit shared scale per block of 32 (17 bytes per 32 weights = 4.25 bits); the tensors the checkpoint keeps in BF16 stay at 16-bit.ggml block layout
F16 / BF16 (unquantized)Unquantized162 bytes per weight. llama.cpp's table lists F16 at 16.0005 bits for the whole file.llama.cpp quantize README
KV-cache typeBytes per valueBasisSource
F16 (llama.cpp default) / BF16216-bit keys and values; llama.cpp's default for --cache-type-k and --cache-type-v.llama.cpp server README
q8_0 (llama.cpp)1.06232 int8 values and one 16-bit scale per block: 34 bytes per 32 values. A quantized V cache needs flash attention, which llama.cpp turns on by default (-fa auto).ggml block layout
q4_0 (llama.cpp)0.562532 4-bit values and one 16-bit scale per block: 18 bytes per 32 values. Needs flash attention for the V cache, as above.ggml block layout
FP8 (vLLM --kv-cache-dtype fp8)11 byte per value; the VRAM calculator's layout, including the FP8 layout vLLM uses for latent-attention (MLA) caches.vLLM quantized KV cache docs

System RAM bandwidth presets

Peak bandwidth is the transfer rate in MT/s times 8 bytes per 64-bit channel times the number of channels: DDR5-5600 on two channels is 89.6 GB/s, the figure Intel states for its Core i9-14900K. Current desktop CPUs from AMD and Intel have two channels (Ryzen 9 9950X up to DDR5-5600; Core Ultra 9 285K up to DDR5-6400; Core Ultra 7 270K Plus up to DDR5-7200). This is a peak: Crucial's table of effective memory bandwidth gives 69.21 GB/s for DDR5-5600, about 77% of it, so real offload speed sits further below the ceiling. Crucial: DDR5 bandwidth. Intel Core i9-14900K specifications. AMD Ryzen 9 9950X. Intel Core Ultra 9 285K. Intel Core Ultra 7 270K Plus.

MemoryTransfer rateChannelsPeak bandwidth
DDR4-3200, two channels3,200 MT/s251.2 GB/s
DDR5-4800, two channels4,800 MT/s276.8 GB/s
DDR5-5600, two channels5,600 MT/s289.6 GB/s
DDR5-6400, two channels6,400 MT/s2102.4 GB/s
DDR5-7200, two channels7,200 MT/s2115.2 GB/s

How the answer is worked out

weights          = parameters x bits per weight / 8
KV cache         = 2 x layers x KV heads x head dim x bytes per value x context tokens
total            = weights + KV cache + runtime allowance
fits on the GPU  when total - token embeddings <= GPU memory
                 (unified memory: memory x the share the GPU may use)
partial offload  layers on GPU = floor((GPU memory - allowance - LM head) / (one layer's weights + its cache))
decode ceiling   = memory bandwidth / bytes read per generated token
bytes per token  = weights x (active / total parameters) + the whole KV cache
offload ceiling  = 1 / (bytes in VRAM / GPU bandwidth + bytes in RAM / RAM bandwidth)

The memory arithmetic for the cache is the VRAM calculator's own code, imported rather than copied, so grouped-query attention, sliding windows, linear attention and latent attention (MLA) are counted the same way on both pages. Model shapes come from each model's config.json.

Two details follow llama.cpp's source code. It keeps the token-embedding table in system RAM ("there is very little benefit to offloading the input layer, so always keep it on the CPU"), so that table never counts against VRAM and only one row of it is read per token. And when you offload with -ngl, the output layer (the LM head) goes to the GPU first, then whole layers. The sources are listed in the data file.

Bits per weight: what a quantization costs

A GGUF file is not "parameters x 4 bits". The k-quant mixes keep some tensors at higher precision, and every block stores scales. llama.cpp's own quantization README lists the bits per weight of each file type, measured on Llama 3.1 8B, and the calculator uses those figures for every model:

Format Bits per weight Llama 3.1 8B weights Llama 3.3 70B weights
Q2_K 3.16 2.95 GiB 25.95 GiB
Q3_K_M 4.00 3.74 GiB 32.82 GiB
Q4_K_M 4.89 4.58 GiB 40.2 GiB
Q5_K_M 5.70 5.33 GiB 46.85 GiB
Q6_K 6.56 6.14 GiB 53.91 GiB
Q8_0 8.50 7.95 GiB 69.82 GiB
F16 / BF16 (unquantized) 16.00 14.96 GiB 131.42 GiB

So Q4_K_M costs 4.89 bits per weight, not 4. Other architectures land within a few percent of these figures, because the mixes are decided tensor by tensor. AWQ and GPTQ 4-bit checkpoints (group size 128) and FP8 are sized like the VRAM calculator sizes them: the quantized weights plus their scales, with the embeddings and LM head left at 16-bit. gpt-oss is published in MXFP4, 4.25 bits per weight for the expert weights. The format table under the calculator has the basis and source for each.

The KV cache has its own precision. llama.cpp's default is F16. Its q8_0 cache type takes 34 bytes per 32 values and q4_0 18 bytes per 32, and a quantized V cache needs flash attention, which current llama.cpp turns on by default (-fa auto). Whether a quantized cache hurts answers depends on the model; check it on your own prompts.

Partial offload: why the speed falls off a cliff

When a model is bigger than VRAM, llama.cpp can keep some layers on the GPU and run the rest on the CPU from system RAM. That works, but every generated token still has to read every layer once, and system RAM is far slower than VRAM. Two-channel desktop memory peaks at 89.6 GB/s with DDR5-5600, against 1008 GB/s on an RTX 4090.

Llama 3.3 70B at Q4_K_M needs 43.7 GiB at an 8k context. On one 24 GB card, 43 of its 80 layers fit. The other 37 run from system RAM, and reading them takes 90% of each token's time. The ceiling moves with the RAM, not the GPU:

System RAM Peak bandwidth Ceiling, Llama 3.3 70B Q4_K_M, 8k context
DDR4-3200, two channels 51.2 GB/s 2.3 tokens/s
DDR5-4800, two channels 76.8 GB/s 3.4 tokens/s
DDR5-5600, two channels 89.6 GB/s 3.9 tokens/s
DDR5-6400, two channels 102.4 GB/s 4.4 tokens/s
DDR5-7200, two channels 115.2 GB/s 4.9 tokens/s

With two 24 GB cards the same model fits entirely in VRAM (44.7 GiB of 48), and the ceiling on two RTX 3090s rises to 20.7 tokens per second. A second card adds memory, not bandwidth: llama.cpp's default layer split passes each token through the cards one after another, so the ceiling is that of one card reading all the bytes. These figures are peaks. Crucial's table of effective DDR5-5600 bandwidth gives 69.21 GB/s, about 77% of the peak, and inference engines lose more on top.

Mixture-of-experts models are the exception worth knowing. Each token reads only the active experts, so a model like gpt-oss-20b or Qwen3 30B-A3B stays usable with part of it in RAM, and llama.cpp can keep just the expert weights in RAM (--n-cpu-moe or --override-tensor) while attention stays on the GPU. That usually beats offloading whole layers. The calculator models whole-layer offload only, so for MoE models it is the cautious estimate.

Unified memory: Macs, Ryzen AI Max and DGX Spark

On a Mac, a Ryzen AI Max mini-PC or a DGX Spark, the CPU and GPU share one pool of memory, so there is no separate VRAM to overflow. The question becomes how much of that pool the GPU may use.

When a model fits in memory but not in the GPU's share, the calculator says "over the default GPU limit" instead of "partial offload", and shows the ceiling you would get after raising the limit.

A worked example

Llama 3.1 8B at Q4_K_M with an 8k context on an RTX 4090:

That is arithmetic, not a measurement. Real engines stay below the bandwidth ceiling, and how far below depends on the engine, the kernels and the GPU. Prompt processing (prefill) is limited by compute rather than bandwidth and is not estimated here. Measured speeds on real hardware are planned, with the scripts and raw results to be published alongside them.

What the calculator does not model

Where the numbers come from

Memory sizes, bandwidth and power are read from each vendor's own spec pages, datasheets and whitepapers, linked row by row in the hardware table above. Where a vendor does not publish a value, the table says "n/a" and the calculator shows no ceiling rather than a guess: NVIDIA publishes no memory bandwidth for the RTX 3080 12 GB or for RTX Spark PCs. Model shapes come from Hugging Face config.json files, and the bits-per-weight figures from llama.cpp's repository. The methodology lists every source and assumption, and the VRAM guide applies the calculator to the models people ask about most.

Also available as Markdown.