# GPUCostLab > GPU and inference costs, worked out by someone who builds inference engines. Written by one author, a Principal ML engineer (10+ years in machine learning, from research to production; Former Staff ML Engineer working on document extraction and OCR; Former Lead ML Platform Engineer for a 5,000-stream computer-vision platform; Contributor to vLLM; creator of Videoflow (1,000+ GitHub stars)). The tools run in the browser and never upload files. Every page has a Markdown version at the same path with .md appended. ## Tools - [Can my computer run this LLM? Memory, CPU offload and speed for GPUs, Macs and mini-PCs](https://gpucostlab.com/can-i-run-this-llm.md): Pick your graphics card, Mac or mini-PC, a model and a quantization (GGUF Q2_K to Q8_0, AWQ, FP8 or MXFP4). See whether it fits in GPU memory, how many layers spill into system RAM, and the most tokens per second your memory bandwidth allows. Specs come from the vendors' own pages, and the arithmetic is on the page. - [GPU cloud prices per GPU-hour: H100, H200, B200, A100, L40S, RTX 4090 and more, with AWS, Google Cloud and Azure](https://gpucostlab.com/gpu-cloud-prices.md): On-demand list prices per GPU-hour from the public pricing pages of RunPod, Lambda, DigitalOcean, CoreWeave, Hyperstack, Nebius, Modal, Crusoe, JarvisLabs, Voltage Park and others, next to AWS, Google Cloud and Azure GPU instances. Sort and filter them, and follow each row to its source. Prices that are not stated stay blank. - [LLM fine-tuning cost estimator: LoRA, QLoRA or full fine-tune, memory and GPU-hours](https://gpucostlab.com/llm-fine-tuning-cost.md): Memory per GPU for full fine-tuning, LoRA and QLoRA, whether it fits on the GPUs you pick, the GPU-hours your tokens and epochs take at a stated throughput assumption, and the cost at on-demand list prices, including the cheapest listing for that GPU. Formulas and sources are on the page. - [LLM VRAM calculator: weights, KV cache, and which GPUs fit](https://gpucostlab.com/llm-vram-calculator.md): How much GPU memory a model needs at your context length and concurrency. Weights at BF16, FP8, INT8, INT4 or FP4, KV cache per token and in total, and which GPUs fit, alone or with tensor parallelism. The formula is on the page, and model shapes come from each model's config.json. - [What does it cost to run an LLM at home? Electricity calculator with EIA prices](https://gpucostlab.com/local-llm-electricity-cost.md): Monthly kWh and electricity cost of running a local LLM on your GPU, Mac or mini-PC, the cost per million generated tokens, and the same tokens priced at an API's list price, with the hardware you bought spread over the months you choose. Electricity prices by state from the U.S. Energy Information Administration. - [Self-hosting an LLM vs paying for an API: break-even calculator](https://gpucostlab.com/self-host-vs-api-cost.md): Monthly cost of serving your token volume on rented GPUs versus an API, the volume where they cross, and the throughput self-hosting would need to match the API. API and GPU list prices are linked to their sources and every input is editable. ## Guides - [Apple Silicon vs NVIDIA for running LLMs at home: memory against bandwidth](https://gpucostlab.com/apple-silicon-vs-nvidia-for-local-ai.md): A Mac holds far larger models than any consumer graphics card, and a high-end NVIDIA card generates faster on the models it can hold. This guide works out where the line falls for every M1 to M6 chip against the RTX 4090, RTX 5090, RTX PRO 6000, DGX Spark and Ryzen AI Max, from vendor specs and the calculator's arithmetic. - [Best GPU for running LLMs at home, at each budget tier](https://gpucostlab.com/best-gpu-for-local-llms.md): Graphics cards and unified-memory computers for running language models at home, grouped into budget tiers by memory size and ranked by memory bandwidth, with what each tier can run and its speed ceiling. Derived from vendor specifications and the calculator's arithmetic, not from benchmarks we have not run, and with no retailer prices. - [How much VRAM do you need to run Llama, Qwen or gpt-oss locally?](https://gpucostlab.com/how-much-vram-to-run-llms-locally.md): Memory needed by the open models people run at home, at every common GGUF quantization and at 8k, 32k and 128k tokens of context, worked out from each model's config.json and llama.cpp's bits-per-weight figures. Includes the smallest GPU or Mac memory size that holds each one. - [How GPUCostLab collects its data and what it assumes](https://gpucostlab.com/methodology.md): Where every number on GPUCostLab comes from, how often it is re-read, which assumptions the calculators make (including the Run AI at home calculators), and the measured benchmarks that are planned. No benchmark results are published yet. - [Run AI at home: will the model fit, how fast can it go, and what will it cost?](https://gpucostlab.com/run-ai-at-home.md): Free calculators and guides for running language models on your own GPU, Mac or mini-PC. Whether a model fits, how many layers spill into system RAM, the speed ceiling your memory bandwidth allows, and what the electricity costs against an API. Built from vendor specifications and the models' own configs, with nothing presented as a benchmark. ## Optional - [About GPUCostLab and its author](https://gpucostlab.com/about.md): Who writes GPUCostLab, the credentials behind it, how the site is funded, and how to report a wrong price or specification. - [Affiliate disclosure](https://gpucostlab.com/affiliate-disclosure.md): How GPUCostLab may earn money from referral and affiliate links, how those links are labeled, the rules that keep commissions out of prices and rankings, and the current list of relationships. - [Privacy](https://gpucostlab.com/privacy.md): GPUCostLab's calculators run in your browser and send nothing anywhere. The site sets no cookies and has no accounts. Analytics, if switched on, is cookie-free.