# How GPUCostLab collects its data and what it assumes

> Where every number on GPUCostLab comes from, how often it is re-read, which assumptions the calculators make (including the Run AI at home calculators), and the measured benchmarks that are planned. No benchmark results are published yet.

By GPUCostLab. Updated 7 October 2026. Canonical URL: https://gpucostlab.com/methodology/

## Rules

1. **Every number has a source and a date.** Prices, specifications and model shapes link to the page they were read from, with the date they were read.
2. **Unknown stays unknown.** When a page does not state a value, the site shows "Not stated", "Check provider" or "n/a". Values are never filled in from memory, from third-party aggregators, or by guessing a form factor from a product name.
3. **Formulas are on the page.** The calculators show their arithmetic with your numbers substituted, and their code is tested before every deploy.
4. **Assumptions are inputs.** Anything that is not arithmetic or a sourced fact, such as runtime memory overhead, throughput or utilization, is an editable field labeled as an assumption.
5. **Measured means measured.** No throughput, latency or cost-per-token result appears on this site until it has been measured with the method below and published with its data and scripts.

## Data sources

### GPU prices

[prices.json](/tools/gpu-price-table/prices.json) holds on-demand list prices per GPU-hour from the public pricing pages of RunPod, Lambda, DigitalOcean (GPU Droplets and Paperspace), Vast.ai, CoreWeave, Hyperstack, Nebius, Modal, Crusoe, JarvisLabs and Voltage Park, and from the public price files of AWS, Google Cloud and Azure.

- **How it is read.** `data/gpu/update_prices.py` fetches each pricing page once per run, after checking the site's robots.txt. It uses no logins and no private APIs. A parser per provider reads the prices from the page's own HTML or its structured data (RunPod publishes schema.org offers).
- **What is recorded.** On-demand prices only. Instance prices are divided by the GPU count, and the note keeps the original. Reserved, spot, committed and promotional prices are noted and left out of the numbers.
- **Cadence.** Weekly. The script prints every price that changed, appeared or disappeared. When a page fails or stops parsing, the previous prices are kept with their original dates and the failure is reported, so a stale price never looks fresh.
- **Not recorded.** Vast.ai's marketplace rates are loaded in the browser, and the HTML of its public pricing pages (the main page and the per-GPU page for the H100 SXM) contained no prices on 7 October 2026.

#### AWS, Google Cloud and Azure

The hyperscalers publish machine-readable prices that need no account, so the script reads those instead of scraping their pages. One representative US region each, on-demand Linux prices only, for a fixed list of GPU instance types (the single-GPU size where one exists and the full-machine size):

| Provider | Region | Source read | GPU names from |
|---|---|---|---|
| AWS | us-east-1 | The [AWS Price List](https://docs.aws.amazon.com/awsaccountbilling/latest/aboutv2/using-the-aws-price-list-bulk-api-fetching-price-list-files-manually.html) file for EC2 in us-east-1 (CSV, about 300 MB), streamed once per run and reduced to the rows for the listed instance types. Each row links to the exact file version read. | AWS's instance-type pages (P6, P5, P4, G6e, G6, G5) |
| Google Cloud | us-central1 | The [accelerator-optimized pricing page](https://cloud.google.com/products/compute/pricing/accelerator-optimized), whose tables show its default region, us-central1; the script checks the region name on every read. Google's Cloud Billing Catalog API needs an API key from a Google Cloud project, so it is not used. | Google's [accelerator-optimized machine documentation](https://cloud.google.com/compute/docs/accelerator-optimized-machines) |
| Azure | eastus | The [Azure Retail Prices API](https://learn.microsoft.com/en-us/rest/api/cost-management/retail-prices/azure-retail-prices) (prices.azure.com), Linux pay-as-you-go rows only. A size the API lists no eastus price for is read in eastus2, and its row says so (ND H200 v5 in October 2026). Each row links to the API query for that size and region. | Azure's GPU size-series documentation |

- **What a row means.** The whole-instance hourly price divided by the GPU count. The note on each row gives the instance price, vCPUs and memory where the source states them, and the price list's effective date.
- **Left null, with the reason in the note.** AWS's p5e.48xlarge has no on-demand Linux row in the us-east-1 price list; Google Cloud's a4-highgpu-8g (B200) shows "N/A" as its on-demand price. Spot, reserved, savings-plan, committed-use, Capacity Block and DWS prices on the same sources are not recorded.
- **Form factor.** Recorded only where the hyperscaler's documentation states it: SXM for Google Cloud's A3 and A3 Ultra, NVL for Azure's NCads H100 v5, PCIe for Azure's NC A100 v4. AWS does not name it for P5 and P5en, and Azure does not for ND H100 v5 and ND H200 v5, so those rows record only the GPU family.
- **Why these prices look high.** The instances bundle vCPUs, system memory, local NVMe and fast networking, and most hyperscaler GPU capacity is bought with commitments. The price table says so next to the comparison.

### API prices

[api-prices.json](/tools/self-host-vs-api-break-even/api-prices.json) holds standard-tier list prices per million tokens from the pricing pages or pricing docs of OpenAI, Anthropic, Google, DeepSeek, Mistral, Together AI, Fireworks AI, DeepInfra, Groq, Hyperstack and Crusoe.

- **Cadence.** Re-read by hand monthly, and whenever a provider announces a change.
- **What is recorded.** Batch, priority, flex, off-peak and long-context tiers go in the notes, not the numbers. Promotional prices are recorded as today's price, with the regular price and the end date in the note.
- **Substitute pages.** Where a provider's main pricing page could not be used, the price table's notes say which page was used instead and why. For example, openai.com returned HTTP 403 to our request, so OpenAI's API documentation pricing page was used.

### GPU specifications

[gpus.json](/tools/llm-vram-calculator/gpus.json) holds GPU specifications from NVIDIA's and AMD's datasheets, product pages and architecture whitepapers.

- **Compute figures.** These are dense tensor-core TFLOPS. When a vendor lists only the figure "with sparsity", it is halved, and the note says so.
- **GeForce cards.** For these, the FP16 figure is the FP32-accumulate rate, which is what BF16 and FP16 inference kernels use.
- **Conflicting sources.** Where the vendor's own documents disagree (B300 memory and bandwidth, for example), the value used and the conflicting one are both in the note.

### Model shapes

[models.json](/tools/llm-vram-calculator/models.json) holds model shapes copied by script from each model's `config.json` on Hugging Face.

- **Parameter counts.** These come from the Hugging Face safetensors metadata: what the checkpoint holds, including any vision encoder or multi-token-prediction layer.
- **Gated repositories.** For Meta's Llama models, the shapes come from an ungated copy and the parameter count from Meta's own repository. The note on each preset says so.
- **Engine layouts.** Some facts come from the inference engine, not the model: vLLM's default memory utilization, KV-head replication under tensor parallelism, and the FP8 layout of the MLA and indexer caches. Each is sourced to vLLM's code or documentation.

### Home hardware (Run AI at home)

[consumer-hardware.json](/tools/local-ai-can-i-run/consumer-hardware.json) holds 94 graphics cards, Apple Silicon chips and unified-memory PCs: NVIDIA GeForce RTX 30, 40 and 50 series and RTX workstation cards, AMD Radeon RX 7000 and 9000 series and Radeon PRO and AI PRO cards, Intel Arc B-series, Apple M1 to M6 chips, AMD Ryzen AI Max and NVIDIA DGX Spark and RTX Spark.

- **Read from the vendor only.** Memory, bandwidth and power come from NVIDIA's, AMD's, Intel's and Apple's own product pages, spec pages, datasheets, whitepapers, launch articles and press releases, never from third-party databases. Each row links to its page.
- **Bandwidth.** Recorded as the vendor states it. Where a vendor states only the memory's data rate and bus width, bandwidth is their product divided by 8, marked "derived", with the inputs in the note (AMD Ryzen AI Max 390 and 385, and the Ryzen AI Max+ PRO 495). NVIDIA publishes no bandwidth for the RTX 3080 12 GB or for RTX Spark PCs, so those rows have none and the calculator shows no speed ceiling for them.
- **Power.** Recorded under the vendor's own name for it: Total Graphics Power or Graphics Card Power (GeForce), Total Board Power (AMD and Intel), Max Power Consumption (NVIDIA workstation cards), the whole computer's maximum at the wall (Apple's desktop Macs, for the one configuration Apple measured), the power figure of the whole DGX Spark, and the top of AMD's configurable TDP range for Ryzen AI Max, which covers the chip only. Laptop-only Apple chips, and the 40-core-GPU versions of M4 Max and M5 Max, have no Apple power figure.
- **Software support.** AMD's ROCm compatibility matrix (ROCm 10.1.0 when read) and HIP SDK for Windows tables, and llama.cpp's backend documentation for Vulkan, HIP and SYCL. The two AMD documents disagree on Windows support for some cards; both readings are in the notes.
- **Unified memory.** Apple documents no rule for how much memory the GPU may use; the default follows the two examples in Apple's 2021 Metal tech talk (two thirds up to 32 GB, three quarters above). AMD's stated Variable Graphics Memory maximums are used for Ryzen AI Max. Where a vendor states no limit (DGX Spark, RTX Spark), the GPU is assumed to use all memory except the RAM kept for the operating system.
- **Conflicts.** Where a vendor's own pages disagree (the RX 7900 XT's board power, the RX 9070 GRE's memory size, the M3 Ultra's 512 GB option), the value used and the other one are both in the note.

### Electricity prices

[electricity-prices.json](/tools/local-ai-electricity-cost/electricity-prices.json) holds the average residential electricity price for each state, DC and the U.S. from the U.S. Energy Information Administration's Electric Power Monthly, Tables 5.6.A (latest month) and 5.6.B (year to date), with the U.S. figure checked against Table 5.3. EIA marks these values as preliminary. They are re-read when a new Electric Power Monthly is released (monthly).

### Quantization formats

[local-ai-assumptions.json](/tools/local-ai-can-i-run/local-ai-assumptions.json) holds the bits per weight of each GGUF type from llama.cpp's quantization README (measured there on Llama 3.1 8B), the block layouts behind llama.cpp's KV-cache types, AutoAWQ's and GPTQModel's default group size, and the RAM bandwidth presets, each with its source and commit.

## Calculator assumptions

| Assumption | Value | Why |
|---|---|---|
| Share of GPU memory the engine may use | 0.92 | vLLM's default `gpu_memory_utilization` |
| Runtime allowance per GPU | 2 GiB | A round number for the CUDA context, CUDA graphs, NCCL buffers and peak activations. Editable; vLLM logs the real figures at startup. Per-model measurements are planned. |
| GPU memory | Marketed GB, treated as GiB | The driver can report slightly less, for example with ECC enabled |
| Quantized weights | Embeddings and LM head at 16-bit; group scales counted | How common AWQ, GPTQ, FP8 and MXFP4 checkpoints are stored |
| Sliding-window layers | Cache at most the window | vLLM V1's hybrid KV-cache manager |
| Linear-attention state | FP32 recurrent state, 16-bit convolution state | The models' `mamba_ssm_dtype` setting |
| Hours in a month | 730 | 8,760 hours a year / 12 |
| Throughput (break-even calculator) | No default you should keep | Placeholders only; measure your own |

### Run AI at home calculators

| Assumption | Value | Why |
|---|---|---|
| GGUF weight size | Parameters x llama.cpp's effective bits per weight for that type | The mixes keep some tensors at higher precision; figures measured on Llama 3.1 8B, so other architectures differ by a few percent |
| Token embeddings | In system RAM, read one row per token (GGUF) | llama.cpp always keeps the input layer on the CPU |
| Partial offload | LM head first, then whole layers, each with its share of the cache | llama.cpp's `-ngl` placement |
| Runtime allowance | 1 GiB per GPU | A round assumption for llama.cpp's compute buffers and the driver context; llama.cpp prints its buffer sizes at load. Editable |
| RAM kept for the OS and apps | 4 GiB | An assumption. Editable |
| System RAM bandwidth | DDR5-5600, two channels: 89.6 GB/s peak | Transfer rate x 8 bytes x channels; the top memory speed AMD lists for the Ryzen 9 9950X. Editable, with presets |
| Decode ceiling | Bandwidth / (active weights + whole KV cache) per token | An upper bound for one conversation; not a measurement |
| Several identical GPUs | Memory adds; ceiling as one GPU reading all the bytes | llama.cpp's default layer split runs the cards one after another |
| Rest of system | 100 W around a graphics card, 30 W around an APU chip figure, 0 W for whole-computer figures | An assumption; a plug-in meter replaces it |
| Days in a month | 365 / 12 | The same 730-hour month as the break-even calculator |

### Fine-tuning cost estimator

| Assumption | Value | Why |
|---|---|---|
| Full fine-tune memory | 2 + 2 + 12 bytes per parameter with mixed-precision AdamW; 8 or 2 bytes of optimizer state for the other optimizer choices | [ZeRO](https://arxiv.org/abs/1910.02054) (K = 12), [8-bit optimizers](https://arxiv.org/abs/2110.02861) |
| LoRA adapters | Rank r on q, k, v and o of every layer; FP32 weights, gradients and AdamW moments (16 bytes per adapter parameter) | [LoRA](https://arxiv.org/abs/2106.09685); PEFT keeps adapters in FP32 by default. The trainable count is editable. |
| QLoRA frozen weights | 4.127 bits per parameter; embeddings and LM head at 16-bit | NF4 plus double-quantized constants, [QLoRA](https://arxiv.org/abs/2305.14314) |
| Activations | 34 x s x b x h bytes per layer, or 2 x s x b x h per layer plus one recomputed layer with gradient checkpointing; FP32 logits | [Korthikanti et al. 2022](https://arxiv.org/abs/2205.05198), attention-score term dropped for FlashAttention-style kernels |
| Training FLOPs per token | 6N + 12 x L x H x Q x T; 4N for LoRA and QLoRA | [PaLM appendix B](https://arxiv.org/abs/2204.02311); adapters skip the frozen weights' gradients |
| MFU | 30% of peak dense BF16, editable | A rule of thumb below the 38-46% reported for large, heavily tuned pre-training runs (PaLM, Llama 3). Not a measurement; enter your own tokens per second. |
| Framework overhead per GPU | 2 GiB | Same round allowance as the VRAM calculator. Editable. |
| Scaling across GPUs | MFU unchanged as GPUs are added | A simplification; the result warns when GPUs have no NVLink |

## Benchmarks (planned, no results yet)

**Status: planned. Nothing below has been measured.** The method is published first so it cannot be bent to fit the results.

- **Throughput per GPU.** vLLM serving throughput (input and output tokens per second) for a fixed set of popular open models at stated precisions, on the GPUs in the price table. Fixed prompt and output length distributions and a sweep of concurrency levels. Engine version, flags and driver versions recorded.
- **Cost per million tokens.** The throughput results combined with the on-demand prices read the same week, at stated utilization levels.
- **Memory overhead.** vLLM's own startup accounting of weights, activations, non-torch memory and CUDA graphs per model and GPU. These will replace the 2 GiB allowance with measured values.
- **Serverless cold starts.** Time from request to first token on serverless GPU platforms after an idle period, repeated at different times of day.
- **Same GPU, different provider.** The same model and benchmark on the same GPU type at several providers, to see whether the hourly price difference survives contact with real throughput.
- **Home hardware.** Generation and prompt-processing speed with llama.cpp and MLX on consumer GPUs and Macs, for the models in the VRAM guide at Q4_K_M and Q8_0, at several context lengths, with the power drawn at the wall measured at the same time. These will be compared with the bandwidth ceilings the calculator shows.
- **Fine-tuning throughput.** Training tokens per second for LoRA, QLoRA and full fine-tuning of the same models on the same GPUs, with the trainer, versions and settings recorded, to replace the fine-tuning estimator's MFU assumption with measured figures.

Every benchmark will publish its scripts, raw results and run dates, and will be re-run when engines or prices change enough to matter.

## Conflicts of interest

This site may earn referral or affiliate commissions from RunPod, Vast.ai, DigitalOcean and Lambda, and from Amazon on the home-hardware guides (see the [affiliate disclosure](/affiliate-disclosure/); none are active today). Hardware is ranked by memory and bandwidth from vendor specifications, never by where it is sold or whether a retailer pays. Every provider in the price table, whether it pays a commission or not, is read the same way, and tables are ordered by price or by the column you choose.

## Corrections

Corrections are dated and described on the page they affect. If a price, specification or model shape here is wrong, the contact details are on the [about page](/about/).
