Find the cheapest cloud GPU for your model. Compare hourly and monthly prices for RTX 4090, A6000, L40S, A100 and H100 on Runpod, Lambda and AWS.
Renting a GPU is two questions in the wrong order. Most pricing pages start with dollars per hour, but the first thing that decides the bill is memory: a model that needs 45 GB of VRAM does not run on a 24 GB card at any price. This calculator answers the memory question first, using the same formula as our LLM VRAM calculator, then lists every cloud GPU configuration the job actually fits on, cheapest first, with the hourly and monthly on-demand price on Runpod (Community Cloud and Secure Cloud), Lambda and AWS EC2.
Prices on this page were checked against each provider's own pricing page on October 8, 2026. They are on-demand list prices in US dollars, before tax, for the GPU itself. Storage volumes, network egress, and spot, interruptible or reserved discounts are not included, and every provider changes prices, so check the provider before committing to a long run.
nvidia-smi or a training run, switch to the VRAM mode and type it in.VRAM needed = weights + KV cache + overhead. Weights are parameter count × bytes per parameter (2 for FP16, about 0.61 for Q4_K_M, calibrated against real GGUF file sizes). The KV cache grows with context length and is computed from each model's real architecture, including grouped-query attention and sliding-window layers. Overhead is the larger of 6% of the weights or 0.75 GB, which covers the CUDA context, activation buffers and allocator fragmentation.
A single GPU fits when the total is no more than its memory. When a model is split across several GPUs with tensor or pipeline parallelism, some buffers are duplicated on every card, so each GPU is counted at about 95% of its memory. That is why a 45 GB model fits on one 48 GB card but needs two 24 GB cards, not a theoretical 1.9.
This is an inference calculation for a single sequence. Serving many users at once multiplies the KV cache, and training needs gradients and optimizer state on top of the weights. For training, size the job with the fine-tuning VRAM calculator and enter the result in VRAM mode here.
| GPU | Runpod Community | Runpod Secure | Lambda | AWS (us-east-1) |
|---|---|---|---|---|
| RTX 3090 24GB | $0.22 | $0.50 | — | — |
| RTX 4090 24GB | $0.34 | $0.74 | — | — |
| L4 24GB | $0.44 | $0.49 | — | $0.80 |
| RTX 5090 32GB | $0.69 | $0.99 | — | — |
| RTX A6000 48GB | $0.33 | $0.53 | $1.09 | — |
| A40 48GB | $0.35 | $0.49 | — | — |
| L40S 48GB | $0.79 | $1.09 | — | $1.86 |
| RTX 6000 Ada 48GB | $0.74 | $0.84 | — | — |
| A100 PCIe 80GB | $1.19 | $1.59 | — | — |
| A100 SXM 80GB | $1.39 | $1.59 | — | — |
| H100 PCIe 80GB | $1.99 | $2.89 | $3.29 | — |
| H100 SXM 80GB | $2.69 | $3.49 | $4.29 | $6.88 |
| H100 NVL 94GB | $2.59 | $3.19 | — | — |
| RTX Pro 6000 96GB | $1.69 | $2.09 | — | — |
| H200 141GB | $3.59 | $4.59 | — | — |
| B200 180GB | $5.98 | $6.79 | $6.99 | — |
A dash means the provider does not sell that GPU as a single-GPU on-demand instance. Runpod prices every GPU individually, so a 4-GPU pod costs four times the per-GPU rate. Lambda sells fixed 1, 2, 4 and 8 GPU instances, and its per-GPU rate falls slightly as instances get bigger: an H100 SXM is $4.29/hr as a single GPU and $3.99 per GPU in an 8-GPU instance. AWS sells whole instances: p5.4xlarge is one H100 at $6.88/hr, p5.48xlarge is eight at $55.04/hr, and the A100 80GB is only available as the 8-GPU p4de.24xlarge.
Vast.ai is not in the table on purpose. It is a marketplace where each host sets its own price and listings change by the hour, so any single number would be out of date before you read it. It is often cheaper still for consumer cards. Check its live listings directly.
Renting wins for bursty work: a fine-tuning run, a batch job, an evaluation, or testing whether a model is good enough before buying hardware. The monthly figure is what to compare against buying. If a job keeps one GPU busy most hours of every day, the rental bill over a year or two can exceed the purchase price of an equivalent card plus electricity. The self-hosted LLM cost calculator does that break-even comparison, and also compares self-hosting with paying per token for an API.
How much does it cost to rent an H100? On October 8, 2026, a single H100 SXM was $2.69/hr on Runpod Community Cloud, $3.49/hr on Runpod Secure Cloud, $4.29/hr on Lambda and $6.88/hr on AWS (p5.4xlarge, us-east-1).
What is the cheapest GPU that can run a 70B model? At 4-bit quantization with 8K context, Llama 3.3 70B needs about 45.0 GB, so any single 48 GB card (RTX A6000, A40, L40S, RTX 6000 Ada) or two 24 GB cards will run it.
Are the prices live? No. They are a dated snapshot, checked against each provider's pricing page on October 8, 2026 and shown on the page with that date. Each provider links to its own pricing page so you can confirm the current rate.
Does the monthly cost include storage? No. It is the GPU hourly rate × hours per day × days per month. Persistent volumes, network storage and egress are billed separately by every provider.
Start from memory, not price. Add the model weights (parameters × bytes per parameter for your quantization), the KV cache for your context length, and roughly 6% overhead. Any GPU, or set of GPUs, with at least that much memory will run it. This calculator does the sum for you and then sorts every configuration that fits by price.
On October 8, 2026, on-demand single-GPU prices ranged from $0.16/hr for a 24 GB RTX A5000 on Runpod Community Cloud to $7.89/hr for a 288 GB B300 on Runpod Secure Cloud. A 24 GB RTX 4090 was $0.34/hr on Runpod Community Cloud, and a single H100 SXM was $2.69-$6.88/hr depending on the provider.
At Q4_K_M with 8K context, Llama 3.3 70B needs about 45.0 GB. A single 48 GB card (RTX A6000, A40, L40S or RTX 6000 Ada) runs it, and those are among the cheapest rentable 48 GB GPUs. Two 24 GB cards also work. At FP16 it needs about 141.8 GB, which is more than even a single 141 GB H200, so it needs several GPUs, for example four 48 GB cards or two H200s.
They are Runpod's two on-demand tiers for the same GPU models. Community Cloud is the cheaper one, and Secure Cloud costs more per hour for the same card. Both are priced per GPU, so a multi-GPU pod costs the per-GPU rate times the number of GPUs. Runpod's pricing page describes what each tier includes.
AWS prices whole instances that bundle large amounts of CPU, RAM, networking and local storage, and its A100 and most H100 capacity only comes in 8-GPU instances. You pay for integration with the rest of AWS. For a standalone model-serving or fine-tuning job, specialist GPU clouds are usually cheaper per GPU-hour.
No. They are a snapshot checked against each provider's official pricing page on October 8, 2026, and the date is shown on the page. Prices change, so follow the provider link to confirm before starting a long run.
No. It is the GPU hourly rate × hours per day × days per month. Network volumes, persistent disks and egress are billed separately, and spot, interruptible or reserved pricing can be well below the on-demand rate shown here.
Rent for bursty work such as fine-tuning runs, batch jobs and evaluations, or to test a model before buying hardware. Buy when a GPU would be busy most hours of most days for a year or more. The self-hosted LLM cost calculator compares renting, buying and paying an API per token over time.
Calculate how much VRAM any LLM needs to run locally. Pick a model, quantization, and context size — see download size, total memory required, and which GPUs it fits on.
Calculate VRAM requirements for fine-tuning LLMs with full fine-tuning, LoRA, or QLoRA — accounting for gradients, optimizer states, activations, and gradient checkpointing.
Is it cheaper to self-host an LLM or use an API? Compare GPT, Claude, and Gemini API costs against running open models on your own hardware or cloud GPUs — with break-even timelines.
Detect your GPU with one click and see which LLMs your computer can actually run — ranked by whether they fit in your VRAM, need CPU offloading, or will not run at all.
Estimate LLM tokens per second from memory bandwidth, model size, quantization, and context window. Compare generation speed across GPUs and understand the memory-bandwidth bottleneck.