GPU Rental Cost Calculator

Find the cheapest cloud GPU for your model. Compare hourly and monthly prices for RTX 4090, A6000, L40S, A100 and H100 on Runpod, Lambda and AWS.

Advertisement

GPU rental cost calculator: which cloud GPU fits your model, and what it costs per hour and per month

Renting a GPU is two questions in the wrong order. Most pricing pages start with dollars per hour, but the first thing that decides the bill is memory: a model that needs 45 GB of VRAM does not run on a 24 GB card at any price. This calculator answers the memory question first, using the same formula as our LLM VRAM calculator, then lists every cloud GPU configuration the job actually fits on, cheapest first, with the hourly and monthly on-demand price on Runpod (Community Cloud and Secure Cloud), Lambda and AWS EC2.

Prices on this page were checked against each provider's own pricing page on October 8, 2026. They are on-demand list prices in US dollars, before tax, for the GPU itself. Storage volumes, network egress, and spot, interruptible or reserved discounts are not included, and every provider changes prices, so check the provider before committing to a long run.

How to use it

  1. Size the job. Pick a model (or search Hugging Face), a quantization such as FP16, Q8_0 or Q4_K_M, and a context length. If you already know your peak memory from nvidia-smi or a training run, switch to the VRAM mode and type it in.
  2. Set your hours. Hours per day and days per month turn the hourly rate into a monthly bill. A workday (8 hours × 22 days) is 176 GPU-hours. Running 24/7 is 720.
  3. Read the list. Each row is the smallest configuration a provider sells that holds the job, so you never see 8 GPUs where 2 would do. The memory-used column shows how much headroom is left.

How the fit is calculated

VRAM needed = weights + KV cache + overhead. Weights are parameter count × bytes per parameter (2 for FP16, about 0.61 for Q4_K_M, calibrated against real GGUF file sizes). The KV cache grows with context length and is computed from each model's real architecture, including grouped-query attention and sliding-window layers. Overhead is the larger of 6% of the weights or 0.75 GB, which covers the CUDA context, activation buffers and allocator fragmentation.

A single GPU fits when the total is no more than its memory. When a model is split across several GPUs with tensor or pipeline parallelism, some buffers are duplicated on every card, so each GPU is counted at about 95% of its memory. That is why a 45 GB model fits on one 48 GB card but needs two 24 GB cards, not a theoretical 1.9.

This is an inference calculation for a single sequence. Serving many users at once multiplies the KV cache, and training needs gradients and optimizer state on top of the weights. For training, size the job with the fine-tuning VRAM calculator and enter the result in VRAM mode here.

Cloud GPU prices per hour (single GPU, on-demand, October 8, 2026)

GPURunpod CommunityRunpod SecureLambdaAWS (us-east-1)
RTX 3090 24GB$0.22$0.50——
RTX 4090 24GB$0.34$0.74——
L4 24GB$0.44$0.49—$0.80
RTX 5090 32GB$0.69$0.99——
RTX A6000 48GB$0.33$0.53$1.09—
A40 48GB$0.35$0.49——
L40S 48GB$0.79$1.09—$1.86
RTX 6000 Ada 48GB$0.74$0.84——
A100 PCIe 80GB$1.19$1.59——
A100 SXM 80GB$1.39$1.59——
H100 PCIe 80GB$1.99$2.89$3.29—
H100 SXM 80GB$2.69$3.49$4.29$6.88
H100 NVL 94GB$2.59$3.19——
RTX Pro 6000 96GB$1.69$2.09——
H200 141GB$3.59$4.59——
B200 180GB$5.98$6.79$6.99—

A dash means the provider does not sell that GPU as a single-GPU on-demand instance. Runpod prices every GPU individually, so a 4-GPU pod costs four times the per-GPU rate. Lambda sells fixed 1, 2, 4 and 8 GPU instances, and its per-GPU rate falls slightly as instances get bigger: an H100 SXM is $4.29/hr as a single GPU and $3.99 per GPU in an 8-GPU instance. AWS sells whole instances: p5.4xlarge is one H100 at $6.88/hr, p5.48xlarge is eight at $55.04/hr, and the A100 80GB is only available as the 8-GPU p4de.24xlarge.

Worked examples

  • Mistral 7B at Q4_K_M, 8K context: 5.9 GB. Cheapest fit: RTX A5000 24GB on Runpod Community Cloud at $0.16/hr, or $28 a month at 8 hours a day for 22 days. A 7B model at 4-bit fits on almost anything, so the cheapest 24 GB card wins.
  • Llama 3.3 70B at Q4_K_M, 8K context: 45.0 GB. Cheapest fit: 2× RTX A5000 24GB on Runpod Community Cloud at $0.32/hr, or $56 a month at 8 hours a day for 22 days. This is the classic case for renting: too big for one 24 GB RTX 4090, but it fits two 24 GB cards or a single 48 GB card.
  • Llama 3.3 70B at FP16, 8K context: 141.8 GB. Cheapest fit: 8× RTX A5000 24GB on Runpod Community Cloud at $1.28/hr, or $225 a month at 8 hours a day for 22 days. Full precision triples the memory. Even a single 141 GB H200 is just too small, so this is a multi-GPU job.
  • Llama 3.1 405B at Q4_K_M, 8K context: 248.3 GB. Cheapest fit: 8× RTX A6000 48GB on Runpod Community Cloud at $2.64/hr, or $1,901 a month at 24 hours a day for 30 days. At this size only multi-GPU nodes apply, and a month of 24/7 use is a serious budget line.

Choosing between providers

  • Runpod Community Cloud is usually the cheapest per hour. Availability of a particular card, especially in multi-GPU pods, varies.
  • Runpod Secure Cloud costs more per hour for the same GPU. Compare the two tiers on Runpod's site if hosting location or uptime terms matter for your job.
  • Lambda sells datacenter GPUs (A100, H100, GH200, B200) as fixed instances with large amounts of CPU RAM and local NVMe.
  • AWS EC2 is the most expensive per GPU-hour on this list, and is the choice when the rest of your stack already lives in AWS and data transfer, IAM and compliance matter more than the hourly rate.

Vast.ai is not in the table on purpose. It is a marketplace where each host sets its own price and listings change by the hour, so any single number would be out of date before you read it. It is often cheaper still for consumer cards. Check its live listings directly.

Hourly or monthly: when renting stops making sense

Renting wins for bursty work: a fine-tuning run, a batch job, an evaluation, or testing whether a model is good enough before buying hardware. The monthly figure is what to compare against buying. If a job keeps one GPU busy most hours of every day, the rental bill over a year or two can exceed the purchase price of an equivalent card plus electricity. The self-hosted LLM cost calculator does that break-even comparison, and also compares self-hosting with paying per token for an API.

Making a model fit a cheaper GPU

  • Quantize. Q4_K_M uses roughly 30% of FP16's weight memory with a small quality loss, which often moves a job from an 80 GB card to a 48 GB one.
  • Shorten the context. The KV cache scales linearly with context length. Running at 8K instead of 128K can save tens of gigabytes on large models.
  • Quantize the KV cache. An 8-bit KV cache halves context memory.
  • Use a smaller model in the same family. A 32B model at Q4 usually fits a single 24 GB card at modest context. A 70B does not.

Frequently asked questions

How much does it cost to rent an H100? On October 8, 2026, a single H100 SXM was $2.69/hr on Runpod Community Cloud, $3.49/hr on Runpod Secure Cloud, $4.29/hr on Lambda and $6.88/hr on AWS (p5.4xlarge, us-east-1).

What is the cheapest GPU that can run a 70B model? At 4-bit quantization with 8K context, Llama 3.3 70B needs about 45.0 GB, so any single 48 GB card (RTX A6000, A40, L40S, RTX 6000 Ada) or two 24 GB cards will run it.

Are the prices live? No. They are a dated snapshot, checked against each provider's pricing page on October 8, 2026 and shown on the page with that date. Each provider links to its own pricing page so you can confirm the current rate.

Does the monthly cost include storage? No. It is the GPU hourly rate × hours per day × days per month. Persistent volumes, network storage and egress are billed separately by every provider.

Frequently Asked Questions

How do I work out which cloud GPU I need?+

Start from memory, not price. Add the model weights (parameters × bytes per parameter for your quantization), the KV cache for your context length, and roughly 6% overhead. Any GPU, or set of GPUs, with at least that much memory will run it. This calculator does the sum for you and then sorts every configuration that fits by price.

How much does it cost to rent a GPU per hour?+

On October 8, 2026, on-demand single-GPU prices ranged from $0.16/hr for a 24 GB RTX A5000 on Runpod Community Cloud to $7.89/hr for a 288 GB B300 on Runpod Secure Cloud. A 24 GB RTX 4090 was $0.34/hr on Runpod Community Cloud, and a single H100 SXM was $2.69-$6.88/hr depending on the provider.

What is the cheapest GPU for running a 70B model?+

At Q4_K_M with 8K context, Llama 3.3 70B needs about 45.0 GB. A single 48 GB card (RTX A6000, A40, L40S or RTX 6000 Ada) runs it, and those are among the cheapest rentable 48 GB GPUs. Two 24 GB cards also work. At FP16 it needs about 141.8 GB, which is more than even a single 141 GB H200, so it needs several GPUs, for example four 48 GB cards or two H200s.

What is the difference between Runpod Community Cloud and Secure Cloud?+

They are Runpod's two on-demand tiers for the same GPU models. Community Cloud is the cheaper one, and Secure Cloud costs more per hour for the same card. Both are priced per GPU, so a multi-GPU pod costs the per-GPU rate times the number of GPUs. Runpod's pricing page describes what each tier includes.

Why is AWS more expensive than Runpod or Lambda?+

AWS prices whole instances that bundle large amounts of CPU, RAM, networking and local storage, and its A100 and most H100 capacity only comes in 8-GPU instances. You pay for integration with the rest of AWS. For a standalone model-serving or fine-tuning job, specialist GPU clouds are usually cheaper per GPU-hour.

Are the prices on this page live?+

No. They are a snapshot checked against each provider's official pricing page on October 8, 2026, and the date is shown on the page. Prices change, so follow the provider link to confirm before starting a long run.

Does the monthly estimate include storage and data transfer?+

No. It is the GPU hourly rate × hours per day × days per month. Network volumes, persistent disks and egress are billed separately, and spot, interruptible or reserved pricing can be well below the on-demand rate shown here.

Should I rent a GPU or buy one?+

Rent for bursty work such as fine-tuning runs, batch jobs and evaluations, or to test a model before buying hardware. Buy when a GPU would be busy most hours of most days for a year or more. The self-hosted LLM cost calculator compares renting, buying and paying an API per token over time.

Related tools

This tool is provided for informational and educational purposes only. All processing happens in your browser — no data is sent to or stored on our servers. While we strive for accuracy, we make no warranties about the completeness or reliability of results.