PyTorch raises this when a single GPU allocation fails:
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 2.00 GiB. GPU 0 has a total capacity of 23.65 GiB of which 1.12 GiB is free. Including non-PyTorch memory, this process has 22.51 GiB memory in use. Of the allocated memory 20.87 GiB is allocated by PyTorch, and 1.19 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.
Older PyTorch versions word it differently and raise a plain RuntimeError, but it is the same failure:
RuntimeError: CUDA out of memory. Tried to allocate 20.00 MiB (GPU 0; 7.79 GiB total capacity; 6.52 GiB already allocated; 9.31 MiB free; 6.57 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.
You may also see it as torch.cuda.OutOfMemoryError. Recent PyTorch versions expose it as torch.OutOfMemoryError, and catching either one works.
Read the Error Before Changing Anything
The message tells you which fix applies. Pull out three numbers:
| Number in the message | What it tells you |
|---|---|
| Tried to allocate | The one allocation that failed. Tiny (MiB) = you are at the edge; huge (GiB) = one oversized tensor, usually activations |
| allocated by PyTorch | Memory held by live tensors. This is your real usage |
| reserved but unallocated | Cached by PyTorch but unused. If this is large, you have fragmentation, not a shortage |
Also compare "this process has X in use" with the card's total. If your process uses far less than the total and there is still no free memory, something else is on the GPU.
Fix 1: Check What Else Is Using the GPU
nvidia-smi
Look at the process list at the bottom. A forgotten Jupyter kernel, an earlier training run that did not exit, or a second notebook that loaded the same model will each hold gigabytes. Kill them:
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv
kill <pid>
In Jupyter, restarting the kernel is the only reliable way to release everything that kernel allocated.
Fix 2: Reduce the Batch Size (Keep the Effective Batch With Accumulation)
Activation memory scales with batch size and sequence length. Halve the batch first:
loader = DataLoader(dataset, batch_size=8) # was 16
To keep the same effective batch for training, accumulate gradients:
accum_steps = 2
for i, (x, y) in enumerate(loader):
loss = model(x.cuda(), y.cuda()) / accum_steps
loss.backward()
if (i + 1) % accum_steps == 0:
optimizer.step()
optimizer.zero_grad(set_to_none=True)
For transformers, sequence length matters as much as batch size. Truncating max_length from 4096 to 2048 roughly halves activation memory in the attention-heavy layers.
Fix 3: Train in bf16 or fp16
Mixed precision halves activation memory and speeds up most modern GPUs:
# bf16 (Ampere and newer): no loss scaling needed
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
loss = model(x, y)
loss.backward()
optimizer.step()
On older cards without bf16, use dtype=torch.float16 and wrap the backward pass with torch.amp.GradScaler to avoid underflowing gradients.
With Hugging Face Trainer, pass bf16=True (Ampere and newer) or fp16=True in TrainingArguments.
Fix 4: Enable Gradient Checkpointing
Gradient checkpointing discards intermediate activations in the forward pass and recomputes them during backward. It costs roughly 20-30% more compute and saves a large share of activation memory:
model.gradient_checkpointing_enable() # Hugging Face models
For your own modules, wrap expensive blocks with torch.utils.checkpoint.checkpoint(block, x, use_reentrant=False).
Fix 5: Stop Holding Memory You Don't Need
These three bugs cause memory to grow every step until it fails:
# Wrong: keeps every iteration's autograd graph alive
total_loss += loss
# Right
total_loss += loss.item()
# Wrong: logits stay on the GPU with their graph
all_logits.append(logits)
# Right
all_logits.append(logits.detach().cpu())
And never run evaluation or inference with autograd on:
model.eval()
with torch.inference_mode():
preds = model(batch)
Fix 6: Free the Cache Properly
torch.cuda.empty_cache() only returns blocks that no tensor references. Delete the references first:
del outputs, loss
import gc
gc.collect()
torch.cuda.empty_cache()
This matters most in notebooks, where a variable from an earlier cell can quietly hold a whole model.
Fix 7: Fix Fragmentation With PYTORCH_CUDA_ALLOC_CONF
If the error reports a lot of memory "reserved by PyTorch but unallocated", the memory exists but is split into blocks too small for the request. Let the allocator grow segments instead:
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
python train.py
Set it before Python starts (or at the very top of the script, before any CUDA call). On older PyTorch versions without expandable_segments, the equivalent knob is max_split_size_mb:
export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:128
Fix 8: Quantize the Model
If the weights are most of your memory, store them in fewer bits. With Hugging Face Transformers and bitsandbytes:
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=bnb, device_map="auto")
4-bit weights take about a quarter of fp16. For fine-tuning, this is the QLoRA setup: 4-bit base model plus small trainable LoRA adapters.
Fix 9: Offload to CPU RAM
device_map="auto" places what fits on the GPU and the rest in CPU RAM. Cap the GPU share so there is room left for activations:
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
max_memory={0: "20GiB", "cpu": "64GiB"},
)
It runs, but every offloaded layer crosses PCIe on every forward pass, so expect it to be several times slower. For training, DeepSpeed ZeRO stage 2 or 3 with CPU offload does the same for optimizer state and gradients.
Fix 10: vLLM and Other Inference Servers
vLLM reserves most of the GPU for its KV cache at startup, so it often fails with CUDA out of memory before serving a single request:
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--gpu-memory-utilization 0.85 \
--max-model-len 8192
Lower --max-model-len if the KV cache for the default (often the model's full context window) does not fit. Lower --gpu-memory-utilization if something else shares the card. Add --tensor-parallel-size 2 to split the model across two GPUs.
Find Out Where the Memory Goes
print(torch.cuda.memory_summary())
print(torch.cuda.max_memory_allocated() / 2**30, "GiB peak")
For a timeline of every allocation, record a memory snapshot and drop the file onto pytorch.org/memory_viz:
torch.cuda.memory._record_memory_history()
# ... run a few steps ...
torch.cuda.memory._dump_snapshot("snapshot.pickle")
When the Model Simply Doesn't Fit
Every fix above trims activations, caches or precision. None of them helps if the weights alone are bigger than the card. Do the arithmetic before tuning:
| Model | fp16 weights | 4-bit weights | Full fine-tune (Adam, mixed precision) |
|---|---|---|---|
| 7-8B | ~15 GB | ~5 GB | ~120 GB+ |
| 13-14B | ~27 GB | ~9 GB | ~220 GB+ |
| 70B | ~140 GB | ~43 GB | ~1.1 TB+ |
Full fine-tuning with Adam needs about 16 bytes per parameter before activations: weights, gradients and two optimizer states. That is why a 7B model that runs comfortably for inference on a 24 GB card cannot be fully fine-tuned on it.
Work out your real number with the LLM VRAM calculator for inference, including the KV cache at your context length, or the fine-tuning VRAM calculator for full, LoRA and QLoRA training. If the total is above your card even at 4-bit, you have three options left:
- Quantize harder or shrink the model: 3-bit or 2-bit, or a smaller model in the same family.
- Offload (Fix 9) and accept the slowdown.
- Use a bigger GPU for the job. A 48 GB or 80 GB cloud GPU rented by the hour is often cheaper than a day spent tuning around a 24 GB limit. The GPU rental cost calculator takes the model, quantization and context, lists every cloud GPU it fits on, and shows the hourly and monthly price across providers.
Verify the Fix
Run a few steps of the real workload and check the peak, not the steady state:
torch.cuda.reset_peak_memory_stats()
# ... run 3-5 training steps or a full-length inference ...
peak = torch.cuda.max_memory_allocated() / 2**30
total = torch.cuda.get_device_properties(0).total_memory / 2**30
print(f"peak {peak:.1f} / {total:.1f} GiB")
Leave 10-15% headroom. The longest sequence in a dataset, or the last uneven batch, often peaks higher than the first few steps.
Prevention
- Sort or bucket training data by length so one long sample does not set the peak for a whole batch.
- Set
max_lengthexplicitly; tokenizers default to the model's full context. - Wrap all evaluation in
torch.inference_mode(). - Log
torch.cuda.max_memory_allocated()once per epoch so a slow leak shows up before it crashes a long run. - In shared environments, set
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:Trueby default.