Skip to main content
DevOpsintermediate

Fix "CUDA out of memory" (torch.OutOfMemoryError) in PyTorch

Fix `torch.OutOfMemoryError: CUDA out of memory. Tried to allocate` in PyTorch: free the GPU, cut batch size, use mixed precision, gradient checkpointing, PYTORCH_CUDA_ALLOC_CONF, quantization or offload.

11 min readUpdated October 2026

PyTorch raises this when a single GPU allocation fails:

torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 2.00 GiB. GPU 0 has a total capacity of 23.65 GiB of which 1.12 GiB is free. Including non-PyTorch memory, this process has 22.51 GiB memory in use. Of the allocated memory 20.87 GiB is allocated by PyTorch, and 1.19 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.

Older PyTorch versions word it differently and raise a plain RuntimeError, but it is the same failure:

RuntimeError: CUDA out of memory. Tried to allocate 20.00 MiB (GPU 0; 7.79 GiB total capacity; 6.52 GiB already allocated; 9.31 MiB free; 6.57 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.

You may also see it as torch.cuda.OutOfMemoryError. Recent PyTorch versions expose it as torch.OutOfMemoryError, and catching either one works.

Read the Error Before Changing Anything

The message tells you which fix applies. Pull out three numbers:

Number in the messageWhat it tells you
Tried to allocateThe one allocation that failed. Tiny (MiB) = you are at the edge; huge (GiB) = one oversized tensor, usually activations
allocated by PyTorchMemory held by live tensors. This is your real usage
reserved but unallocatedCached by PyTorch but unused. If this is large, you have fragmentation, not a shortage

Also compare "this process has X in use" with the card's total. If your process uses far less than the total and there is still no free memory, something else is on the GPU.

Fix 1: Check What Else Is Using the GPU

nvidia-smi

Look at the process list at the bottom. A forgotten Jupyter kernel, an earlier training run that did not exit, or a second notebook that loaded the same model will each hold gigabytes. Kill them:

nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv
kill <pid>

In Jupyter, restarting the kernel is the only reliable way to release everything that kernel allocated.

Fix 2: Reduce the Batch Size (Keep the Effective Batch With Accumulation)

Activation memory scales with batch size and sequence length. Halve the batch first:

loader = DataLoader(dataset, batch_size=8)  # was 16

To keep the same effective batch for training, accumulate gradients:

accum_steps = 2
for i, (x, y) in enumerate(loader):
    loss = model(x.cuda(), y.cuda()) / accum_steps
    loss.backward()
    if (i + 1) % accum_steps == 0:
        optimizer.step()
        optimizer.zero_grad(set_to_none=True)

For transformers, sequence length matters as much as batch size. Truncating max_length from 4096 to 2048 roughly halves activation memory in the attention-heavy layers.

Fix 3: Train in bf16 or fp16

Mixed precision halves activation memory and speeds up most modern GPUs:

# bf16 (Ampere and newer): no loss scaling needed
with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
    loss = model(x, y)
loss.backward()
optimizer.step()

On older cards without bf16, use dtype=torch.float16 and wrap the backward pass with torch.amp.GradScaler to avoid underflowing gradients.

With Hugging Face Trainer, pass bf16=True (Ampere and newer) or fp16=True in TrainingArguments.

Fix 4: Enable Gradient Checkpointing

Gradient checkpointing discards intermediate activations in the forward pass and recomputes them during backward. It costs roughly 20-30% more compute and saves a large share of activation memory:

model.gradient_checkpointing_enable()  # Hugging Face models

For your own modules, wrap expensive blocks with torch.utils.checkpoint.checkpoint(block, x, use_reentrant=False).

Fix 5: Stop Holding Memory You Don't Need

These three bugs cause memory to grow every step until it fails:

# Wrong: keeps every iteration's autograd graph alive
total_loss += loss
# Right
total_loss += loss.item()

# Wrong: logits stay on the GPU with their graph
all_logits.append(logits)
# Right
all_logits.append(logits.detach().cpu())

And never run evaluation or inference with autograd on:

model.eval()
with torch.inference_mode():
    preds = model(batch)

Fix 6: Free the Cache Properly

torch.cuda.empty_cache() only returns blocks that no tensor references. Delete the references first:

del outputs, loss
import gc
gc.collect()
torch.cuda.empty_cache()

This matters most in notebooks, where a variable from an earlier cell can quietly hold a whole model.

Advertisement

Fix 7: Fix Fragmentation With PYTORCH_CUDA_ALLOC_CONF

If the error reports a lot of memory "reserved by PyTorch but unallocated", the memory exists but is split into blocks too small for the request. Let the allocator grow segments instead:

export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
python train.py

Set it before Python starts (or at the very top of the script, before any CUDA call). On older PyTorch versions without expandable_segments, the equivalent knob is max_split_size_mb:

export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:128

Fix 8: Quantize the Model

If the weights are most of your memory, store them in fewer bits. With Hugging Face Transformers and bitsandbytes:

from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=bnb, device_map="auto")

4-bit weights take about a quarter of fp16. For fine-tuning, this is the QLoRA setup: 4-bit base model plus small trainable LoRA adapters.

Fix 9: Offload to CPU RAM

device_map="auto" places what fits on the GPU and the rest in CPU RAM. Cap the GPU share so there is room left for activations:

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    max_memory={0: "20GiB", "cpu": "64GiB"},
)

It runs, but every offloaded layer crosses PCIe on every forward pass, so expect it to be several times slower. For training, DeepSpeed ZeRO stage 2 or 3 with CPU offload does the same for optimizer state and gradients.

Fix 10: vLLM and Other Inference Servers

vLLM reserves most of the GPU for its KV cache at startup, so it often fails with CUDA out of memory before serving a single request:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --gpu-memory-utilization 0.85 \
  --max-model-len 8192

Lower --max-model-len if the KV cache for the default (often the model's full context window) does not fit. Lower --gpu-memory-utilization if something else shares the card. Add --tensor-parallel-size 2 to split the model across two GPUs.

Find Out Where the Memory Goes

print(torch.cuda.memory_summary())
print(torch.cuda.max_memory_allocated() / 2**30, "GiB peak")

For a timeline of every allocation, record a memory snapshot and drop the file onto pytorch.org/memory_viz:

torch.cuda.memory._record_memory_history()
# ... run a few steps ...
torch.cuda.memory._dump_snapshot("snapshot.pickle")

When the Model Simply Doesn't Fit

Every fix above trims activations, caches or precision. None of them helps if the weights alone are bigger than the card. Do the arithmetic before tuning:

Modelfp16 weights4-bit weightsFull fine-tune (Adam, mixed precision)
7-8B~15 GB~5 GB~120 GB+
13-14B~27 GB~9 GB~220 GB+
70B~140 GB~43 GB~1.1 TB+

Full fine-tuning with Adam needs about 16 bytes per parameter before activations: weights, gradients and two optimizer states. That is why a 7B model that runs comfortably for inference on a 24 GB card cannot be fully fine-tuned on it.

Work out your real number with the LLM VRAM calculator for inference, including the KV cache at your context length, or the fine-tuning VRAM calculator for full, LoRA and QLoRA training. If the total is above your card even at 4-bit, you have three options left:

  1. Quantize harder or shrink the model: 3-bit or 2-bit, or a smaller model in the same family.
  2. Offload (Fix 9) and accept the slowdown.
  3. Use a bigger GPU for the job. A 48 GB or 80 GB cloud GPU rented by the hour is often cheaper than a day spent tuning around a 24 GB limit. The GPU rental cost calculator takes the model, quantization and context, lists every cloud GPU it fits on, and shows the hourly and monthly price across providers.

Verify the Fix

Run a few steps of the real workload and check the peak, not the steady state:

torch.cuda.reset_peak_memory_stats()
# ... run 3-5 training steps or a full-length inference ...
peak = torch.cuda.max_memory_allocated() / 2**30
total = torch.cuda.get_device_properties(0).total_memory / 2**30
print(f"peak {peak:.1f} / {total:.1f} GiB")

Leave 10-15% headroom. The longest sequence in a dataset, or the last uneven batch, often peaks higher than the first few steps.

Prevention

  • Sort or bucket training data by length so one long sample does not set the peak for a whole batch.
  • Set max_length explicitly; tokenizers default to the model's full context.
  • Wrap all evaluation in torch.inference_mode().
  • Log torch.cuda.max_memory_allocated() once per epoch so a slow leak shows up before it crashes a long run.
  • In shared environments, set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True by default.

Frequently Asked Questions

Find answers to common questions

Work down this list: check nothing else is holding the GPU (nvidia-smi), halve the batch size and use gradient accumulation to keep the effective batch, train in bf16 or fp16 with autocast, enable gradient checkpointing, run inference under torch.inference_mode(), and set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True if the error reports a lot of reserved-but-unallocated memory. If the weights alone are bigger than the card, quantize, offload, or use a bigger GPU.

It is the size of the single allocation that failed, not the total your job needs. A small number such as 20 MiB on a nearly full card means you are right at the limit; a large number such as 8 GiB usually points at one big tensor, typically activations for a large batch or a long sequence.

Rarely on its own. It returns cached blocks that no live tensor uses back to the driver, which helps other processes and nvidia-smi readings, but it cannot free memory still referenced by a tensor. Delete the references first (del tensor; gc.collect()), then call empty_cache().

An allocator setting that lets PyTorch grow existing memory segments instead of carving new fixed-size ones, which reduces fragmentation. Use it when the error says a large amount is 'reserved by PyTorch but unallocated'. Set it in the environment before Python starts, or before the first CUDA allocation.

Either the free memory is fragmented into blocks too small for the requested allocation (fix with expandable_segments), or another process grabbed it between your check and the allocation. nvidia-smi also shows memory per process, so a stale Jupyter kernel holding memory is easy to spot there.

Validation is usually run without torch.no_grad() or torch.inference_mode(), so PyTorch builds an autograd graph it never uses, or the validation batch size is larger than the training one. Wrap the evaluation loop in torch.inference_mode() and call model.eval().

Something is keeping tensors alive across iterations. The classic case is accumulating the loss tensor (total_loss += loss) instead of its value (total_loss += loss.item()), which keeps every iteration's graph in memory. Appending outputs or logits to a list without .detach().cpu() does the same.

vLLM pre-allocates most of the GPU for its KV cache. Lower --gpu-memory-utilization (default 0.9) if another process shares the card, lower --max-model-len so the KV cache fits, or serve a quantized checkpoint. If the weights alone do not fit, add GPUs with --tensor-parallel-size.

For inference, roughly parameters times bytes per parameter (2 for fp16, about 0.6 for 4-bit), plus the KV cache for your context length, plus some overhead. A 7B model needs about 15 GB at fp16 and about 5-6 GB at 4-bit. Full fine-tuning with Adam needs roughly 16 bytes per parameter before activations. The LLM VRAM calculator gives exact numbers per model.

Yes. Unlike some fatal errors it is an ordinary Python exception (a subclass of RuntimeError). Catch it, drop references to the failed batch, call torch.cuda.empty_cache(), and retry with half the batch size. Libraries such as Hugging Face Accelerate's find_executable_batch_size do exactly this.