HIP Out of Memory on AMD? Every Fix for PyTorch, ComfyUI, llama.cpp, and Strix Halo (2026)

hiprocmamdtroubleshootingstrix-halorx-7900-xtxcomfyuiollamalocal-llm

TL;DR: HIP out of memory is ROCm’s version of the CUDA OOM error, but on AMD it has two very different causes. On a discrete card (RX 7900 XTX, RX 9070 XT) it usually means fragmentation or a workload that genuinely doesn’t fit — fixed with PYTORCH_HIP_ALLOC_CONF=expandable_segments:True, ComfyUI’s --reserve-vram, or smaller context. On a Ryzen AI Max (Strix Halo) APU it’s usually artificial: the OS is hiding most of your unified memory behind a GTT limit or a bad Variable Graphics Memory split, and the fix is a kernel parameter or an Adrenalin setting — not a smaller model.

What you’ll be able to do after this fix:

  • Decode which of the two failure modes you’re in from the error text and one amd-smi command
  • Clear fragmentation OOMs in PyTorch-based stacks (ComfyUI, diffusers, fine-tuning scripts) with one environment variable
  • Unlock the memory your 64GB or 128GB Strix Halo machine actually has, on both Linux and Windows

Honest take: if a discrete 24GB card throws this error on a model that comfortably fits — a 13GB Q4 GGUF, an SDXL workflow — the fix below will clear it in minutes. If you’re on an 8GB or 12GB AMD card and hitting it on every model you actually want to run, no allocator flag changes the arithmetic; check the VRAM calculator before you spend another evening on environment variables.

The same error text, two different problems

The error arrives in slightly different clothes depending on the stack. From PyTorch-based tools (ComfyUI, diffusers, Axolotl, anything Hugging Face):

torch.OutOfMemoryError: HIP out of memory. Tried to allocate 2.25 GiB.
GPU 0 has a total capacity of 23.98 GiB of which 1.42 GiB is free.
...
If reserved but unallocated memory is large try setting
PYTORCH_HIP_ALLOC_CONF=expandable_segments:True to avoid fragmentation.

From llama.cpp’s ROCm backend (and Ollama’s, which is built on it), the typical form is a buffer allocation failure:

failed to allocate ROCm0 buffer

— typically right after a line reporting the size it tried to allocate (one documented Strix Halo case: a ~38GB model on a machine with ~120GB nominally available).

Note that PyTorch’s message is nearly identical to the CUDA one — the Hugging Face Accelerate project literally patched its code because AMD’s OOM text differs just enough from NVIDIA’s to break their error matching. The flags it recommends differ too: PYTORCH_HIP_ALLOC_CONF, not PYTORCH_CUDA_ALLOC_CONF.

Before touching anything, find out what the GPU thinks it has. On Linux:

amd-smi metric --mem
# or on older ROCm installs:
rocm-smi --showmeminfo vram gtt

If you’re on a discrete card and “total capacity” matches the sticker (24GB on an RX 7900 XTX), you have a real memory problem — go to the next two sections. If you’re on a Strix Halo / Ryzen AI Max machine and the total is roughly half your installed RAM — 64GB machine showing ~32GB, 128GB showing ~64GB — your memory isn’t full. It’s fenced off. Skip to the GTT section.

Fix 1: PyTorch stacks — the fragmentation flag

The error message’s own suggestion is the right first move, and it’s the AMD-spelled variable:

export PYTORCH_HIP_ALLOC_CONF=expandable_segments:True
python main.py

PyTorch’s default caching allocator carves VRAM into fixed segments, and a long session of variable-sized allocations — exactly what a ComfyUI workflow or a fine-tuning run produces — leaves holes too small to satisfy the next request. The telltale is an OOM where “reserved” is gigabytes larger than “allocated.” Expandable segments let allocations grow without needing physically contiguous blocks, and the setting is not on by default. Reports on the PyTorch tracker and the PyTorch forums confirm the HIP spelling is what ROCm builds read.

Two non-obvious additions from the field:

  • A reboot is a legitimate fix. At least one forum-documented HIP OOM showed plenty of free VRAM and cleared only after a restart — a stuck driver state, not fragmentation. If amd-smi shows free memory but allocation fails, reboot before you refactor.
  • Set it in the service, not your shell. If ComfyUI or a training job runs under systemd or a launcher, an export in your terminal never reaches it — the same trap as OLLAMA_KEEP_ALIVE in our Ollama reloading guide.

Fix 2: ComfyUI on Radeon — reserve VRAM instead of starving the desktop

AMD’s own ComfyUI-on-Radeon guide lists “HIP out of memory” as the most common failure on RX 9000/7000 cards, concentrated in SDXL, Flux, and WAN 2.2 video workflows. Their recommended fix is to stop ComfyUI from grabbing the whole card — your compositor and browser live in the same VRAM budget:

python main.py --reserve-vram 3

That holds 3GB back for the system. On a 24GB RX 7900 XTX that still leaves 21GB for the workflow, and in AMD’s testing it resolves the intermittent OOMs that appear mid-generation rather than at model load. --lowvram exists too, but AMD frames it for 8–12GB cards; on a 24GB card it costs speed for nothing.

Also read your log before assuming failure. ComfyUI’s VAE stage degrades gracefully:

Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.

That line is a successful recovery, not an error — your image still renders, just slower. If you see it on every run and generation times bother you, that’s the signal your resolution or batch size is past the card, not a bug to fix. (Running image workflows on AMD at all is its own topic — our ROCm 7.2 Ubuntu setup guide covers the install side, and ComfyUI’s broader ecosystem is reviewed at aifoss.dev.)

Fix 3: llama.cpp and Ollama on discrete cards

For GGUF inference on a discrete Radeon, HIP out of memory responds to the same levers as the NVIDIA version — fewer GPU layers, smaller context, quantized KV cache — and we’ve covered those flag-by-flag in the CUDA OOM guide; the llama.cpp flags are identical on ROCm. What’s AMD-specific:

  • Try Vulkan before you shrink the model. llama.cpp’s Vulkan backend frequently reports and uses memory differently than the ROCm backend on the same card, and it’s a build flag away (-DGGML_VULKAN=ON). In Ollama, Vulkan shipped as an experimental backend in v0.12.11 (November 2025) and is still flagged experimental in the official GPU docs as of October 2026: set OLLAMA_VULKAN=1 and restart the server.
  • GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 is not a reliable fix on AMD. It’s the documented CUDA spillover switch, and Framework forum threads testing it on ROCm report everything from “ignored under HIP” to “works with a small penalty.” Treat it as an experiment, not a solution.
  • If Ollama says insufficient VRAM to load any model layers on an APU even after you’ve raised memory limits, you’re in the integrated-GPU detection mess — one documented NixOS case raised GTT to 32GB, watched dmesg confirm it, and still had ollama-rocm refuse; llama.cpp loaded the same model fine. Switching runner is a fix. So is the HSA override covered in our ROCm detection guide when the GPU isn’t being picked up at all.

Fix 4: Strix Halo on Linux — the GTT limit is the whole problem

This is the case that fills forums in 2026, because the hardware pitch (“128GB of unified memory for models!”) collides with a kernel default nobody mentions at checkout. ROCm allocates through GTT — GPU-accessible system memory — and the kernel’s Translation Table Manager caps how much of your RAM that pool can pin. The default cap is roughly half of RAM. The result, documented in a Level1Techs thread: a 128GB machine failing to allocate a 38GB buffer.

AMD’s own Strix Halo optimization docs point at the TTM limit, with one critical detail: the value is in 4KiB pages, not bytes. 1GiB = 262,144 pages. Set it on the kernel command line (GRUB’s GRUB_CMDLINE_LINUX_DEFAULT):

ttm.pages_limit=32505856 ttm.page_pool_size=32505856

Then update-grub and reboot. That value — used by the kyuz0 strix-halo-toolboxes project, the most battle-tested llama.cpp environment for this hardware — pins up to 124GiB on a 128GB machine. The Lychee llama-cpp-for-strix-halo build recommends the more conservative 25600000 (~97.6GiB), leaving more headroom for the OS. For a 64GB machine, scale down: 14680064 pins 56GiB. Verify after reboot:

cat /sys/module/ttm/parameters/pages_limit

Expected output: the number you set. Two gotchas, both field-reported: on some setups a /etc/modprobe.d/ entry silently fails where the kernel command line works, so verify rather than trust; and some kernel/driver combinations expose the module as amdttm instead of ttm — check with ls /sys/module | grep ttm and use whichever prefix exists.

One more Strix Halo-specific crash worth knowing: llama.cpp used to abort on these APUs (gfx1151) with a hipMemAdviseSetCoarseGrain failure; a 2026 patch demoted it from fatal to a warning. If you’re on an old build and seeing that string, update llama.cpp before touching anything else.

Fix 5: Strix Halo on Windows — Variable Graphics Memory, and the trap inside it

Windows splits the same silicon differently: a “dedicated” VRAM slice set in firmware plus shared GPU memory on top. In AMD Adrenalin under Performance → Tuning → Variable Graphics Memory, a 128GB machine allows up to 96GB dedicated (32GB stays with Windows), for roughly 112GB total addressable once shared memory is counted. The setting survives reboots and driver updates.

Here’s the trap: maxing the dedicated slice can cause out-of-memory errors instead of fixing them. LM Studio’s Vulkan backend treats Strix Halo like an iGPU and fills memory in its own order — one tested 64GB configuration crashed at model load with a 48GB dedicated / 16GB system split because system RAM ran out during loading, and ran stable with the split reversed to 16GB dedicated. The overflow into shared memory carried no measurable speed penalty, which makes sense — it’s all the same physical LPDDR5X behind a 256GB/s controller either way. If LM Studio shows the wrong “available” memory or crashes loading models that should fit: lower the dedicated slice, disable “keep model in memory” in the model load settings, and turn off mmap.

Which fix for which setup

Your setupThe error’s real causeThe fix
RX 7900 XTX / 9070 XT + ComfyUI or PyTorchFragmentation or desktop competing for VRAMPYTORCH_HIP_ALLOC_CONF=expandable_segments:True, --reserve-vram 3
Discrete Radeon + llama.cpp/OllamaModel + context genuinely over budgetSame levers as CUDA OOM; try Vulkan backend
Strix Halo, LinuxKernel TTM cap at ~half of RAMttm.pages_limit on kernel cmdline, in 4KiB pages
Strix Halo, WindowsBad dedicated/shared VGM splitAdrenalin VGM Custom — and smaller dedicated is often better
Any AMD, free VRAM but still OOMStuck driver stateReboot. Really.

When the fix is a bigger card

Allocator flags recover memory you already own; they don’t mint more. If your errors come from 27B-class models on a 12GB or 16GB card, the honest answer in October 2026 is the one in our 24GB-tier model guide: the RX 7900 XTX remains the cheapest new 24GB card at $900–$1,358 street (verified October 9, 2026) — now well under a used RTX 3090, which has climbed to $1,399–$1,450. Whether its ROCm trade-offs suit you is a decision we’ve mapped separately. And if you only need big memory occasionally, renting a 3090 from $0.07/hr on Vast.ai costs less than a month of debugging.

FAQ

Is HIP out of memory the same as CUDA out of memory? Functionally yes — ROCm implements the CUDA-style API, and PyTorch raises the same torch.OutOfMemoryError with HIP wording. The fixes differ in spelling (PYTORCH_HIP_ALLOC_CONF) and in one big structural way: on AMD APUs the error is usually a configurable memory split, not real exhaustion.

Why does my 128GB Strix Halo machine show only 64GB to ROCm? The Linux kernel’s TTM pages limit defaults to roughly half of system RAM. Raise it with ttm.pages_limit (in 4KiB pages) on the kernel command line and reboot. It’s a limit, not a hardware partition.

Should I set dedicated VRAM to the maximum on Windows? Usually no. Dedicated memory is walled off from Windows, and loaders need system RAM to stage the model — a maxed slice can crash loads that a modest one survives. Start at 16–32GB dedicated and only raise it if a specific application demands it.

Does GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 work on AMD? Unreliably. Field reports from Strix Halo owners disagree, and it’s documented as a CUDA feature. Fix the GTT limit or use Vulkan instead of counting on it.

My VRAM is free but I still get the error. What now? Reboot first — stuck driver states produce exactly this signature. If it recurs, check whether another process (compositor, browser, a zombie Python) is holding memory with amd-smi before blaming the model.

Sources

Last updated October 11, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.