Framework Desktop 192GB Preorder in 2026: What $6,799 Actually Unlocks Over the 128GB Model

strix-halomini-pcamdlocal-llmhardware

TL;DR: Framework’s 192GB Desktop (Ryzen AI Max+ PRO 495) opened preorders September 30 at $6,799 DIY — nearly double the 128GB model — and batch 1 sold out anyway. The extra $3,350 buys a 40GB capacity window (a ~160GB GPU pool vs ~120GB) and zero extra speed. That window holds exactly three models most people want: Qwen3-235B at Q4, MiniMax M3 at 2-bit, and DeepSeek R1 at 1.58-bit.

Framework Desktop 192GBFramework Desktop 128GBMac Studio M5 Ultra 256GB
Best for130–160GB MoE models at ~11 tok/sEverything up to gpt-oss-120bThe same giant models, 4.4× faster
Price / Cost$6,799 DIY (ships Nov 2026)$3,449$10,799
The catchSame 273 GB/s — capacity, not speed~120GB usable ceiling$4,000 more, macOS-only stack

Honest take: If you weren’t already refreshing frame.work waiting for this, the 128GB model at $3,449 is still the right buy. The 192GB config is for the specific person who needs Qwen3-235B at Q4 or DeepSeek R1 resident at home, knows it will decode at ~11 tok/s, and wants it anyway.

Framework opened preorders on September 30, 2026 for the machine it had been teasing all fall: a Desktop built on AMD’s Ryzen AI Max+ PRO 495 (“Gorgon Halo”) with 192GB of soldered LPDDR5X-8533. The DIY Edition is $6,799; a prebuilt with a 2TB NVMe drive and Fedora preinstalled is $7,449, with first units shipping in November 2026. Framework says constrained memory inventory limits it to a single initial batch — and batch 1 sold out anyway, sky-high price and all.

When we compared the 128GB Framework Desktop against the GMKtec EVO-X2 in late September, this machine was still “coming soon” with no price, and our standing advice was don’t wait for it. Now there’s a price, a ship date, and a sellout. The advice needs numbers behind it.

What $6,799 buys, in hardware terms

The chip upgrade is real but small. The PRO 495 keeps the same 16 Zen 5 cores as the Ryzen AI Max+ 395, boosting to 5.2 GHz instead of 5.1. The iGPU becomes the Radeon 8065S — the same 40 RDNA 3.5 compute units, clocked at 3.0 GHz instead of 2.9. The NPU moves from 50 to 55 TOPS, which matters as little as it did before — LLM decode is bound by memory bandwidth, not TOPS.

The memory is the actual product. 192GB of LPDDR5X-8533 on the same 256-bit bus delivers 273 GB/s of bandwidth, of which up to 160GB can be allocated to the GPU. Against the 128GB model’s LPDDR5X-8000 that’s a 6.6% bandwidth bump — from 256 GB/s — and a 50% capacity jump. Framework also throws in a pre-installed Noctua NF-A12x25 fan and an open-ended PCIe x4 slot, which are nice and irrelevant to the buying decision.

Here’s the comparison that matters:

128GB model192GB modelDelta
Price (DIY)$3,449$6,799+$3,350
Memory128GB LPDDR5X-8000192GB LPDDR5X-8533+64GB
Bandwidth256 GB/s273 GB/s+6.6%
Usable GPU pool (Linux)~120GB~160GB+40GB
gpt-oss-120b decode31–56 tok/s~33–60 tok/s (est.)~7%
$ per GB of memory$26.9$35.4marginal GB: $52

Two numbers deserve a hard look. First, the price per gigabyte goes the wrong way: the 64GB you’re adding costs $52/GB, nearly double the $26.9/GB the base machine charges. That’s the DRAM crisis showing up in the invoice — Framework warned pricing would be a “pretty substantial jump,” and LPDDR5X supply is exactly why there’s only one preorder batch.

Second, the usable window is 40GB, not 64GB. On a 128GB Strix Halo machine running Linux with the GTT fix, the GPU can address roughly 120GB; Framework quotes 160GB allocatable on the 192GB config. Every model between ~120GB and ~160GB is the entire value proposition of this machine.

The 40GB window: what now fits that didn’t

That window turns out to be prime real estate. Three of the most-requested open-weight models land inside it:

ModelQuantSizeFits 128GB machine?Fits 192GB machine?Expected speed
Qwen3-235B-A22BQ4_K_M~141GB❌ (Q2/Q3 only)✅ with ~19GB for context~11 tok/s
MiniMax M3 (428B MoE)Q2_K_XL~143GB❌✅ tightsingle-digit tok/s
DeepSeek R1 671BUD-IQ1_S (1.58-bit)~131GB❌✅ weights fit, context tightsingle-digit tok/s
gpt-oss-120bMXFP4~61GB✅✅ (no change)31–56 tok/s
GLM 5.2 (744B MoE)UD-IQ1 (smallest)~217GB❌❌ still doesn’t fit—

The Qwen3-235B row is the headline. On a 128GB machine, the 235B MoE runs only at 2-bit (Unsloth’s 88GB dynamic quant) or squeezed 3-bit — real quality loss for a model you bought the box to run. At Q4_K_M, the quant most people consider full-strength, it’s ~141GB and simply doesn’t fit. On the 192GB machine it fits with ~19GB left for KV cache, and measured Strix Halo throughput on this model is around 11 tok/s — consistent with the bandwidth math, since 22B active parameters at 4-bit means reading ~12GB per token from a pool that moves 273 GB/s.

The honest row is the last one. If you’re eyeing this machine for GLM 5.2, the best open-weight coding model of mid-2026, it does not help you. The smallest coherent quant is ~217GB; even 192GB of RAM with a 160GB GPU pool is 57GB short. Framework’s own marketing leans on DeepSeek-V4-Flash at Q8 — a vendor-picked fit, not the model you were probably thinking of.

And the MiniMax M3 and DeepSeek R1 rows come with the same caveat we gave at 96GB: a 1.58-to-2-bit quant of a frontier MoE is a capability demo. It runs, it’s genuinely impressive that it runs, and the quality loss versus the hosted full-precision model is real. You’re buying the ability to do it at all, air-gapped, on 120–140W of wall power — not an experience that competes with the API.

What you’re not buying: speed

This is the same lesson as the 64GB vs 128GB decision one tier down, and it bears repeating because $3,350 is riding on it. Every model that already fits the 128GB machine runs at effectively the same speed on the 192GB machine. The bandwidth bump from LPDDR5X-8533 is 6.6%, which moves gpt-oss-120b from the ~31 tok/s ServeTheHome measured (tuned llama.cpp runs reach 55–56) to maybe 33. You will not feel it.

Capacity doesn’t make tokens faster; it makes bigger models possible, and bigger models are slower. The machine’s ceiling on a 141GB model is set by the same arithmetic as every Strix Halo box: 273 GB/s divided by bytes-read-per-token. At ~11 tok/s, Qwen3-235B Q4 is usable for chat and painful for agentic loops that burn thousands of tokens per step. Check your own model-and-context combination in the VRAM calculator before you commit $6,799 to a number you haven’t sat with.

Setup is also unchanged. Out of the box, Linux will strand most of the pool behind the default GTT limit — models over ~64GB fail to load with out-of-memory errors while free -h shows plenty. The fix is the same kernel-parameter carve-out we documented for the 128GB boxes, scaled up; after it, dmesg should report the bigger pool:

$ sudo dmesg | grep "amdgpu.*GTT"
[    3.211] [drm] amdgpu: 160000M of GTT memory ready.

On Windows, AMD’s Variable Graphics Memory caps allocation at 75% of system RAM — 144GB on this machine — so Linux remains the way to reach the full quoted 160GB.

The alternatives at $6,799

The 192GB Framework sits in an awkward price band, and the comparison shopping is where the preorder decision actually gets made.

Mac Studio M5 Ultra 256GB — $10,799. Four thousand dollars more, and it changes the category: 256GB of unified memory at 1.2 TB/s — 4.4× the bandwidth — turns the same Qwen3-235B Q4 from an ~11 tok/s experience into a comfortable daily driver, and the 256GB pool holds it at Q6 with room to spare. We ran this machine against NVIDIA’s RTX PRO 6000 workstation card earlier this month; against the Framework, the question is simpler: if giant-model inference is the actual job, the Mac is slower to buy and faster to use. If $10,799 is out of reach, the Framework is the cheapest new machine that holds this model class at all — that’s its niche, and it’s a real one.

The 128GB Strix Halo boxes — $3,449–$3,649. The Framework 128GB at $3,449 and the GMKtec EVO-X2 128GB at $3,499–$3,649 run everything through gpt-oss-120b at identical speed to the 192GB machine. If your largest model is 120B-class — and for most home labs in 2026, it is — the extra $3,350 buys you nothing you’ll use. Our full 128GB comparison covers that tier.

A used RTX 3090 build — ~$1,150–$1,350 for the card. If your real workload is 8–35B models for coding and chat, none of these unified-memory boxes is the right platform. A used RTX 3090 has 3.7× the memory bandwidth (936 GB/s) and remains the value king for everything that fits in 24GB.

The DGX Spark — $4,699. Same 273 GB/s bandwidth, 128GB of memory, $1,250 more than the Framework 128GB, and it can’t hold the 141GB models that justify the 192GB tier. We’ve covered why it loses this fight already; nothing about the 192GB Framework changes it.

Renting first. If you’re not sure the ~11 tok/s giant-MoE experience is something you’ll live with, spend $20 finding out before you spend $6,799. Vast.ai rents RTX 3090s from $0.07/hr and bigger iron by the hour — load the exact quant you’re considering and feel the latency yourself.

Order now, or wait for batch 2?

The case for ordering now is supply, not product. Framework says memory inventory constrains it to one batch; DRAM pricing has been climbing all year with no forecast relief before late 2027, so batch 2 — whenever it exists — is at least as likely to cost more as less. If you’re the buyer this machine is for, waiting probably doesn’t save you money.

The case for waiting is everything else. No independent reviews exist yet; first units ship in November. The 160GB GPU-allocation figure is Framework’s, not yet community-verified the way the 128GB machines’ ~120GB GTT ceiling is. And Apple’s M5 Ultra 512GB configuration was slated to open orders in late October — if it lands anywhere near its expected pricing, it compresses the space above the Framework further.

What to actually buy

Prices as of October 2026, all taken from the comparison above:

Your situationThe machinePriceWhere
You specifically need Qwen3-235B Q4 / R1-class resident at home, cheapest possibleFramework Desktop 192GB$6,799frame.work (batch 2 waitlist)
Giant models are the daily job and budget stretchesMac Studio M5 Ultra 256GB$10,799Check price
Your ceiling is 120B-class MoE (most people)Framework Desktop 128GB or GMKtec EVO-X2$3,449–$3,649Check price
You run 8–35B models for coding/chatUsed RTX 3090 24GB$1,150–$1,350Check price
Undecided — want to feel ~11 tok/s before buyingRented GPU, from $0.07/hrpay per hourVast.ai

FAQ

Is the 192GB Framework Desktop faster than the 128GB one? Not meaningfully. Same 256-bit bus, 6.6% more bandwidth from LPDDR5X-8533 (273 vs 256 GB/s). Any model that fits both machines runs within a couple of tokens per second on either. You’re paying for capacity.

What can the 192GB model run that the 128GB can’t? Models between ~120GB and ~160GB: Qwen3-235B-A22B at Q4_K_M (~141GB), MiniMax M3 at Q2_K_XL (~143GB), and DeepSeek R1 671B at 1.58-bit (~131GB). The 128GB machine caps out around 120GB of usable GPU pool on Linux.

Can it run GLM 5.2? No. The smallest coherent GGUF is ~217GB — 57GB past even the 192GB machine’s 160GB GPU pool. That model still needs a 256GB+ box or the API.

How fast is Qwen3-235B on this machine? Expect roughly 11 tok/s at Q4 — measured on Strix Halo silicon, and consistent with the bandwidth ceiling (22B active parameters × 4-bit ≈ 12GB read per token from a 273 GB/s pool). Fine for chat, slow for agents.

Why is it $6,799 when the 128GB is $3,449? LPDDR5X pricing. The DRAM super-cycle has roughly quadrupled memory prices since mid-2025, the 192GB config’s marginal 64GB works out to ~$52/GB, and constrained inventory is why Framework limited preorders to one batch.

For running Ollama or llama.cpp on whichever box you pick, aifoss.dev’s self-hosting guides cover the software stack; if the plan is pointing Cursor or Cline at a local 120B+ backend, aicoderscope.com covers the BYOK setup side.

Sources

Last updated October 4, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.