Framework Desktop 192GB Preorder in 2026: What $6,799 Actually Unlocks Over the 128GB Model
TL;DR: Framework’s 192GB Desktop (Ryzen AI Max+ PRO 495) opened preorders September 30 at $6,799 DIY — nearly double the 128GB model — and batch 1 sold out anyway. The extra $3,350 buys a 40GB capacity window (a ~160GB GPU pool vs ~120GB) and zero extra speed. That window holds exactly three models most people want: Qwen3-235B at Q4, MiniMax M3 at 2-bit, and DeepSeek R1 at 1.58-bit.
| Framework Desktop 192GB | Framework Desktop 128GB | Mac Studio M5 Ultra 256GB | |
|---|---|---|---|
| Best for | 130–160GB MoE models at ~11 tok/s | Everything up to gpt-oss-120b | The same giant models, 4.4× faster |
| Price / Cost | $6,799 DIY (ships Nov 2026) | $3,449 | $10,799 |
| The catch | Same 273 GB/s — capacity, not speed | ~120GB usable ceiling | $4,000 more, macOS-only stack |
Honest take: If you weren’t already refreshing frame.work waiting for this, the 128GB model at $3,449 is still the right buy. The 192GB config is for the specific person who needs Qwen3-235B at Q4 or DeepSeek R1 resident at home, knows it will decode at ~11 tok/s, and wants it anyway.
Framework opened preorders on September 30, 2026 for the machine it had been teasing all fall: a Desktop built on AMD’s Ryzen AI Max+ PRO 495 (“Gorgon Halo”) with 192GB of soldered LPDDR5X-8533. The DIY Edition is $6,799; a prebuilt with a 2TB NVMe drive and Fedora preinstalled is $7,449, with first units shipping in November 2026. Framework says constrained memory inventory limits it to a single initial batch — and batch 1 sold out anyway, sky-high price and all.
When we compared the 128GB Framework Desktop against the GMKtec EVO-X2 in late September, this machine was still “coming soon” with no price, and our standing advice was don’t wait for it. Now there’s a price, a ship date, and a sellout. The advice needs numbers behind it.
What $6,799 buys, in hardware terms
The chip upgrade is real but small. The PRO 495 keeps the same 16 Zen 5 cores as the Ryzen AI Max+ 395, boosting to 5.2 GHz instead of 5.1. The iGPU becomes the Radeon 8065S — the same 40 RDNA 3.5 compute units, clocked at 3.0 GHz instead of 2.9. The NPU moves from 50 to 55 TOPS, which matters as little as it did before — LLM decode is bound by memory bandwidth, not TOPS.
The memory is the actual product. 192GB of LPDDR5X-8533 on the same 256-bit bus delivers 273 GB/s of bandwidth, of which up to 160GB can be allocated to the GPU. Against the 128GB model’s LPDDR5X-8000 that’s a 6.6% bandwidth bump — from 256 GB/s — and a 50% capacity jump. Framework also throws in a pre-installed Noctua NF-A12x25 fan and an open-ended PCIe x4 slot, which are nice and irrelevant to the buying decision.
Here’s the comparison that matters:
| 128GB model | 192GB model | Delta | |
|---|---|---|---|
| Price (DIY) | $3,449 | $6,799 | +$3,350 |
| Memory | 128GB LPDDR5X-8000 | 192GB LPDDR5X-8533 | +64GB |
| Bandwidth | 256 GB/s | 273 GB/s | +6.6% |
| Usable GPU pool (Linux) | ~120GB | ~160GB | +40GB |
| gpt-oss-120b decode | 31–56 tok/s | ~33–60 tok/s (est.) | ~7% |
| $ per GB of memory | $26.9 | $35.4 | marginal GB: $52 |
Two numbers deserve a hard look. First, the price per gigabyte goes the wrong way: the 64GB you’re adding costs $52/GB, nearly double the $26.9/GB the base machine charges. That’s the DRAM crisis showing up in the invoice — Framework warned pricing would be a “pretty substantial jump,” and LPDDR5X supply is exactly why there’s only one preorder batch.
Second, the usable window is 40GB, not 64GB. On a 128GB Strix Halo machine running Linux with the GTT fix, the GPU can address roughly 120GB; Framework quotes 160GB allocatable on the 192GB config. Every model between ~120GB and ~160GB is the entire value proposition of this machine.
The 40GB window: what now fits that didn’t
That window turns out to be prime real estate. Three of the most-requested open-weight models land inside it:
| Model | Quant | Size | Fits 128GB machine? | Fits 192GB machine? | Expected speed |
|---|---|---|---|---|---|
| Qwen3-235B-A22B | Q4_K_M | ~141GB | ❌ (Q2/Q3 only) | ✅ with ~19GB for context | ~11 tok/s |
| MiniMax M3 (428B MoE) | Q2_K_XL | ~143GB | ❌ | ✅ tight | single-digit tok/s |
| DeepSeek R1 671B | UD-IQ1_S (1.58-bit) | ~131GB | ❌ | ✅ weights fit, context tight | single-digit tok/s |
| gpt-oss-120b | MXFP4 | ~61GB | ✅ | ✅ (no change) | 31–56 tok/s |
| GLM 5.2 (744B MoE) | UD-IQ1 (smallest) | ~217GB | ❌ | ❌ still doesn’t fit | — |
The Qwen3-235B row is the headline. On a 128GB machine, the 235B MoE runs only at 2-bit (Unsloth’s 88GB dynamic quant) or squeezed 3-bit — real quality loss for a model you bought the box to run. At Q4_K_M, the quant most people consider full-strength, it’s ~141GB and simply doesn’t fit. On the 192GB machine it fits with ~19GB left for KV cache, and measured Strix Halo throughput on this model is around 11 tok/s — consistent with the bandwidth math, since 22B active parameters at 4-bit means reading ~12GB per token from a pool that moves 273 GB/s.
The honest row is the last one. If you’re eyeing this machine for GLM 5.2, the best open-weight coding model of mid-2026, it does not help you. The smallest coherent quant is ~217GB; even 192GB of RAM with a 160GB GPU pool is 57GB short. Framework’s own marketing leans on DeepSeek-V4-Flash at Q8 — a vendor-picked fit, not the model you were probably thinking of.
And the MiniMax M3 and DeepSeek R1 rows come with the same caveat we gave at 96GB: a 1.58-to-2-bit quant of a frontier MoE is a capability demo. It runs, it’s genuinely impressive that it runs, and the quality loss versus the hosted full-precision model is real. You’re buying the ability to do it at all, air-gapped, on 120–140W of wall power — not an experience that competes with the API.
What you’re not buying: speed
This is the same lesson as the 64GB vs 128GB decision one tier down, and it bears repeating because $3,350 is riding on it. Every model that already fits the 128GB machine runs at effectively the same speed on the 192GB machine. The bandwidth bump from LPDDR5X-8533 is 6.6%, which moves gpt-oss-120b from the ~31 tok/s ServeTheHome measured (tuned llama.cpp runs reach 55–56) to maybe 33. You will not feel it.
Capacity doesn’t make tokens faster; it makes bigger models possible, and bigger models are slower. The machine’s ceiling on a 141GB model is set by the same arithmetic as every Strix Halo box: 273 GB/s divided by bytes-read-per-token. At ~11 tok/s, Qwen3-235B Q4 is usable for chat and painful for agentic loops that burn thousands of tokens per step. Check your own model-and-context combination in the VRAM calculator before you commit $6,799 to a number you haven’t sat with.
Setup is also unchanged. Out of the box, Linux will strand most of the pool behind the default GTT limit — models over ~64GB fail to load with out-of-memory errors while free -h shows plenty. The fix is the same kernel-parameter carve-out we documented for the 128GB boxes, scaled up; after it, dmesg should report the bigger pool:
$ sudo dmesg | grep "amdgpu.*GTT"
[ 3.211] [drm] amdgpu: 160000M of GTT memory ready.
On Windows, AMD’s Variable Graphics Memory caps allocation at 75% of system RAM — 144GB on this machine — so Linux remains the way to reach the full quoted 160GB.
The alternatives at $6,799
The 192GB Framework sits in an awkward price band, and the comparison shopping is where the preorder decision actually gets made.
Mac Studio M5 Ultra 256GB — $10,799. Four thousand dollars more, and it changes the category: 256GB of unified memory at 1.2 TB/s — 4.4× the bandwidth — turns the same Qwen3-235B Q4 from an ~11 tok/s experience into a comfortable daily driver, and the 256GB pool holds it at Q6 with room to spare. We ran this machine against NVIDIA’s RTX PRO 6000 workstation card earlier this month; against the Framework, the question is simpler: if giant-model inference is the actual job, the Mac is slower to buy and faster to use. If $10,799 is out of reach, the Framework is the cheapest new machine that holds this model class at all — that’s its niche, and it’s a real one.
The 128GB Strix Halo boxes — $3,449–$3,649. The Framework 128GB at $3,449 and the GMKtec EVO-X2 128GB at $3,499–$3,649 run everything through gpt-oss-120b at identical speed to the 192GB machine. If your largest model is 120B-class — and for most home labs in 2026, it is — the extra $3,350 buys you nothing you’ll use. Our full 128GB comparison covers that tier.
A used RTX 3090 build — ~$1,150–$1,350 for the card. If your real workload is 8–35B models for coding and chat, none of these unified-memory boxes is the right platform. A used RTX 3090 has 3.7× the memory bandwidth (936 GB/s) and remains the value king for everything that fits in 24GB.
The DGX Spark — $4,699. Same 273 GB/s bandwidth, 128GB of memory, $1,250 more than the Framework 128GB, and it can’t hold the 141GB models that justify the 192GB tier. We’ve covered why it loses this fight already; nothing about the 192GB Framework changes it.
Renting first. If you’re not sure the ~11 tok/s giant-MoE experience is something you’ll live with, spend $20 finding out before you spend $6,799. Vast.ai rents RTX 3090s from $0.07/hr and bigger iron by the hour — load the exact quant you’re considering and feel the latency yourself.
Order now, or wait for batch 2?
The case for ordering now is supply, not product. Framework says memory inventory constrains it to one batch; DRAM pricing has been climbing all year with no forecast relief before late 2027, so batch 2 — whenever it exists — is at least as likely to cost more as less. If you’re the buyer this machine is for, waiting probably doesn’t save you money.
The case for waiting is everything else. No independent reviews exist yet; first units ship in November. The 160GB GPU-allocation figure is Framework’s, not yet community-verified the way the 128GB machines’ ~120GB GTT ceiling is. And Apple’s M5 Ultra 512GB configuration was slated to open orders in late October — if it lands anywhere near its expected pricing, it compresses the space above the Framework further.
What to actually buy
Prices as of October 2026, all taken from the comparison above:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| You specifically need Qwen3-235B Q4 / R1-class resident at home, cheapest possible | Framework Desktop 192GB | $6,799 | frame.work (batch 2 waitlist) |
| Giant models are the daily job and budget stretches | Mac Studio M5 Ultra 256GB | $10,799 | Check price |
| Your ceiling is 120B-class MoE (most people) | Framework Desktop 128GB or GMKtec EVO-X2 | $3,449–$3,649 | Check price |
| You run 8–35B models for coding/chat | Used RTX 3090 24GB | $1,150–$1,350 | Check price |
| Undecided — want to feel ~11 tok/s before buying | Rented GPU, from $0.07/hr | pay per hour | Vast.ai |
FAQ
Is the 192GB Framework Desktop faster than the 128GB one? Not meaningfully. Same 256-bit bus, 6.6% more bandwidth from LPDDR5X-8533 (273 vs 256 GB/s). Any model that fits both machines runs within a couple of tokens per second on either. You’re paying for capacity.
What can the 192GB model run that the 128GB can’t? Models between ~120GB and ~160GB: Qwen3-235B-A22B at Q4_K_M (~141GB), MiniMax M3 at Q2_K_XL (~143GB), and DeepSeek R1 671B at 1.58-bit (~131GB). The 128GB machine caps out around 120GB of usable GPU pool on Linux.
Can it run GLM 5.2? No. The smallest coherent GGUF is ~217GB — 57GB past even the 192GB machine’s 160GB GPU pool. That model still needs a 256GB+ box or the API.
How fast is Qwen3-235B on this machine? Expect roughly 11 tok/s at Q4 — measured on Strix Halo silicon, and consistent with the bandwidth ceiling (22B active parameters × 4-bit ≈ 12GB read per token from a 273 GB/s pool). Fine for chat, slow for agents.
Why is it $6,799 when the 128GB is $3,449? LPDDR5X pricing. The DRAM super-cycle has roughly quadrupled memory prices since mid-2025, the 192GB config’s marginal 64GB works out to ~$52/GB, and constrained inventory is why Framework limited preorders to one batch.
Recommended Gear
- GMKtec EVO-X2 128GB — the 128GB Strix Halo alternative that covers most workloads for roughly half the price
- Apple Mac Studio M5 Ultra — the faster path to the same giant-model tier, at a premium
- Used RTX 3090 24GB — still the bandwidth-per-dollar king for models under 24GB
For running Ollama or llama.cpp on whichever box you pick, aifoss.dev’s self-hosting guides cover the software stack; if the plan is pointing Cursor or Cline at a local 120B+ backend, aicoderscope.com covers the BYOK setup side.
Sources
- Framework Desktop with Ryzen AI Max+ PRO 495 and 192GB memory starts at $6,799 — VideoCardz
- Framework Desktop With AMD Ryzen AI Max+ PRO 495 Starts Out At $6,799 — Phoronix
- AMD Ryzen AI Max+ Pro 495 Framework Desktop Batch 1 Sells Out Despite Sky-High Prices — TechPowerUp
- Framework Desktop with 192GB unified memory gets pre-order date — NotebookCheck
- Framework Desktop with 192GB unified memory opens pre-orders September 30 — TweakTown
- Framework Desktop gets a 192GB Ryzen AI Max config built for local LLMs — Hardware Busters
- Qwen3 235B A22B hardware requirements (Q4_K_M ~141GB) — llmrun.dev
- DeepSeek R1 Dynamic 1.58-bit (131GB) — Open WebUI docs
- AMD Strix Halo (Ryzen AI Max+ 395) GPU performance — llm-tracker.info
- Beelink GTR9 Pro review: gpt-oss-120b at ~31 tok/s, 120–128W — ServeTheHome
- RAM Price Index 2026 — Tom’s Hardware
Last updated October 4, 2026. Prices and specs change; verify current rates before purchasing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.