Strata Runs a 125B Model on a 12GB Gaming GPU: Real Benchmarks, the 64GB RAM Catch, and Who Should Install It (2026)

strataqwenmoelocal-llminference-enginegpu

TL;DR: Strata is an MIT-licensed inference engine (21k GitHub stars, front-paged Hacker News in early October 2026) that runs the 125B-parameter Qwen3.8-Flash-Next MoE on a single 12GB gaming GPU — 94 tok/s on an RTX 5070 at Q2_0, per the developer’s own benchmarks. The catch: it wants 64GB of system RAM, ~80GB of SSD, and its fastest numbers come from an aggressive 2-bit quant. If you already own a gaming PC that qualifies, it’s free frontier-class AI. If you’d have to buy the RAM, the math gets ugly fast.

Strata on a 12-16GB gaming GPUQwen3.8-27B on a used RTX 3090Flash-Next on a 128GB unified box
Best forPC you already own: 12GB+ GPU, 32-64GB RAMBest quality-per-dollar if buying hardware todayFull-quality Flash-Next (4-bit), 262K context
Price / Cost$0 if you qualify; 64GB DDR5 is $680+ if you don’t$1,399-$1,450 used (Oct 2026)GMKtec EVO-X2 128GB $3,499-$3,649
The catchQ2/Q3 quants, one model family, young engine (v0.1.x crashes)27B-class ceiling, not 125BPrice nearly doubled since launch

Honest take: Install Strata tonight if your PC already has a 12GB+ GPU and 32GB+ RAM — it’s the most impressive free upgrade in local AI this year. Don’t buy a single component for it: at October 2026 DRAM prices, a 64GB kit costs more than a used RTX 3060, and a used RTX 3090 running Qwen3.8-27B is still the better purchase.

Five weeks ago, our Qwen3.8-Flash-Next hardware guide said a 24GB card “is not a path to it — full stop,” and pointed everyone below the 96GB tier back to Qwen3.8-27B. Strata is the first thing that genuinely changes that verdict — not by shrinking the model, but by rebuilding the memory hierarchy around what this specific MoE architecture actually touches per token.

It went viral for a reason: “125B parameters on a 12GB card at 94 tokens per second” sounds like a scam. It isn’t — the benchmarks are real, reproduced by community issue reports on different cards. But the headline leaves out three things you’ll care about before installing: the RAM requirement, the quantization quality tax, and what a v0.1.x engine does when you push a 100K-token prompt through it. All three below.

What Strata actually is

The verifiable facts, from the GitHub repo (checked October 11, 2026):

SpecStrata
LicenseMIT (engine; each model keeps its own license)
Current versionv0.1.40.x, October 2026
Stars~21,000 (up from ~16,600 at the HN launch a week earlier)
Built onllama.cpp / GGML, heavily modified
OSWindows 10/11, Linux (Docker included)
GPUsNVIDIA RTX 20/30/40/50-series, AMD RX 6800/6900, RX 7700 XT and newer — 12GB+ VRAM
ExperimentalIntel Arc (Linux, source build), Strix Halo (Linux, source build), pre-AVX2 CPUs
System RAM32GB minimum; 64GB “runs every model size”
Disk~80GB free, SSD recommended (~70GB model download)
ModelsQwen3.8-Flash-Next only (Q2_0 / IQ2_XS / IQ3_XXS / IQ3_S, a Coder variant, Unsloth 4-bit experimental)
APIOpenAI- and Anthropic-compatible on localhost, browser chat UI at 127.0.0.1:8080, MCP server

Install is deliberately non-technical: START-HERE.bat on Windows, ./setup.sh on Linux. That’s a different ambition from llama.cpp, and it shows in who’s filing the GitHub issues — a lot of gamers running their first local model.

Note the one-model catch before you get attached: Strata is not a general engine. It runs Qwen3.8-Flash-Next and its variants, period. If you want to swap in Gemma or Llama, you’re back to Ollama or llama.cpp.

How 125B fits in 12GB — and why “125B” undersells it

Qwen3.8-Flash-Next is the weirdest open-weight release of 2026, and Strata is essentially a purpose-built exploit of its architecture. The model card math, from our Flash-Next deep dive: ~180B parameters on disk (a 125B main model — the number in Strata’s headline — plus a 51B n-gram “engram” lookup table and a ~4B multi-token-prediction head), but only ~6B active per token. Its 48 layers carry 512 experts each — 24,576 tiny experts total — and each token routes through just 10 of them.

Strata splits that across four tiers of your machine:

  1. GPU VRAM holds the attention layers plus the few thousand most frequently routed experts — the “hot set.”
  2. System RAM holds the full expert pool; the CPU computes cold-expert contributions in parallel with the GPU rather than just feeding it.
  3. SSD holds the engram lookup table. Those 51B parameters are sparse reads — the model looks rows up rather than computing through them — so they never need to touch RAM.
  4. A draft head speculates 1.6-1.8× ahead (the same trick as speculative decoding in llama.cpp, built into the model), and long prompts prefill in 8,192-token chunks at 1,000+ tok/s.

llama.cpp can do pieces of this (--n-cpu-moe expert offload, and it streams the engram table from disk too — that’s how the Strix Halo deployments work). Strata’s difference is aggression and defaults: hot-expert caching by routing frequency, CPU-as-coprocessor instead of CPU-as-overflow, and the SSD tier wired in out of the box.

The numbers, and how much to trust them

The developer’s published benchmarks (README, engine v0.1.26-0.1.36; 32K-token prompts, 4K-token answers — honest test lengths, not 512-token sprints):

GPUCPU / RAMQuantDecode (tok/s)Prefill (tok/s)
RTX 5070 12GBRyzen 5 7600 / 64GBQ2_0942,650
RTX 5070 12GBRyzen 5 7600 / 64GBIQ2_XS792,090
RTX 5070 12GBRyzen 5 7600 / 64GBIQ3_XXS621,750
RTX 5070 12GBRyzen 5 7600 / 64GBIQ3_S531,620
RTX 5070 12GBRyzen 5 7600 / 64GBCoder552,180
RX 9070 XT 16GBRyzen 9 3900X / 47GBQ2_0601,160
RX 9070 XT 16GBRyzen 9 3900X / 47GBIQ2_XS521,110

These are self-reported, but community issue reports corroborate the shape: an RTX 4090 owner with 192GB RAM measured 70-85 tok/s sustained at IQ3_S on Linux (issue #29), and an RTX 5070 Ti 16GB owner reported ~94 tok/s at IQ3_S on Windows (issue #922). The README’s RTX 3090 figure of 100-140 tok/s is explicitly an estimate, not a measurement — treat it as such.

Two things the headline number hides:

The 94 tok/s is Q2_0. That’s a 2-bit quantization — the tier our quantization quality-loss guide tells you to treat as a last resort on dense models. MoE models with huge total parameter counts degrade more gracefully at low bits than dense ones, and a 2-bit 125B MoE genuinely outclasses a 4-bit 8B — but the honest comparison for quality is the IQ3_S row at 53 tok/s, and the full-quality 4-bit path (Unsloth UD-Q4_K_XL, ~94GB) drops to 7-8.5 tok/s on a 64GB machine because most of it reads from SSD. You don’t get 94 tok/s and 4-bit quality on 12GB. Pick one.

The CPU and RAM matter as much as the GPU. Both benchmark rigs pair the GPU with 47-64GB of RAM, and the AMD card’s lower numbers track its older CPU (Ryzen 9 3900X) as much as the GPU. Cold-expert math runs on your CPU; RAM bandwidth is the second axle. This is not a workload where the GPU does everything and the rest of the box is a power supply.

There’s also a Coder variant with half the experts removed so the whole pool fits 32GB RAM — the README itself warns it’s weaker outside code, including on non-English text. Reasonable trade if coding is the use case; see local coding backends for where a localhost OpenAI-compatible endpoint like Strata’s plugs into Cline or Cursor.

When it breaks: a v0.1.x engine in the wild

Strata is six months from first commit and it shows. Two failure modes worth knowing before you file your own issue:

The long-prompt crash. On prompts in the 54K-226K-token range, multiple users hit a hard CUDA/HIP fault and engine restart (issue #1468, open as of October 11):

prefill copy_i32: an illegal memory access was encountered

No confirmed root cause yet. The workarounds that reporters verified: cap the prefill chunk with --prefill auto:16384 (stable in the original reporter’s synthetic runs, not a guaranteed fix), and STRATA_PF_STEP_SYNC=1 to reduce the failure rate — one tester logged zero faults in 40 runs at STRATA_PF_STEP_SYNC=2, at some prefill speed cost. If your use case is feeding entire codebases into the 262K context, you’re on the engine’s bleeding edge; expect restarts.

The mid-generation freeze (fixed, instructively). Early versions would stop producing tokens at random — GPU pinned at 100% drawing ~91W, nothing happening. The cause was a race in the CPU expert pool: a worker thread waking late could steal a job from the next batch, and the engine lost count and waited forever. v0.1.12 tied jobs to batches and added a watchdog (issue #29). If you see freeze reports in old threads, they’re mostly this — one more reason to stay current via UPDATE.bat/update.sh.

Neither of these is disqualifying. Both are what “v0.1.40” means.

What a Strata rig costs in October 2026 — the part the hype skips

Here’s where the viral framing dies on contact with DRAM crisis pricing. The pitch writes itself in 2024 prices: “a $549 midrange GPU and $130 of RAM.” In October 2026, with street prices verified against our tracking:

ComponentOct 2026 street priceNote
RTX 5070 12GB$799-$900 (MSRP $549)The benchmark card
RTX 4070 12GB$485-$585Cheapest current NVIDIA entry that qualifies
Used RTX 3060 12GB$260-$296Qualifies on paper; no published Strata benchmark — expect well under the 5070’s numbers on 360 GB/s
64GB DDR5 kit$680-$1,070More than the RTX 4070 itself
1TB NVMe SSD headroom~80GB of itYou likely have this

So a from-scratch “budget” Strata build — RTX 4070 + 64GB DDR5 + CPU/board/PSU — lands at $1,900-$2,500, with the RAM as the single biggest line item after the GPU. For that money you could instead buy a used RTX 3090 24GB ($1,399-$1,450, October 2026) plus a 32GB kit, run Qwen3.8-27B at ~41 tok/s at honest 4-bit quality with the whole GGUF resident in VRAM, and keep a clean upgrade path to every other model family. Our September verdict on the 3090 survives Strata.

The economics only flip when the hardware is sunk cost. A gamer who bought a 12-16GB card for games and has 32GB of RAM from before the price surge pays $0 and gets a 125B-class model at usable speed. That’s the actual audience, and it’s enormous — it’s why this repo gained ~4,400 stars in a week. Run the numbers for your own setup with our VRAM calculator.

What to actually buy

Prices as of October 2026, all verified in the comparison above:

Your situationWhat to doPriceWhere
Own a 12GB+ GPU and 32GB+ RAMBuy nothing — install Strata$0github.com/Niko1221/strata
Own a 12GB+ GPU but only 16GB RAM64GB DDR5 kit — only if you want this specific model$680-$1,070Check price
Buying a GPU today for local AIUsed RTX 3090 24GB + Qwen3.8-27B, not a 12GB card for Strata$1,399-$1,450Check price
Want full-quality Flash-Next (4-bit, 262K ctx)GMKtec EVO-X2 128GB unified memory$3,499-$3,649Check price
Undecided — test the workload on rented hardware firstRented GPU, RTX 3090 from $0.07/hrpay per hourVast.ai

If you just want to taste Flash-Next’s output quality before touching any of this, the hosted API ran $0.15/M input and $0.47/M output on OpenRouter as of September 2026 — an evening of testing costs less than a coffee.

FAQ

Does Strata run models other than Qwen3.8-Flash-Next? No. It’s a single-model engine: Flash-Next in four quant tiers, a code-focused variant with half the experts removed, a faster fine-tune, and experimental Unsloth 4-bit GGUFs. For anything else, use llama.cpp or Ollama.

Is 8GB of VRAM enough? Not supported. The floor is 12GB, and the supported list starts at RTX 20-series and AMD RX 6800. Intel Arc and Strix Halo work experimentally via Linux source builds. On 8GB cards, see the best models for 8GB VRAM instead.

Is Q2_0 output actually good? Better than 2-bit instincts suggest — very large sparse MoEs tolerate low-bit quantization far better than dense models — but it is measurably below the 4-bit model, and the developer’s own quant ladder (Q2_0 → IQ3_S) exists because people noticed. For code, benchmark your own tasks against the Coder variant before trusting it.

Can I point Cursor or Cline at it? Yes — Strata serves OpenAI- and Anthropic-compatible APIs on localhost, which is exactly the BYOK local-backend pattern. Expect the long-prompt crash above if your agent stuffs 100K+ tokens of context per request; cap the prefill chunk first.

Does this make the 128GB unified-memory boxes pointless? No. A Strix Halo-class box runs the 4-bit model with the full 262K context at measured 22-82 tok/s — full quality, no Q2 compromise. Strata gets you the 2-3-bit tiers on hardware you already own. Different quality points, different budgets.

Sources

Last updated October 11, 2026. Prices and specs change; verify current rates before purchasing. Strata benchmark figures are developer-published unless noted as community-reported; no independent lab benchmarks exist yet.

Products linked in this guide:

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.