RTX Spark Pricing Is Real: Every Surface Laptop Ultra Config From $2,599 to $5,899, and the Only One That Makes Sense for Local AI

rtx-sparksurface-laptop-ultranvidialocal-llmunified-memorylaptophardwarebuying-guidepricing

TL;DR: RTX Spark pricing is no longer a rumor — Surface Laptop Ultra preorders opened October 7 at $2,599 to $5,899 across eight configs, shipping October 16. Only the $5,899 128GB config matters for local AI, and it costs $2,250 more than a GMKtec EVO-X2 with the same capacity. Buy it for CUDA-in-a-backpack or not at all.

Surface Laptop Ultra 128GBGMKtec EVO-X2 128GBMac Studio M5 Max 128GB
Best for128GB + CUDA in a laptopCheapest 128GB that runs 120B modelsFastest 128GB machine under $5,500
Memory / bandwidth128GB / ~300 GB/s128GB / 256 GB/s128GB / 614 GB/s
Price$5,899 (preorder, ships Oct 16)$3,499–$3,649$5,099
The catchv1 drivers, Windows takes a memory cutNo CUDA, slow prompt processingNo CUDA, it’s a desktop

Honest take: If the question is “cheapest way to run gpt-oss-120b at home,” the Surface loses to three machines that already shipped. The $5,899 config exists for exactly one buyer: someone who needs 100GB+ of model weights and the CUDA stack and a battery. Everyone else should read the price ladder below, notice that five of the eight configs can’t hold more model than a $1,300 used RTX 3090, and close the preorder tab.

Yesterday we told you not to preorder the Surface Laptop Ultra because Microsoft hadn’t named a price and the only public benchmark came from a leaked prototype whose CUDA path didn’t work. Half of that changed within hours: preorders went live October 7 at $2,599 to $5,899, with shipping starting October 16 (TechPowerUp, VideoCardz). The driver question is still open — nobody outside NVIDIA has published a retail-hardware llama.cpp run yet — but the money question now has exact answers, and they change the buying math enough to deserve their own breakdown.

This is that breakdown: what each of the eight configs costs, what each one can actually hold, why the memory bandwidth is identical from $2,599 to $5,899, and where the one interesting config lands against the 128GB machines you can already buy.

The full price ladder

Microsoft listed eight configurations at launch (TweakTown, cross-checked against Windows Central). Two N1X chip tiers exist — an 18-core CPU with a 5,120-CUDA-core GPU, and a 20-core CPU with 6,144 CUDA cores (HWBusters) — and the memory ceiling is what you’re really paying for:

ConfigMemorySSDPrice$/GB of memoryLocal AI verdict
18-core / 5,120 CUDA24GB512GB$2,599$108Pass — a used RTX 3090 holds the same models at 3× the bandwidth
18-core / 5,120 CUDA24GB1TB$2,999$125Pass — same ceiling, $400 for storage
18-core / 5,120 CUDA32GB1TB$3,299$103Pass — 27B-class ceiling
20-core / 6,144 CUDA32GB1TB$3,699$116Pass — faster prefill, same 32GB wall
20-core / 6,144 CUDA48GB1TB$3,999$83Weak — 70B Q4 fits but decodes at ~5-7 tok/s
20-core / 6,144 CUDA64GB1TB$4,299$67Marginal — gpt-oss-120b doesn’t fit after Windows takes its cut
20-core / 6,144 CUDA64GB2TB$4,699$73Marginal — the 2TB is the only path to big storage
20-core / 6,144 CUDA128GB1TB$5,899$46The only local-AI config

Three things jump out of that table.

The 64GB-to-128GB step costs $1,600. That sounds like Apple-grade memory pricing until you remember what DDR5 costs in late 2026: a 64GB desktop kit runs $680–$1,070 retail right now, and LPDDR5X soldered on a package is the more expensive kind. The upcharge is steep, not insane — but it does mean the 128GB config carries the best per-gigabyte price on the whole ladder, $46/GB versus $108/GB at the bottom. Microsoft priced the ladder to pull AI buyers upward, and for once the top rung is the rational one.

The 128GB config is 1TB-only. If you keep a few quantized 70B–120B models plus a ComfyUI model folder, 1TB fills fast, and there’s no 128GB/2TB option at any price. Budget for an external drive.

And pricing came in under the leaks. Earlier retail chatter pegged N1X systems “above $2,899” at entry and floated $7,000+ for the top config (VideoCardz). $2,599 at the bottom and $5,899 at the top undercut both — against the MacBook Pro 16” M5 Max at $6,999 for 128GB, the Surface is $1,100 cheaper for the same capacity (at less than half the bandwidth; more on that below).

The Surface isn’t alone anymore, either. MSI’s Prestige N16 Flip AI+ is listed for preorder at $3,299 with the full 20-core chip, 32GB, and 1TB, with a November 6 release (per Newegg’s own preorder listing — citation in Sources) — $400 less than Microsoft charges for the same silicon and memory. It doesn’t change the local-AI verdict (32GB is 32GB), but it signals that OEM competition on the same chip will compress prices fast. Microsoft also announced a Surface RTX Spark Dev Box — a mini PC with a fixed 128GB and a 100W thermal budget, bundling WSL2 GPU passthrough and full CUDA out of the box (Engadget, Tom’s Hardware). One outlet reports Dev Box preorders at $5,999 shipping in November (Pulse2), but Microsoft’s own page showed no price at announcement — treat that figure as unconfirmed until it’s on microsoft.com.

Every config has the same bandwidth, so the ladder only buys capacity

All eight configs use LPDDR5X at roughly 300 GB/s — the $2,599 machine and the $5,899 machine move weights at the same speed. (The DGX Spark desktop, built on the same GB10 silicon family, is rated at 273 GB/s.) Decode speed on local LLMs is bandwidth-bound, so the ceiling math is fixed across the whole ladder:

  • A dense 70B at Q4_K_M reads ~42.5GB per token: 300 ÷ 42.5 ≈ 7 tok/s ceiling, before real-world losses. The 48GB config can hold that model; nothing on this platform runs it pleasantly.
  • gpt-oss-120b is a MoE that reads ~3GB of active experts per token: 300 ÷ 3 ≈ 100 tok/s ceiling. This is the model class the 128GB config exists for.

How close does GB10-family hardware get to that MoE ceiling in practice? The best public data is the llama.cpp community benchmark thread for the DGX Spark, the Surface’s 273 GB/s desktop cousin (llama.cpp discussion #16578). On the stock DGX OS kernel, build 6816:

$ llama-bench -m gpt-oss-120b-mxfp4.gguf -fa 1 -ub 2048
| model                   |      size | test   |             t/s |
| ----------------------- | --------- | ------ | --------------- |
| gpt-oss 120B MXFP4 MoE  | 59.02 GiB | pp2048 | 1737.17 ± 81.66 |
| gpt-oss 120B MXFP4 MoE  | 59.02 GiB | tg32   |    45.87 ± 0.74 |

That’s 46 tok/s decode on a 120B-class model — about half the theoretical ceiling, which tells you the software stack still has headroom. The same thread proves the point: one user swapped the stock kernel for a mainline 6.17 build on Fedora 43 and the same benchmark jumped to 60.57 tok/s decode and 1,956 tok/s prefill — a 32% speedup from a kernel change alone. Expect the Surface’s Windows numbers to start below that and climb with driver updates, the same trajectory. For reference, the only leaked Surface prototype ran Qwen3.5 9B at 22.45 tok/s on Vulkan because CUDA wasn’t working at all — we covered that leak in detail yesterday. Nothing published since changes that caution: preordering still means betting the retail drivers land in better shape than the September prototype.

Windows takes a cut: 128GB isn’t 128GB

There’s a subtlety buried in NVIDIA’s own developer documentation that matters more on this machine than any spec-sheet number: Windows splits the unified pool, and the GPU can’t touch all of it (NVIDIA RTX Spark Porting Guide — Unified Memory Architecture).

The pool is carved into three regions: a dedicated carveout reserved for the GPU (what Windows reports as “dedicated GPU memory” — it’s ordinary DRAM, not separate VRAM), a shared region both CPU and GPU can use, and a CPU-only remainder. The shared region is sized by formula: post-carveout capacity minus 16GB, clamped between 50% and 80% of post-carveout capacity. On a 128GB machine with a hypothetical 16GB carveout, that’s 112GB post-carveout, and the 80% clamp caps the shared region at ~89.6GB — roughly 105GB GPU-touchable in total, not 128GB. Microsoft says it’s raising the GPU-accessible limit on high-memory unified systems in current Windows builds (Windows Experience Blog), but no exact retail numbers exist yet.

The practical fallout, config by config: on the 128GB machine, a ~60GB gpt-oss-120b plus KV cache fits with room to spare even after the split — fine. On the $4,299 64GB config, it doesn’t: apply the same formula and the GPU-accessible pool lands in the high-40s of GB, short of the model’s 59GB of weights before you allocate a single token of context. That’s why the table above calls 64GB “marginal” — the config that looks like the budget path to 120B-class models probably isn’t one on launch-day Windows. If a 70B Q4 at ~7 tok/s doesn’t excite you (it shouldn’t), the 64GB tiers buy nothing the 32GB tiers don’t. Check your exact model-plus-context arithmetic in our VRAM calculator before you pick a tier.

This is the same class of problem Strix Halo owners hit with Windows GTT limits — and the fix there was Linux. On RTX Spark, Linux support exists (DGX OS on the sibling desktops), but the Surface ships with Windows 11 and an agentic-Windows pitch; buying it to immediately install Linux defeats the point of buying a Surface.

What $5,899 buys against the 128GB machines that already shipped

So the real decision is the top config or nothing. Here’s where $5,899 lands in the 128GB-class field we’ve been benchmarking all year:

MachinePriceBandwidthgpt-oss-120b decodeCUDAPortable
Surface Laptop Ultra 128GB$5,899~300 GB/sunproven (sibling DGX Spark: 46–61 tok/s)YesYes
GMKtec EVO-X2 128GB$3,499–$3,649256 GB/s~31 tok/sNo (ROCm/Vulkan)No
ASUS Ascent GX10$3,099–$4,150273 GB/s46–61 tok/s (GB10, llama.cpp #16578)YesNo
NVIDIA DGX Spark$4,699273 GB/s46–61 tok/sYesNo
Mac Studio M5 Max 128GB$5,099614 GB/s65–88 tok/sNo (Metal/MLX)No
MacBook Pro 16” M5 Max 128GB$6,999614 GB/s65–88 tok/s classNoYes

Read that table cold and the Surface’s problem is obvious: it’s the second-most-expensive machine in the class, with the second-slowest proven silicon family, and the slowest option — the EVO-X2 — costs $2,250 less for the same capacity. The GX10 runs the same GB10-family numbers for $2,800 less if you can live without a screen and battery. And if raw decode speed per dollar is the metric, the Mac Studio M5 Max wins outright: $800 cheaper than the Surface, double the bandwidth, measured 65–88 tok/s on the same model class.

What the table can’t show is the two-word combination no other row has: CUDA, portable. The MacBook Pro is portable and fast but locks you out of the CUDA ecosystem — fine for llama.cpp and MLX, a wall the moment your workflow touches vLLM features, NVFP4 ComfyUI pipelines, CUDA-only fine-tuning stacks, or local coding stacks that assume an NVIDIA backend. The EVO-X2, GX10, and DGX Spark have no battery. If your actual requirement is “develop against CUDA with a 100GB model in the seat next to me,” the Surface Laptop Ultra 128GB is — for now — the only product on Earth that does it, and $5,899 is what a monopoly config costs. That’s a real buyer. It’s just a much rarer buyer than Microsoft’s launch-event framing suggests.

For everyone else, the machines above it in value already exist, ship today, and have months of community benchmarks behind them. Our $3K/$6K/$10K build guide covers how the desktop paths slot into real budgets.

What to actually buy

Prices as of October 2026, all verified in the comparison above:

Your situationThe machinePriceWhere
Need 128GB + CUDA + a battery, accept v1 driversSurface Laptop Ultra 128GB$5,899Microsoft Store (no affiliate relationship — go direct)
Cheapest 128GB that runs 120B-class MoE todayGMKtec EVO-X2 128GB~$3,499–$3,649Check price
Fastest 128GB machine under $5,500Mac Studio M5 Max 128GB$5,099Check price
CUDA desktop on the same GB10 silicon, shipping nowASUS Ascent GX10$3,099+Check price
Everything you run fits in 24GBUsed RTX 3090$1,190–$1,400Check price
Want to test 120B-class models before spending $3,500+Rented GPUfrom $0.07/hr (3090), $0.25/hr (5090)Vast.ai

FAQ

Is the $2,599 base config good value for local AI? No. 24GB of capacity at ~300 GB/s is strictly worse for inference than a used RTX 3090 at 936 GB/s for around $1,300 — the 3090 holds the same models and decodes roughly three times faster. The base Surface is a nice Arm laptop that happens to run small models; it is not an AI purchase.

Why skip the 48GB and 64GB middle tiers? Bandwidth. Every config decodes at the same ~300 GB/s, so the middle tiers only add capacity for dense 70B models that decode at ~7 tok/s — below comfortable reading speed. And after the Windows memory split, the 64GB tier likely can’t fit gpt-oss-120b at all. The ladder’s useful rungs are the bottom (as a general laptop) and the top (as an AI machine); the middle buys capacity you can’t enjoyably use.

Should I preorder the 128GB config now or wait? Wait for the first retail reviews showing llama.cpp or Ollama running on CUDA — not Vulkan — at a sane fraction of the 100 tok/s MoE ceiling. The last public data point (a September prototype) had the CUDA path failing outright, and the DGX Spark’s post-launch kernel-swap speedup shows this platform’s software is still moving fast. Shipping starts October 16; reviews will exist within days of that. Two weeks of patience against $5,899 is cheap insurance.

Does the MSI Prestige N16 Flip AI+ change anything? It undercuts Microsoft by $400 for the same 20-core chip at 32GB/1TB ($3,299, November 6 per Newegg’s listing), which is good news for RTX Spark pricing generally and irrelevant for local AI specifically — no announced MSI config reaches 128GB yet.

What about the Surface RTX Spark Dev Box instead? If the reported ~$5,999 price holds, you’d be paying Surface-laptop money for a desktop with the same ~300 GB/s ceiling — while the DGX Spark ($4,699) and GX10 (from $3,099) do the same job on the same silicon family for less. Its one distinctive feature is shipping Windows 11 Pro with WSL2 GPU passthrough and CUDA preconfigured. Wait for a confirmed price.

Sources

Last updated October 8, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.