Surface Laptop Ultra for Local AI: RTX Spark's 128GB Debut, the Leaked Benchmarks, and Why You Shouldn't Preorder Today
TL;DR: Microsoft fully reveals the Surface Laptop Ultra today (October 7) — the first laptop with NVIDIA’s RTX Spark chip: up to 128GB unified memory at ~300 GB/s with a real CUDA stack. A leaked engineering sample ran Qwen3.5 9B at just 22.45 tok/s on Vulkan because the CUDA path failed outright. Capacity is the pitch; the drivers are the risk.
| Surface Laptop Ultra 128GB | GMKtec EVO-X2 128GB | MacBook Pro 16” M5 Max 128GB | |
|---|---|---|---|
| Best for | CUDA + 128GB in a laptop | Cheapest 128GB box today | Fastest dense-model laptop |
| Memory / bandwidth | 128GB / ~300 GB/s | 128GB / 256 GB/s | 128GB / 614 GB/s |
| Price | Unannounced (rumored ~$2,999 start, 128GB likely far higher) | $3,499–$3,649 | $6,999 |
| The catch | Prototype showed CUDA AI workloads failing | No CUDA, slower prompt processing | $6,999, no CUDA |
Honest take: Don’t preorder. The one thing this machine offers that nothing else does — 128GB plus CUDA in a laptop — is exactly the part that wasn’t working on the leaked prototype a few weeks ago. Wait for launch reviews showing llama.cpp running on CUDA, not Vulkan. If you need 128GB for local AI today, the GMKtec EVO-X2 and Mac Studio M5 Max already exist.
Microsoft holds its Windows and Surface event in San Francisco today, October 7, and the headline device is the Surface Laptop Ultra — the machine Microsoft teased alongside NVIDIA’s RTX Spark platform but never priced (Windows Central, GSMArena). NVIDIA says RTX Spark laptops and compact desktops start arriving this month, with over 30 laptop models due by year-end (PCWorld, NVIDIA blog).
If you run local models, this is the most interesting Windows laptop announcement in years, for one specific reason: it’s the first time 100GB+ of unified memory and the CUDA software stack ship in the same portable machine. It’s also a machine whose only public hands-on — a leaked engineering sample tested for a month — showed the AI stack in rough shape. Both halves matter before you give Microsoft your money. We covered the platform’s specs when it was announced in our RTX Spark overview and the Computex reality check; this is what’s changed now that a real product is hours from a price tag.
What Microsoft is actually shipping
The Surface Laptop Ultra is a 15-inch machine built around NVIDIA’s RTX Spark N1X — a GB10-family Arm SoC with a Blackwell GPU on the same package, connected over NVLink C2C, sharing one pool of LPDDR5X that either side can use (Tom’s Hardware). Two chip configurations exist (Windows Report, Windows Central):
| Config | CPU | GPU | Memory options | Bandwidth |
|---|---|---|---|---|
| N1X 18-core | 18 Arm cores | 5,120 CUDA cores | 24GB or 32GB | LPDDR5X, up to ~300 GB/s |
| N1X 20-core | 20 Arm cores | 6,144 CUDA cores | 32 / 48 / 64 / 128GB | LPDDR5X, up to ~300 GB/s |
NVIDIA rates the platform at 1 petaflop of sparse FP4 AI compute — the same class of number it quotes for the DGX Spark desktop, which uses the same GB10 silicon at 273 GB/s. The rest of the Surface package: a mini-LED “PixelSense Ultra” touchscreen rated at 2,000 nits peak HDR, plus HDMI, USB-C, USB-A, an SD slot, and a headphone jack (Thurrott).
Pricing is the open question today’s event is supposed to answer. Microsoft has said nothing official; retail listings peg the expected US start around $2,999, and Windows Central expects the entry 18-core/24GB config to land “well above $2,000.” A 128GB/20-core machine will cost meaningfully more — the comparable MacBook Pro line runs $3,000 to $7,000 as you climb the memory ladder. Treat every dollar figure for this laptop as unconfirmed until the event ends.
If you don’t need the Surface badge: Lenovo’s Yoga Pro 9n (15.3”) and Yoga 9n 2-in-1 (16”), ASUS’s ProArt P16 and P14, a Dell XPS 16 refresh, and HP OmniBook models are all launching on the same two N1X configs in the coming weeks.
The leaked prototype: real numbers, real problems
The only public hands-on data for this machine comes from an unusual source. In mid-June, a TechPowerUp forum member going by Fouquin picked up what turned out to be a Surface Laptop Ultra EV1.5 engineering sample — a pre-production validation unit with the top 20-core/6,144-core chip, 24GB of unified memory, and a 512GB SSD — and spent about a month testing it before publishing results in late July (TechPowerUp, Tom’s Hardware).
The LLM numbers, from llama-bench running Qwen3.5 9B on the Vulkan backend (the secondary coverage doesn’t state the quantization) (kocpc):
$ llama-bench -m qwen3.5-9b.gguf
| test | t/s |
| ------ | ---------: |
| pp512 | 1,543.16 |
| pp1024 | 1,582.68 |
| pp2048 | 1,741.19 |
| tg128 | 22.45 |
Two things stand out. First, prompt processing north of 1,500 tok/s on a 9B model is genuinely strong for a thin laptop — that’s the Blackwell compute doing its job, and it matters for long-context work like feeding a codebase into a local model. Second, 22.45 tok/s generation is low for the hardware. Run the bandwidth ceiling check we apply to every number on this site: a 9B model at Q4_K_M reads roughly 5.6GB of weights per token, so 300 GB/s puts the ceiling near 53 tok/s — the prototype hit about 42% of it (if the test used Q8, the ceiling is ~31 tok/s and 22.45 is a normal 72%). Healthy hardware with mature drivers lands at 70–95% of ceiling. Either the test ran a heavier quant, or the drivers were leaving a lot on the table.
The rest of the report points firmly at drivers. The CUDA path — the entire reason to buy NVIDIA silicon — failed outright on AI workloads: on the same machine, Vulkan was 54–78× faster at prompt processing and 5.3× faster at generation than the broken CUDA backend (TechTimes). Games stuttered, GPU clocks were unstable, the 616.00 pre-release driver was visibly immature (Igor’s Lab), and the Arm CPU sat at 100°C under sustained load (Hardware Busters). CPU performance itself was fine — Cinebench 2026 scores of 5,771 multicore / 540 single-core put it in 14-core Apple M3 Max territory (VideoCardz).
An engineering sample on pre-release drivers is not the shipping product, and NVIDIA has had months since. But “CUDA doesn’t work yet” a few weeks before launch is exactly the kind of thing you let reviewers confirm is fixed before paying $3,000+.
The 300 GB/s math hasn’t changed
Bandwidth decides generation speed, and ~300 GB/s is this platform’s hard ceiling — the same math we walked through when RTX Spark was announced:
- Dense 70B models fit, but crawl. Llama 3.3 70B at Q4_K_M is 42.5GB of weights; 300 ÷ 42.5 ≈ 7 tok/s theoretical ceiling, with real-world numbers below that. The DGX Spark — same GB10 family at 273 GB/s — delivers 2.7–6 tok/s on 70B in practice. The 128GB Surface will hold the model easily and generate at reading-aloud pace.
- MoE models are where the capacity pays off. A mixture-of-experts model only reads its active experts per token: the DGX Spark runs Qwen3 30B A3B at roughly 89 tok/s in llama.cpp, and 100B+ MoE models that read ~3GB per token have a ceiling near 100 tok/s on this bandwidth. 128GB of unified memory means gpt-oss-120b-class models fit entirely in memory with room for context — something no 24GB or 32GB discrete GPU can claim.
- A used desktop GPU still wins on raw speed. A $1,190–$1,400 used RTX 3090 has 936 GB/s — three times this platform’s bandwidth — and demolishes it on any model that fits in 24GB. Our used RTX 3090 guide is the tokens-per-dollar case; the Surface’s case is capacity and portability, not speed.
That’s also why the 24GB and 32GB configs deserve a hard pass for AI work. They pair the platform’s modest bandwidth with less capacity than a used desktop card that’s 3× faster. The entire argument for RTX Spark is holding models a discrete GPU can’t — if you’re not buying the 64GB or 128GB tier, buy something else.
The trap to avoid at launch (learned from DGX Spark)
The same silicon has been shipping in the DGX Spark desktop since winter, and its community found the failure mode you should expect Surface reviewers to trip over: run a 70B model at the default precision and performance looks broken. NVIDIA’s NIM container serves Llama 3.3 70B as FP8 by default — ~70GB of weight reads per token through a ~273–300 GB/s pipe — and users reported exactly the ~3 tok/s the math predicts (NVIDIA Developer Forums). The fix is the same on the Surface as on the desktop: use Q4_K_M GGUFs (roughly doubles decode speed), or better, run MoE models that only read active experts. If a launch review shows a single dense-70B-at-FP8 number and calls the machine slow, it measured the configuration, not the hardware.
The second thing to check before ordering: at least one review running llama.cpp or Ollama on the CUDA backend at sane fractions of the bandwidth ceiling. Vulkan-only results in October would mean the July prototype situation hasn’t been fixed — and without working CUDA you’ve bought a quieter, more expensive Strix Halo. NVIDIA’s IFA llama.cpp optimizations and the PAIR local-network router are built around this platform working as advertised, so NVIDIA has every incentive to land the drivers — but incentive isn’t evidence.
What to actually buy
Prices as of October 2026, all verified in the comparison above or in the linked articles:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| Need 128GB + CUDA in a laptop and can wait 2–4 weeks | Surface Laptop Ultra 128GB | TBA today | Wait for launch reviews confirming CUDA inference works |
| Want a 128GB local AI box today at the lowest price | GMKtec EVO-X2 128GB (Strix Halo) | $3,499–$3,649 | Check price |
| Want the same GB10 CUDA stack today, desktop is fine | ASUS Ascent GX10 | $3,099–$4,150 | Check price |
| Fastest laptop for dense models, budget stretches | MacBook Pro 16” M5 Max 128GB | $6,999 | Check price |
| Everything you run fits in 24GB — speed over capacity | Used RTX 3090 desktop | $1,190–$1,400 (card) | Check price |
| Undecided — test your workload on rented hardware first | Rented 3090, from $0.07/hr | pay per hour | Vast.ai |
The deeper comparisons behind those rows: DGX Spark vs Mac Studio M5 Max, Strix Halo vs DGX Spark, ASUS GX10 vs DGX Spark, and MacBook Pro vs Mac Studio at M5 Max. If a discrete GPU is back on the table after reading that table, start from the GPU buying guide — especially with the RTX 5090 gone from US retail.
One audience genuinely should care about this laptop: developers running local coding agents. Agentic workloads are long-prompt, modest-generation — exactly the shape this chip’s strong prefill and weak decode favor — and CUDA means the vLLM/fine-tuning ecosystem works where it never will on Apple Silicon or Strix Halo. If that’s you, the setup half of the story is on our sister site: running Continue.dev against a local Ollama backend.
FAQ
Is the Surface Laptop Ultra faster than a MacBook Pro M5 Max for local LLMs? For dense models, no — 614 GB/s vs ~300 GB/s means the Mac generates roughly twice as fast on anything that reads all its weights per token. For MoE models the gap narrows sharply, prompt processing favors the Blackwell GPU, and the Surface has CUDA. It’s a workload question, not a spec-sheet one.
How much will it cost? Microsoft announces pricing at today’s event. Retail leaks suggest around $2,999 to start, with the entry 18-core/24GB config expected “well above $2,000” and the 128GB tier likely several thousand more. Nothing is official until the event ends.
Can it run 70B models? The 64GB and 128GB configs hold a 70B Q4_K_M (42.5GB) comfortably. Expect single-digit tokens per second — the ~300 GB/s ceiling allows about 7 tok/s on dense 70B before overhead. MoE models in the 100B+ class are the better fit for this bandwidth.
Should I buy the 24GB config? For local AI, no. At 24GB you’re paying unified-memory prices for less capacity than a used RTX 3090 with 3× the bandwidth. This platform only makes sense at 64GB+.
How is this different from the DGX Spark? Same GB10 chip family. The DGX Spark ($4,699) is a desktop at 273 GB/s with NVIDIA’s curated DGX OS stack and ConnectX networking for clustering; the Surface is a consumer Windows laptop at ~300 GB/s. If you don’t need it to be a laptop, the desktop boxes are proven and discounted — the ASUS GX10 starts around $3,099.
Recommended Gear
Products linked in this article:
- GMKtec EVO-X2 128GB — cheapest 128GB unified-memory box shipping today
- ASUS Ascent GX10 — the same GB10 platform in a desktop, available now
- MacBook Pro 16” M5 Max 128GB — the dense-model speed king among 128GB laptops
- Used RTX 3090 — still the tokens-per-dollar benchmark everything here is measured against
Sources
- Microsoft Surface Laptop Ultra wields Nvidia’s RTX Spark superchip — Tom’s Hardware
- Microsoft Announces Surface Laptop Ultra With Nvidia RTX Spark Chip — Thurrott
- Surface Laptop Ultra Specs Detailed, Up to 128GB RAM and 20-Core RTX Spark SoC — Windows Report
- NVIDIA confirms RTX Spark configurations and availability — Windows Central
- What to expect at Microsoft’s October 7 Windows & Surface event — Windows Central
- Microsoft announces October 7 Windows and Surface event — GSMArena
- Nvidia’s RTX Spark PCs launch in October — PCWorld
- Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026 — NVIDIA Blog
- NVIDIA RTX Spark Laptop Found on Roadside Gets Reviewed Before Official Launch — TechPowerUp
- Mystery reviewer ‘finds’ Nvidia RTX Spark prototype laptop — Tom’s Hardware
- RTX Spark N1X laptop early prototype performance — kocpc
- Surface Laptop Ultra: Initial RTX Spark performance in prototype — Igor’s Lab
- Surface Laptop Ultra Prototype Leak Reveals RTX Spark CUDA Failures — TechTimes
- An Unreleased RTX Spark Laptop Got a Month of Testing — and the CPU Lives at 100°C — Hardware Busters
- Surface Laptop Ultra with NVIDIA N1X matches 14-core Apple M3 Max in Cinebench 2026 — VideoCardz
- Trouble with Llama 70B 3.3 FP8 at 3 tok/s on DGX Spark — NVIDIA Developer Forums
Last updated October 7, 2026. Prices and specs change; verify current rates before purchasing. Surface Laptop Ultra pricing was unannounced at publication time — Microsoft’s October 7 event may have since confirmed it.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.