MicroLLM Lab and Browser LLMs in 2026: What WebGPU AI With No Install Actually Delivers

webgpubrowser-llmlocal-llmwebllmgpu

TL;DR: MicroLLM Lab, which hit the Hacker News front page in late September 2026, runs seven tiny LLMs (26M–362M parameters, 15–216 MB downloads) entirely in a browser tab via WebGPU — zero install, zero server, and genuinely fast (~115 tok/s on an Apple M4). But micro-models are demos and autocomplete engines, not chatbots; for real models in a browser you want WebLLM, and for real work you still want Ollama on a real GPU.

MicroLLM Lab (in-tab)WebLLM (in-tab)Ollama + dedicated GPU
Best forTrying WebGPU inference in 30 seconds, demos, autocomplete researchReal 1B–8B chat models with zero installDaily-driver local AI: 7B–70B, RAG, agents, image gen
Model ceiling362M params (216 MB Q4)~8B–13B Q4, limited by browser GPU budgetWhatever your VRAM holds — 24GB runs 27B+
Speed~115 tok/s (PetitGPT 124M, Apple M4)~41 tok/s (Llama 3.1 8B Q4, M3 Max) ≈ 80% of native~95 tok/s (7B Q4, used RTX 3090)
The catchMicro-models ramble; 2048-token context capMulti-GB download per browser profile; tab owns your GPUCosts money: $260+ used, $300+ new

Honest take: Browser LLMs graduated from party trick to legitimate entry point in 2026 — but MicroLLM Lab is a physics classroom, not a workshop. Play with it to understand what a 124M model is, then run real models: WebLLM if you can’t install anything, a used RTX 3060 12GB (~$260–$296) if you can.

MicroLLM Lab (“Try 7 tiny LLMs in the browser”) is exactly what it says: open a web page, click Load, and a quantized language model streams into your browser’s storage and starts generating tokens on your GPU — through WebGPU, with no Python, no Ollama, no Docker, and no data leaving your machine. It’s a fun 30-second demo. It is also a useful lens on the bigger 2026 story: every major browser now ships WebGPU, which quietly turned “local AI” into something you can hand to a colleague as a URL.

Here’s what the lab actually contains, what the numbers say, and — because this is runaihome — what it changes (and doesn’t) about the hardware you should buy.

What MicroLLM Lab actually is

First, a correction to the framing you may have seen in social threads: these are not “7 LLMs” in the sense of the models you run in Ollama. They are micro language models, 26M to 362M parameters — between 20x and 300x smaller than a 7B model. The project’s GitHub repo lists the lineup:

ModelParametersQ4 downloadLicense
MiniMind2 Small26M15 MBApache-2.0
MiniMind2104M62 MBApache-2.0
PetitGPT research-v1124.6M74 MBApache-2.0
GPT-2 (baseline)124M77 MBMIT
SmolLM 135M Instruct134.5M80 MBApache-2.0
SmolLM2 135M Instruct134.5M80 MBApache-2.0
L20-Edu 135M134.5M80 MBApache-2.0
SmolLM2 360M Instruct362M216 MBApache-2.0

The engineering is the interesting part. Weights are quantized to 4-bit (group-32), cutting memory roughly 75% versus 16-bit — which is how a 124M-parameter model fits in a 74 MB download. The lab caches weights in the browser’s IndexedDB, so the second load is instant. Custom WebGPU decoders (Metal-oriented and fused-NVIDIA paths) cover Llama-style and GPT-2 architectures, and context is capped at 2,048 tokens to stay inside the browser’s GPU memory budget. The lab itself is Apache-2.0, a WebGPU port of the PetitGPT project — a 124.6M Llama-style model that was trained on a single RTX 4090, which is a nice data point for anyone who thinks training only happens on clusters.

Speed, measured by the project on an Apple M4 with greedy decoding: PetitGPT generates around 115 tokens/second, and SmolLM2 135M around 66 tokens/second — in a browser tab. For scale, comfortable human reading speed is roughly 7–10 tok/s. Micro-models in a tab are not slow.

What a 124M model is actually good for

Speed was never the question at this size — quality is. A 124M–362M model completes sentences plausibly, classifies short text, does simple structured extraction, and powers instant autocomplete (sub-10ms time-to-first-token is the pitch for this class of model as an edge layer). What it does not do is hold a coherent multi-turn conversation, follow non-trivial instructions, or write working code. Ask SmolLM2 135M a factual question and you’ll get fluent, confidently wrong prose more often than not.

That’s not a flaw in the lab — it’s the point of it. Loading seven models side by side and watching a 26M model versus a 362M model on the same prompt teaches you more about what parameter count buys than any benchmark table. Treat it as an interactive textbook.

The practical uses that survive contact with reality:

  • Demos you can share as a link. A model embedded in a web page needs no install and sends nothing to a server — the whole inference runs client-side.
  • Latency-critical micro-tasks. Autocomplete, tagging, intent detection, where a tiny model’s instant response beats a big model’s round trip.
  • Understanding quantization and scale. The Q4 group-32 compression on display here is the same technique your Ollama GGUF files use — see our quantization quality-loss guide for what it costs at real model sizes.

Why this works now: WebGPU shipped everywhere

Browser LLMs weren’t blocked on model quality — they were blocked on GPU access. That’s over. Per web.dev’s announcement, WebGPU now ships in every major browser: Chrome and Edge have had it since version 113 (Windows via Direct3D 12, macOS, ChromeOS) with Android support since Chrome 121; Firefox shipped it on Windows in version 141; and Safari 26 brought it to macOS Tahoe, iOS 26, and iPadOS 26.

You can check your own browser in five seconds. Open DevTools (F12) and run:

await navigator.gpu.requestAdapter()
// Expected on a working setup:
// GPUAdapter { features: GPUSupportedFeatures, limits: GPUSupportedLimits, ... }
// If you get null or "navigator.gpu is undefined", WebGPU is off or unsupported.

A problem you’ll actually hit: if that returns null on Linux, WebGPU is still behind flags in some Chrome builds — enable chrome://flags/#enable-unsafe-webgpu and ensure Vulkan drivers are installed. And if you clone MicroLLM Lab to run it locally, opening index.html directly fails — WebGPU requires an HTTP origin, not file://, so serve it with python3 -m http.server first. On work laptops with locked-down driver policies, requestAdapter() returning null usually means GPU access is administratively disabled, and no flag will fix it.

The serious version: WebLLM runs real models in a tab

MicroLLM Lab is the demo tier. The production tier of browser inference is WebLLM, the MLC project’s in-browser engine, and its numbers are the ones that should update your mental model: per the WebLLM paper, Llama 3.1 8B at 4-bit quantization generates about 41 tokens/second in the browser on an Apple M3 Max — roughly 80% of what the same model does natively via MLC-LLM on the same machine. The prebuilt catalog spans over 160 model builds, from SmolLM2 360M up through the Llama family and 13B-class models.

Forty-one tokens/second is four times reading speed, running in a tab, from a shared link. The 20% tax versus native inference is real but no longer disqualifying.

What still separates a WebLLM tab from your Ollama box:

  1. The GPU memory budget. A browser tab doesn’t get your whole card. An 8B model at Q4 needs roughly 5–6 GB of GPU memory, which rules out most integrated GPUs and 4–6GB laptop cards for anything past the 1B–3B class. Micro and small models run anywhere; 8B needs a machine that could run Ollama anyway.
  2. Downloads per browser profile. Multi-GB weights cached in IndexedDB, per browser, per machine — and the cache can be evicted.
  3. No stack. No persistent server API for other apps, no RAG pipeline, no model routing, no agents watching your filesystem. A tab is a sandbox; that’s its privacy feature and its ceiling.
  4. Context limits. Browser engines cap context well below what the same model handles natively — KV cache competes for the same constrained GPU budget. Check your model-plus-context fit with our VRAM calculator.

Use a browser LLM if / use Ollama if

A browser LLM is the right call when:

  • You can’t install software — locked-down work laptop, borrowed machine, Chromebook.
  • You’re sharing a demo with someone who will never touch a terminal.
  • You want to sanity-check a small model’s behavior before committing to a local setup.
  • Your only GPU is integrated. WebGPU runs on Intel Arc iGPUs, AMD APUs, and Apple base chips — hardware that had no practical local-AI story before. A 1B–3B Q4 model fits comfortably in most iGPU memory budgets; see what’s realistic at small VRAM tiers.

Ollama and a real GPU remain non-negotiable when:

  • You want daily-driver quality: 7B is entry level, and the models worth building a workflow on live in the 12–27B range — see best models for 8GB VRAM and up the tiers from there.
  • You need an API other tools can call — a local coding stack in Cline or Continue.dev points at an Ollama endpoint, not a browser tab.
  • You want RAG over your documents, agents, long contexts, or image generation. None of that exists in the browser-only world yet.
  • You care about the last 20% of speed and the first 100% of reliability. (Ollama itself is free — the cost is the hardware.)

Does WebGPU change what GPU to buy? Slightly — at the bottom.

The honest hardware takeaway: WebGPU doesn’t change the top of the market at all, and it softens the bottom. If you were about to spend $200–$300 on an entry card only to try local AI, a browser tab now gives you the trial for free on the iGPU you already own. The upgrade trigger is unchanged, though: the moment you want a model that’s actually good (7B+ at speed, or 12B+ at all), you need dedicated VRAM.

At September 2026 street prices, the entry rungs look like this: a used RTX 3060 12GB runs $260–$296 and its 12GB of VRAM at 360 GB/s makes it the cheapest card that treats 7B–13B models as routine — our full verdict on the used 3060 still stands. A new RTX 4060 costs about the same ($297–$330) with a warranty but only 8GB — fine for 7B Q4 in-browser and in Ollama, cramped beyond it. The next meaningful jump is the RTX 5060 Ti 16GB at $679–$805 (MSRP $429 — nobody pays MSRP in this market), which opens the 14B–27B class. Full tier math lives in the GPU buying guide.

What to actually buy

Prices as of September 2026, all verified above:

Your situationThe movePriceWhere
Just curious — want to try local AI todayNothing. MicroLLM Lab / WebLLM on your current machine$0Your browser
iGPU laptop, ready for real 7B–13B modelsUsed RTX 3060 12GB$260–$296Check price
Same budget, need a warranty, 7B is enoughRTX 4060 8GB (new)$297–$330Check price
Want the 14B–27B class without going usedRTX 5060 Ti 16GB$679–$805Check price
Undecided — want to test big models before buyingRented RTX 3090, from $0.07/hrpay per hourVast.ai

FAQ

Does MicroLLM Lab send my prompts to a server? No. Weights download once (15–216 MB), cache in IndexedDB, and inference runs entirely on your GPU via WebGPU. Nothing you type leaves the machine — that’s the architecture, not a policy promise.

Can my integrated GPU really run this? If your browser passes the navigator.gpu.requestAdapter() check, yes for the micro-models — they need well under 300 MB of GPU memory. For WebLLM’s 1B–3B models, most recent iGPUs work; 8B-class models generally need a dedicated card’s memory budget.

Why does the same model run ~20% slower in a browser than in Ollama? WebGPU adds an abstraction layer over Metal/Vulkan/D3D12, and browser security constraints limit the kernel tricks native runtimes use. The WebLLM paper measured ~80% of native throughput on identical hardware — a real tax, but decode speed is bandwidth-bound either way, so the ranking of your hardware doesn’t change.

Is a 360M model useful for anything real? Autocomplete, classification, short extraction — tasks where instant latency beats quality. For chat, code, or anything factual, no: that starts around 7B, which is exactly where browser convenience ends and dedicated VRAM begins.

Should I wait to buy a GPU since browsers keep improving? Browser engines will keep closing the software gap, but they can’t add memory to your machine — model size is a hardware wall. If your workload fits 8GB–12GB today, waiting buys you nothing; GPU prices have been rising, not falling, through 2026.

Sources

Last updated September 30, 2026. Prices and specs change; verify current rates before purchasing.

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.