VibeThinker-3B for Local AI in 2026: A 2GB Download That Claims 94.3 on AIME — and the VRAM Math That Actually Matters

vibethinkerlocal-llmgpusmall-modelsreasoning

TL;DR: VibeThinker-3B is a 1.93GB Q4_K_M download (MIT license) that Weibo’s AI team says scores 94.3 on AIME 2026 — the same ballpark as DeepSeek V3.2 at 671B parameters. The weights fit any GPU made this decade; the 60K–100K-token thinking budget is what actually sets your hardware floor. Almost nobody needs to buy anything to run it.

Any 8GB+ GPU you ownUsed RTX 3060 12GBRent first
Best forTrying it tonight, $0Cheapest full-context (131K) cardTesting before any purchase
Price / Cost$0~$260–$300 used (Oct 2026)Vast.ai 3090 from $0.07/hr
The catchFull 131K context is tight on 8GB360 GB/s caps you near ~130 tok/sSetup time for a 2GB model

Honest take: Don’t buy hardware for VibeThinker-3B — it runs on whatever you have, including a MacBook Air. Buy the used RTX 3060 12GB only if you have no GPU at all and want the cheapest card that holds the model plus its entire 131K context in VRAM. And treat the AIME number as a benchmark-specialist result, not proof it replaces your 27B daily driver.

VibeThinker-3B resurfaced on r/LocalLLaMA this week, months after its June 2026 release, because the pitch is irresistible in a year when a used RTX 3090 costs $1,150–$1,350: a 3.09B-parameter model, post-trained from Qwen2.5-Coder-3B by a nine-person team at Sina Weibo, claiming competition-math scores that match models 200× its size. The weights are real, MIT-licensed, and on Hugging Face (WeiboAI/VibeThinker-3B) with the technical report on arXiv (2606.16140).

This guide does the part the hype threads skip: what it actually takes to reproduce that reasoning on your desk — in VRAM, in tokens per second, and in minutes per answer. Run your own card through our VRAM calculator if you want to check a different model+context combination.

What Weibo actually released

The verified facts, as of October 3, 2026:

  • Architecture: 3.09B dense transformer (2.77B non-embedding), fine-tuned from Qwen2.5-Coder-3B — 36 layers, grouped-query attention with 16 query heads and 2 KV heads. No MoE tricks; every token reads every weight.
  • License: MIT. Commercial use, fine-tuning, redistribution all allowed.
  • Context: the model card configures 131,072 tokens, with output budgets up to ~102K tokens. Weibo recommends 60K–100K max tokens for hard problems.
  • Training method: a four-stage “Spectrum-to-Signal” pipeline — curriculum SFT, multi-domain RL, and offline self-distillation. The predecessor VibeThinker-1.5B put its entire post-training bill at $7,800; Weibo hasn’t disclosed the 3B’s cost.
  • Self-reported scores: 94.3 on AIME 2026 (97.1 with claim-level test-time scaling), 89.3 on HMMT25, 76.4 on IMO-AnswerBench, 80.2 Pass@1 on LiveCodeBench v6, and a 96.1% acceptance rate on unseen LeetCode weekly/biweekly contests from late April through late May 2026.
  • The asterisk: 70.2 on GPQA-Diamond, where frontier models clear 80. The model card itself says it was not trained for tool calling or agentic workflows.

That last pair of numbers is the paper’s actual thesis, and it’s more interesting than the headline: logical reasoning compresses into 3B parameters surprisingly well; broad world knowledge does not.

The VRAM math nobody does: weights are the cheap part

The Q4_K_M GGUF is 1.93GB (constructai/VibeThinker-3B-GGUF; community quants from mradermacher and prithivMLmods land between 1.8 and 2.0GB). Full BF16 weights are about 6GB. On weights alone, a 4GB card from 2016 qualifies.

But this is a long-thinking model. Its benchmark scores come from answers that run 60,000–100,000 tokens, and every one of those tokens sits in the KV cache. With 36 layers, 2 KV heads, and 128-dim heads, the cache costs about 36KB per token at FP16 — computed from the architecture, so here’s what context actually costs:

Context filledKV cache (FP16)Total VRAM at Q4_K_M*
16,384 tokens~0.6GB~3.0GB
32,768 tokens~1.2GB~3.6GB
65,536 tokens~2.4GB~4.8GB
131,072 tokens (max)~4.8GB~7.2GB

*1.93GB weights + KV cache + ~0.5GB compute buffers. Quantizing the KV cache to Q8_0 roughly halves the cache column.

So the real tiers are:

  • 4GB GPU: loads the model, but caps you around 16K–24K context — enough for chat, not for benchmark-grade reasoning chains.
  • 8GB GPU (RTX 4060, $297–$330 as of late September 2026): handles ~90K+ context, or the full 131K with Q8_0 KV quantization. This is the honest minimum for reproducing what the paper measured.
  • 12GB GPU (used RTX 3060, ~$260–$300 per ResalePrices’ October tracker, fair-ask range $279–$301 across 336 listings): full context at FP16 cache with room to spare. Cheapest no-compromise card.

If you already own anything on this list, the correct purchase is nothing. Our 8GB VRAM model guide covers what else runs at that tier.

How fast it runs: computed ceilings, not vendor numbers

We couldn’t find credible community tok/s benchmarks for this model yet, so we’re doing what we always do instead of guessing: decode speed on a dense model is bounded by memory bandwidth divided by bytes read per token. At Q4_K_M, every token reads the full 1.93GB. These are hard ceilings — real-world speeds land below them, and small models land further below than big ones because per-token launch overhead stops being negligible when the math is this light:

HardwareBandwidthCeiling (tok/s)Realistic expectation
RTX 4060 8GB288 GB/s149~90–130
RTX 3060 12GB360 GB/s186~110–160
RTX 5060 Ti 16GB448 GB/s232~140–200
RTX 3090 24GB936 GB/s485~250–400
M4 Pro Mac273 GB/s141~85–125
Dual-channel DDR4 CPU~50 GB/s26~15–22

Any site quoting more than the ceiling column for this model at Q4 is publishing a number that physics rejects.

The real cost is thinking time, not VRAM

Here’s what the benchmark setup means at your desk. A hard competition problem at the recommended 60K-token budget takes:

  • ~7–9 minutes on an RTX 3060 at ~120 tok/s
  • ~5 minutes on an RTX 5060 Ti at ~180 tok/s
  • ~3 minutes on an RTX 3090 at ~320 tok/s
  • 45+ minutes on a CPU at ~20 tok/s

And the paper’s scores are averages over 64 sampled attempts per math problem (8 for coding, 16 for knowledge), at temperature 1.0. Your single run is one draw from that distribution — sometimes it nails AIME problems, sometimes it doesn’t. A 3B model being this cheap to run is exactly why that evaluation style works; just don’t expect the 94.3 experience on every prompt.

The upside: at 1.93GB, this model loads from NVMe in about two seconds and you can keep it resident permanently next to a bigger model. It costs almost nothing to have around.

Running it tonight

Any llama.cpp-family stack works. Pulling a community GGUF straight from Hugging Face:

ollama run hf.co/mradermacher/VibeThinker-3B-GGUF:Q4_K_M
>>> Find all primes p such that p^2 + 8 is also prime.
<think>
We need p^2 + 8 prime. Check p = 3: 9 + 8 = 17, prime. For p ≠ 3,
p^2 ≡ 1 (mod 3), so p^2 + 8 ≡ 0 (mod 3)...
</think>
The only such prime is p = 3.

llama.cpp directly:

llama-cli -hf mradermacher/VibeThinker-3B-GGUF:Q4_K_M \
  -c 65536 --temp 1.0 --top-p 0.95 -n 60000

Two settings matter more than usual, and both come from the model card:

Keep temperature at 1.0. This model was RL-trained to explore diverse reasoning paths at high temperature; Weibo explicitly warns that lowering it degrades reasoning quality. The usual local-LLM habit of dropping to 0.2–0.6 for “reliability” actively hurts here.

The problem you will actually hit: the default context window truncates its brain. Ollama’s default num_ctx is 4,096 tokens. VibeThinker regularly thinks for 20,000+ tokens before answering, so out of the box the window fills mid-thought, earlier reasoning slides out of context, and output degrades into repetition loops or an answer that ignores its own work. The fix is one line:

/set parameter num_ctx 65536

or OLLAMA_CONTEXT_LENGTH=65536 on the server. Per the table above, 64K costs ~2.4GB of KV cache — cheap on anything 8GB+. If you’re on 6GB, set 32K and add --cache-type-k q8_0 --cache-type-v q8_0 in llama.cpp. This same truncation trap applies to every long-thinking model; we covered its cousin in the Ollama reloading fix guide.

For coding use, it speaks the standard OpenAI-compatible API, so it plugs into Cline or Cursor as a local backend the same way our sister site covers for other local models — though read the next section before making it your daily coding model. For the Ollama/self-hosting deep dive, aifoss.dev has a dedicated setup guide.

The benchmark argument, honestly

VentureBeat’s headline called it “the AI world arguing over benchmarks again,” and the skeptics’ case is specific: AIME 2026 and HMMT25 problems were public before the model’s release, and the scores are self-reported with no independent harness — comparison numbers come from each vendor’s own reports. If competition problems leaked into training data, the headline collapses. Weibo’s counter is a three-tier decontamination process: N-gram filtering against benchmarks, LLM-based query filtering, and verifier-checked answers. Sebastian Raschka’s post-training notes walk through the pipeline without resolving the contamination question — because from outside, nobody can.

What you can verify at home is the part that matters for actual use:

  • LiveCodeBench v6 at 80.2 is the more meaningful score for most readers — it tests executable code on problems published after training, which is harder to contaminate. The 96.1% LeetCode acceptance on contests from April–May 2026 points the same direction.
  • GPQA-Diamond at 70.2 and the explicit “not trained for tool calling” warning mean this is not a general assistant. Ask it about PCIe bifurcation or your homelab’s Docker setup and you’ll feel the missing 600B parameters immediately.
  • It’s a specialist: verifiable math, competition-style coding, STEM problem-solving. For an everyday local coding model, the 2026 coding LLM rankings still apply.

Our verdict on the controversy: the contamination question is unresolvable from outside, but it also doesn’t change the practical calculus. The model costs 2GB of disk and nothing else to evaluate on your problems — the cheapest benchmark audit you’ll ever run.

What to actually buy

Prices as of October 2026, all verified above:

Your situationThe movePriceWhere
Any 8GB+ GPU, or a 16GB+ MacBuy nothing — pull the GGUF$0—
No GPU; want this + the whole 3B–13B tierUsed RTX 3060 12GB~$260–$300Check price
Want fast 100K-token chains + headroom for 27B modelsRTX 5060 Ti 16GB$679–$805 (MSRP $429)Check price
Undecided — test the workload firstRented 3090, billed hourlyfrom $0.07/hrVast.ai

The honest framing: VibeThinker-3B is an argument against spending money. If your workload is verifiable reasoning, a $260 used card now runs what required a $1,200 card two years ago. The 5060 Ti row is there for people who want this model’s speed ceiling and the 12GB-to-16GB tier jump for larger models — don’t buy it for a 2GB model alone.

FAQ

Does VibeThinker-3B really match DeepSeek V3.2? On AIME 2026 specifically, the self-reported 94.3 is in the same range. On general knowledge (GPQA-Diamond 70.2) and anything agentic, no — and Weibo’s own paper says so. It’s a reasoning specialist, not a frontier-model replacement.

How much VRAM does it need? ~1.93GB for Q4_K_M weights. Budget ~3.6GB total at 32K context, ~4.8GB at 64K, ~7.2GB at the full 131K (FP16 KV cache). An 8GB card covers realistic use; 12GB covers everything.

Can it run on CPU only? Yes — it’s the rare reasoning model where CPU inference is tolerable (~15–22 tok/s on dual-channel DDR4). But 60K-token answers take 45+ minutes, so interactive use wants any GPU.

Why does it ramble or loop forever on my machine? Almost always the context window. Raise num_ctx to 65536 (Ollama defaults to 4,096, which truncates its thinking chain) and keep temperature at 1.0 — both lower temperature and small windows degrade this model specifically.

Is it good for coding in Cursor or Cline? It scores 80.2 on LiveCodeBench v6, but it wasn’t trained for tool calling or agentic workflows, so as an agent backend it will disappoint. Use it for algorithm-style problems; use a dedicated local coding model for agent work.

MIT license — can I ship it in a product? Yes. Commercial use, modification, and redistribution are allowed; you keep the copyright notice.

Sources

Last updated October 3, 2026. Prices and specs change; verify current rates before purchasing.

Products linked in this guide:

Was this article helpful?

Get the numbers before you buy

New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.