Aleph Alpha Kolibri-1 for Local AI in 2026: The 78B Apache 2.0 MoE Your Tools Can't Load Yet
TL;DR: Aleph Alpha released Kolibri-1 on October 3, 2026 — a 78.1B-parameter MoE with only 3.46B active, Apache 2.0, German and English. Community GGUFs exist (Q4_K_M is 47.45 GB), but mainline llama.cpp, Ollama, and LM Studio can’t load it. Don’t buy hardware for this model; run it if you already own 64GB+.
| Kolibri-1 (patched llama.cpp) | Qwen3.6-35B-A3B (mainline) | Rent first | |
|---|---|---|---|
| Best for | German-language work, 64GB+ owners | Coding + everything else on 24GB | Testing before any purchase |
| Memory needed | 47.5GB (Q4_K_M) + context | ~22GB, fits one 24GB card | none |
| Price / Cost | $0 if you own the RAM | Used RTX 3090: $1,399–$1,450 | 3090s from $0.07/hr |
| The catch | Patched llama.cpp only; CUDA untested | Lower ceiling on German tasks | Setup time per session |
Honest take: Kolibri-1 is the most interesting European open-weight release of 2026, and almost nobody should buy hardware for it — it trails Qwen3.6-35B-A3B on coding benchmarks while needing twice the memory, and the local toolchain is held together by a community patch. Run it tonight if you already own the RAM; otherwise keep your money.
Aleph Alpha — the Heidelberg lab that has spent years positioning itself as Europe’s sovereign-AI answer — shipped open weights on October 3, 2026: Kolibri-1, a 78.1B-parameter mixture-of-experts model under Apache 2.0. The pitch is unusual for this site’s beat: a bilingual German-English reasoning model, validated out to a 1,048,576-token context, from a lab that explicitly isn’t American or Chinese.
The question that matters here is narrower: can the hardware in your office actually run it, and should you spend money to make that happen? Short answers: only with a patched build, and almost certainly not. The longer answers involve real measured numbers — 11.9–14.9 tok/s on a desktop CPU, 64 tok/s on a 48GB Mac — and a memory map you can check against your own machine with our VRAM calculator.
What Kolibri-1 actually is
The architecture, from the model card and Aleph Alpha’s tech report:
- 78.1B total parameters, 3.46B active per token. That’s a sparser ratio than almost anything else in the consumer-adjacent MoE class — 4.4% activation, versus ~8.6% for Qwen3.6-35B-A3B.
- 50 layers, 384 routed experts per layer, top-6 sigmoid routing plus one shared expert.
- Hybrid attention: sliding-window layers mixed with full-attention layers, which is how the long context stays affordable.
- Context: trained to 262,144 tokens, validated by Aleph Alpha up to 1,048,576.
- License: Apache 2.0 for the weights and configs. Training code and data recipes are not released.
- Languages: German and English, by design. This is the first serious Apache 2.0 model where German is a first-class target rather than an afterthought.
Why the 3.46B active figure matters: decode speed on every machine is memory-bandwidth-bound, and a MoE only reads its active parameters per token. At Q4_K_M (~0.61 bytes per weight), Kolibri-1 reads roughly 2.1 GB per token generated — about the same traffic as Qwen3.6-35B-A3B. The physics says this model can be fast on modest hardware. The problem is fitting it there in the first place, and getting any software to load it.
The catch: nothing you already use can load it
Kolibri-1’s architecture is registered as kolibri1, and as of October 10, 2026, that architecture is not in mainline llama.cpp, not in any Ollama release, and not in LM Studio. Download one of the community GGUFs and point stock llama.cpp at it, and you get the classic failure:
$ ./llama-cli -m Kolibri-1-Q4_K_M.gguf -p "Hallo, wer bist du?"
llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'kolibri1'
That’s the same error class we covered in our unknown-architecture fix guide, and the usual fix — update llama.cpp — doesn’t work here, because there’s nothing to update to yet. Three real paths exist today:
1. The Hob-forge patch (the practical one). The Hob-forge/Kolibri-1-GGUF repo ships quantizations from Q2_K to Q8_0 plus a kolibri1-llama.cpp.patch that applies against upstream commit 836d571, and a prebuilt release if you don’t want to compile. The converter’s author has validated CPU inference; CUDA, Vulkan, and Metal are explicitly listed as untested by that project. A separate Prompt48 Q4_K_M build (47.45 GB) was likewise validated CPU-only.
2. The CWBudde port (the rigorous one). An independent effort, kolibri-llama-cpp, keeps its changes as ordered patches against a pinned upstream commit, with tokenizer golden tests and a goal of bit-exact parity with Aleph Alpha’s vLLM reference. It already runs the full 50-layer graph and has produced the fastest local number anyone has published so far (more below). Numerical validation against the vLLM reference is still open — treat outputs as provisional.
3. vLLM with Aleph Alpha’s plugin (the official one). Aleph Alpha’s supported serving path is vLLM via their aleph-alpha-inference package; a native vLLM PR has been open since October 5. The launch command per their docs:
vllm serve Aleph-Alpha/Kolibri-1 --tensor-parallel-size 2 \
--kv-cache-dtype fp8 --reasoning-parser kolibri1 \
--tool-call-parser kolibri1 --enable-auto-tool-choice
The official minimum for the FP8 checkpoint (~78 GB of weights) is 2× A100 80GB, 2× H100, 1× H200, or one B200/B300. That’s datacenter hardware — the vLLM path is not a home-lab path, which is why the patched llama.cpp builds matter. If you want the vLLM background anyway, aifoss.dev’s vLLM review covers the server side.
There is no hosted Aleph Alpha API for Kolibri-1 either. Open weights or nothing — which is honest, but it means the patch situation above is the entire local story right now.
Memory map: what fits where
File sizes are from the Hob-forge and Prompt48 model cards and the CWBudde repo; what-fits verdicts assume you leave room for context (the KV cache is small per token thanks to the hybrid attention, but 64K+ contexts still want several GB).
| Quantization | File size | Runs on | Verdict |
|---|---|---|---|
| BF16 | 149 GiB (156 GB checkpoint) | 4× A100 80GB class | Datacenter only |
| Q8_0 | 79 GiB | 96GB RTX PRO 6000, 128GB unified memory | Fits, pointless for most |
| Q4_K_M | 47.45 GB | 64GB+ RAM (CPU), 64GB/128GB unified memory | The sane default |
| CWBudde mix (IQ3_XXS/IQ4_XS experts, Q8_0 rest) | 33.1 GiB | 48GB Macs, 48GB dual-GPU | Tightest working fit |
| Q2_K | ~26 GB | 32GB unified / 2×16GB | Exists; quality unverified |
Read that table against the cards people actually own and the shape of the problem is clear:
- A 24GB card — used RTX 3090 or RTX 4090 — cannot hold any usable quant. Even Q2_K plus context overflows 24GB. You’d be in CPU-offload territory, reading experts from system RAM, at which point the GPU is barely participating.
- 48GB is the entry ticket, and today that means a 48GB Mac (where the 33.1 GiB mix is proven to run) or dual 24GB cards (where nobody has published a verified CUDA run — remember, CUDA is listed untested). Our 48GB VRAM guide covers what else that tier buys you.
- 64GB of system RAM runs it CPU-only — genuinely, not theoretically, with measured numbers below. The sting: a DDR5 64GB kit costs $680–$1,070 as of September 2026 in the middle of the DRAM supply crisis, so “just add RAM” is no longer the cheap advice it was in 2025.
- 128GB unified-memory machines — GMKtec EVO-X2 ($3,499–$3,649), Framework Desktop ($3,449), Mac Studio M5 Max (from $2,499; $5,099 with 128GB) — hold Q4_K_M with room for the full trained context. That’s the comfortable tier, covered in our 128GB unified memory guide.
Measured speeds, and the ceiling check
Every number here is from a published run, with the bandwidth ceiling (bandwidth ÷ bytes-read-per-token) sanity-checked the way we always do:
$ python3 -c "
bw = {'DDR5 dual-ch desktop': 89.6, 'Strix Halo': 256, 'M5 Max': 614, 'RTX 3090': 936}
gb_tok = 2.1 # 3.46B active params at Q4_K_M, ~0.61 bytes/weight
for k, v in bw.items(): print(f'{k}: decode ceiling ~{v/gb_tok:.0f} tok/s')"
DDR5 dual-ch desktop: decode ceiling ~43 tok/s
Strix Halo: decode ceiling ~122 tok/s
M5 Max: decode ceiling ~292 tok/s
RTX 3090: decode ceiling ~446 tok/s
(The RTX 3090 row is the MoE punchline: the bandwidth is there for 400+ tok/s, but 24GB can’t hold the weights. Kolibri-1 is capacity-bound, not speed-bound, on every consumer GPU.)
| Hardware | Setup | Measured decode | Source |
|---|---|---|---|
| Ryzen 7 7800X3D, 64GB DDR5, no GPU | Q4_K_M, patched llama.cpp, 8 threads | 11.9–14.9 tok/s (83 tok/s prompt, 46.6GB RAM used) | Hob-forge card |
| 48GB unified-memory Mac | 33.1 GiB quant mix, Metal, 64K context | 64 tok/s | CWBudde port |
| Intel Arc Pro B70, Vulkan | Q4_K_M, partial offload | 22.4 tok/s (runs: 22.5 / 24.0 / 20.6) | intelinside PR #98 |
| 2× H100 (vendor, serving) | FP8, vLLM plugin | ~18 concurrent 256K-token requests | Aleph Alpha tech report |
All measured figures sit comfortably under their ceilings, so none of them trip our fake-benchmark alarm. The 64 tok/s Mac number is the one that should make 48GB-Mac owners sit up: that is a very usable speed for a 78B-class model, from a port that didn’t exist two weeks ago. It’s also provisional — that project itself says output parity with the vLLM reference hasn’t been confirmed yet.
One number circulating on social media fails the check: a claimed ~90 tok/s on a single RTX 3090 with a 4-bit quant. A 47.5 GB file does not fit in 24GB, so that run — if real — was mostly reading experts from system RAM, where the ceiling math above caps decode in the low 40s. We couldn’t trace the claim to a reproducible setup, so we’re not citing it as a data point. Treat it as noise.
The benchmark problem: Qwen already does this, smaller
Aleph Alpha’s own reported numbers, from their harness: SWE-bench Verified 66.4, MMLU-Pro 80. Respectable — and the model card’s own comparison shows Qwen3.6-35B-A3B at 73.8 on the same SWE-bench Verified run, with Qwen3.5-35B-A3B at 71.6. Kolibri-1 also trails Qwen3.6-35B-A3B on closed-book knowledge and multi-turn tool use in Aleph Alpha’s published tables. Credit where due: a vendor publishing a comparison its own model loses is rare, and it makes these numbers more trustworthy, not less.
But the implication for buyers is brutal. Qwen3.6-35B-A3B scores higher on coding, runs at 107 tok/s in stock Ollama on a used RTX 3090 (120+ on a 4090), fits in 24GB, and needs zero patches. Kolibri-1 needs twice the memory, a patched build, and loses the English benchmark race. If you’re assembling a local coding stack — say, Continue.dev against a local backend — Kolibri-1 is not the model that earns the hardware.
Where it genuinely has no local rival: German. Kolibri-1 was trained bilingual from scratch, and no Apache 2.0 model at any runnable size treats German as a primary language. If your prompts, documents, or users are German — a real consideration for the EU slice of the home-lab world — the Qwen comparison stops being the relevant one. The same goes if “trained outside the US and China, Apache 2.0, no API dependency” is itself the requirement, which for some European businesses it legally is.
What to actually buy
Prices as of October 2026, all verified in the sections above:
| Your situation | The machine | Price | Where |
|---|---|---|---|
| You want this class of MoE running tonight, in stock Ollama, with better coding scores | Used RTX 3090 24GB + Qwen3.6-35B-A3B | $1,399–$1,450 | Check price |
| You need Kolibri-1 itself (German work, sovereignty requirement) in one quiet box | GMKtec EVO-X2 128GB | $3,499–$3,649 | Check price |
| You already own a 48GB+ Mac or a 64GB-RAM desktop | Nothing — apply the patch and run | $0 | Hob-forge GGUF |
| You want to test it before spending anything | Rented GPUs, by the hour | 3090s from $0.07/hr (you’ll need two, or one 80GB card) | Vast.ai |
One honesty note on the EVO-X2 row: Kolibri-1 on Strix Halo means Vulkan or ROCm through the patched builds, and both are currently listed untested by the Hob-forge converter. The Arc B70 Vulkan run (22.4 tok/s, partial offload) suggests the Vulkan path works, but if you buy that machine today, you’re buying the proven 128GB-class workloads — gpt-oss-120b at ~31 tok/s, the big Qwen MoEs — with Kolibri-1 as an expected-soon bonus, not a guarantee. If none of these rows fit, start from the GPU buying guide instead of forcing this model into your budget.
FAQ
Will Ollama and LM Studio support Kolibri-1?
Both need the kolibri1 architecture to land in their bundled llama.cpp first, and mainline llama.cpp support is still at the community-patch stage as of October 10, 2026. History says weeks-to-months: popular architectures with this much attention (two independent ports, multiple GGUF uploads in the first week) tend to get merged. Nothing is announced.
Can I run it on a single RTX 4090 or RTX 5090? No quant that preserves usable quality fits in 24GB or 32GB. With experts offloaded to system RAM you’re capped around the low-40s tok/s by DDR5 bandwidth before overheads — and the CUDA path in the community patches is untested anyway. This model wants unified memory or 48GB+.
Is the 1M-token context real on local hardware? The weights were trained to 262K and validated by Aleph Alpha to 1M — on datacenter serving. Locally, the proven configuration so far is 64K context on a 48GB Mac with ~2 GiB to spare. Long-context KV cache is the next thing to eat your memory headroom after the weights; size it with the VRAM calculator.
Why does a 78B model decode faster than a 22B dense model? It only reads its 3.46B active parameters per token (~2.1 GB at Q4_K_M), while a dense 22B at Q4 reads ~13 GB. Decode speed tracks bytes-read-per-token, not total parameters. Capacity (fitting all 47.5 GB of weights) is what you pay for — once they fit, the speed comes almost free.
Should German-speaking users buy hardware for this? If you were already shopping in the 128GB-unified-memory tier, Kolibri-1 strengthens that case meaningfully — it’s the first Apache 2.0 model where German isn’t a compromise. If you weren’t, rent two 3090s (or one 80GB card) for an evening on Vast.ai and see whether the German quality difference matters for your documents before committing $3,500.
Recommended Gear
- Used RTX 3090 24GB — $1,399–$1,450 used; still the value play for the Qwen3.6-35B-A3B class that beats Kolibri-1 on coding
- GMKtec EVO-X2 128GB — $3,499–$3,649; the capacity box where Q4_K_M fits with full-context headroom
- Mac Studio M5 Max — from $2,499 ($5,099 at 128GB); 614 GB/s of unified bandwidth and the platform where the fastest Kolibri-1 run so far happened
Sources
- Kolibri-1 model card — Aleph Alpha / Hugging Face
- Kolibri: A Sovereign European Model on the Pareto Frontier (tech report) — Aleph Alpha
- Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE With 3.46B Active Parameters — MarkTechPost
- Kolibri 1: Aleph Alpha’s Sovereign Open-Weight LLM — DataCamp
- Kolibri-1-GGUF: quantizations, llama.cpp patch, and CPU benchmarks — Hob-forge / Hugging Face
- Kolibri-1-GGUF Q4_K_M (47.45 GB) — Prompt48 / Hugging Face
- kolibri-llama-cpp: independent llama.cpp port with tokenizer golden tests — CWBudde / GitHub
- Kolibri-1-78B-A3B Q4_K_M on Intel Arc Pro B70, llama.cpp Vulkan — labscommunity/intelinside PR #98
- Aleph Alpha Kolibri-1: EU Open-Weight Model, Benchmarks & Setup — Connic
- How to Run Kolibri Locally: GGUF, MLX and Hardware Picks — Developers Digest
- Kolibri-1: Aleph Alpha’s Open-Weight German-English MoE Model — MindStudio
- RAM price index: DDR5 64GB kits $680–$1,070 — Tom’s Hardware
Last updated October 10, 2026. Prices and specs change; verify current rates before purchasing. Kolibri-1 toolchain status (llama.cpp mainline support, Ollama/LM Studio availability) is moving fast — check the linked repos for the current state.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.