Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a focus on German and English. The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency.
| Model | Kolibri 1 |
|---|---|
| Model Provider | Aleph Alpha GmbH |
| Model Developer | Aleph Alpha Research GmbH |
| Architecture | Mixture-of-Experts |
| Total parameters | 78B (78,103,074,560) |
| Active parameters / token | 3.46B (3,457,573,120) |
| Languages | German, English |
| Context length | 1,048,576 tokens; we recommend β€262,144 tokens for serving efficiency and complex tasks |
| Precision | float8_e4m3fn weights in 128Γ128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in bfloat16 |
| Reasoning mode | Yes |
| Tool calling | Yes |
| License | Apache 2.0 |
| Knowledge cutoff | EN: June 18, 2026, DE: June 18, 2026 This only affects implicit knowledge, the model may use more recent information through tool use. |
| Hardware requirements | Model memory footprint: ~78 GB (FP8 weights). Minimum: 2Γ A100 80 GB, 2Γ H100 SXM5, 1Γ H200, 1Γ B200 or 1Γ B300. Recommended: 2Γ H100 SXM5, 2Γ H200, 1Γ B200 or 1Γ B300. |
| Best for | Multi-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant |
| Release Date | 3rd of October 2026 |
| Code of Practice | Aleph Alpha is a signatory of the EU GPAI Code of Practice, see https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai. |
| Training Data | Pre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension. Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases. |
| Training method | We trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed. |
| Computing Resources | Pre-training (based on actual measurements), excluding mid-training and long-context: Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes); Parallelism: EP8 FSDP16 DP6; Time: 21 days (511h, 392k GPUh) Mid-training: 5 days, 90k GPUh (same setup as above) Long-Context: 13h, 10k GPUh (same setup as above except parallelism: FSDP128 DP6) FLOPS: 6.4e23 |
| Tech Report | https://aleph-alpha.com/downloads/tech-report.pdf |
Kolibri is intended to process text input and output in German and English and to perform a wide range of tasks beyond natural-language generation, for example multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and agentic tool calling.
Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length. Because positional encoding is applied only in the sliding-window layers, the context can be extended beyond that length without any position scaling, in principle to arbitrary lengths. We have validated quality and serving efficiency up to 1,048,576 tokens. For latency- or throughput-sensitive deployments and for complex tasks, we recommend contexts of at most 262,144 tokens. See the technical report for details.
For the kinds of systems Kolibri is meant to be integrated into, see AI system types below; for uses that we encourage users to refrain from, see Responsible Use.
Kolibri is designed to deliver strong German and English performance at low serving cost. Its mixture-of-experts architecture activates only a small fraction of its parameters for each token, keeping compute per token low while retaining the capacity of a much larger model. The trade-off is memory: the full model must be held in memory even though only part of it is active at any time. To keep long contexts affordable, most attention layers focus on nearby text, while a smaller number attend across the whole context. We also developed a tokenizer tailored to German word structure, so German text is processed efficiently without sacrificing English. Supporting two languages rather than many is a deliberate choice of depth over breadth.
Kolibri is intended for integration into conversational assistants and agentic workflows in German and English, in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed. It suits document-processing and drafting systems, question-answering systems over an organisation's own material, and internal knowledge and research tools. Its tool-calling and structured-output capabilities make it appropriate for orchestration layers that call APIs, execute code or run searches, provided the calling system validates the results. In decision-support systems it belongs on the advisory side, surfacing evidence and drafting options for a human to weigh, and it is not intended as the deciding component. More broadly, it is built for human-AI collaboration rather than unsupervised operation.
9.5Γ10Β² MWh (estimated), including node power and data-centre overhead (PUE). This includes pre-training, mid-training and long-context training. It excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models.
For pre-training, we obtained the average power usage per node and the power usage effectiveness (PUE) from the cluster provider. We multiplied these with the number of nodes and the runtime to obtain a total energy estimate.
Kolibri requires the aleph-alpha-inference package that provides the
Kolibri vLLM plugin. You can either use the provided container image
ghcr.io/aleph-alpha/aleph-alpha-inference, or install the package from
Aleph-Alpha/aleph-alpha-inference,
which also installs the vLLM version it supports:
pip install 'aleph-alpha-inference>=1'
Serve the model with reasoning and tool-calling enabled:
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.
The recommended sampling parameters for the model are temperature=1.0,
top_p=0.97 and top_k=128.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Aleph-Alpha/Kolibri-1",
messages=[
{
"role": "user",
"content": "ErklΓ€re kurz, was ein Mixture-of-Experts-Modell ist.",
},
],
extra_body={
"chat_template_kwargs": {
"reasoning_effort": "high",
"enable_thinking": True,
}
},
)
print(response.choices[0].message.content)
Kolibri supports explicit thinking effort levels that need to be configured
through the chat template. You can pass reasoning_effort values low,
medium and high to configure the amount of effort our model puts into
finding the answer. You can disable thinking altogether by setting
reasoning_effort to none or passing enable_thinking=false. The model
will then immediately respond.
The serving command enables Hermes-style tool calling (--tool-call-parser kolibri1 --enable-auto-tool-choice). Pass your function schemas via the
standard tools field of the chat completions request; the model will emit
tool calls that the parser converts into structured output. Tool calling can
be combined with reasoning mode.
import json
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
MODEL = "Aleph-Alpha/Kolibri-1"
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, e.g. Heidelberg",
},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["city"],
},
},
}
]
def get_weather(city, unit="celsius"):
return {"city": city, "temperature": 18, "unit": unit, "conditions": "cloudy"}
messages = [{"role": "user", "content": "Wie ist das Wetter gerade in Heidelberg?"}]
first = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
msg = first.choices[0].message
messages.append(msg)
for call in msg.tool_calls or []:
args = json.loads(call.function.arguments)
result = get_weather(**args)
messages.append(
{
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
}
)
final = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
print(final.choices[0].message.content)
The best value per row is bolded. In each table, the best value of each group with more than one model is underlined, and both marks compare the MoE models only. The dense models activate several times as many parameters per token, so they are greyed out and unmarked.
| Type | MoE | Dense | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Active parameters | 3B | 4-6B | 12B | 27B | 70B | |||||||||
| Ours | Baseline models | |||||||||||||
| Eval | Kolibri | Kolibri Origin | GLM-4.7 Flash 30B-A3B | Nemotron 3 Nano 30B-A3B | Qwen3.5 35B-A3B | Qwen3.6 35B-A3B | Qwen3-Next 80B-A3B Thinking | Gemma 4 26B-A4B IT | GPT-OSS 120B | Mistral Small 4 119B-A6B | GLM-4.5 Air 106B-A12B | Nemotron 3 Super 120B-A12B | Qwen3.8 27B | Apertus 70B Instruct |
| Overall (EN) | 75.5 | 54.1 | 64.7 | 65.6 | 74.7 | 71.4 | 62.4 | 71.9 | 72.3 | 63.1 | 64.4 | 73.0 | 80.2 | β |
| Overall (DE) | 70.8 | 46.4 | 50.4 | 59.3 | 69.8 | 67.3 | 58.0 | 66.3 | 70.2 | 61.4 | 64.8 | 67.9 | 79.9 | β |
| Knowledge | ||||||||||||||
| Average (EN) | 50.1 | 39.7 | 45.5 | 46.0 | 52.7 | 52.1 | 48.4 | 51.4 | 50.0 | 47.5 | 45.7 | 52.0 | 56.8 | 22.8 |
| Average (DE) | 57.6 | 44.9 | 46.5 | 41.2 | 61.3 | 61.0 | 55.7 | 61.5 | 58.0 | 51.4 | 52.2 | 59.5 | 69.2 | 24.8 |
| GPQA Diamond (EN) | 84.3 | 68.1 | 73.1 | 73.9 | 83.8 | 83.4 | 76.1 | 81.1 | 76.4 | 74.7 | 73.2 | 78.0 | 89.2 | 29.5 |
| GPQA Diamond (DE) | 81.3 | 58.5 | 59.8 | 49.6 | 84.2 | 80.6 | 72.2 | 80.1 | 76.0 | 72.9 | 71.1 | 76.6 | 88.1 | 31.4 |
| Humanity's Last Exam (EN) | 21.5 | 9.4 | 15.4 | 12.1 | 20.4 | 21.1 | 11.6 | 19.2 | 19.4 | 9.7 | 8.7 | 20.6 | 35.6 | 5.2 |
| Humanity's Last Exam (DE) | 15.9 | 10.4 | 9.1 | 13.1 | 18.1 | 20.5 | 15.6 | 23.4 | 20.7 | 10.5 | 10.5 | 22.3 | 37.2 | 5.7 |
| AA-Omniscience Accuracy (public set) | 14.8 | 11.3 | 17.0 | 19.5 | 22.2 | 21.0 | 24.2 | 20.7 | 23.3 | 25.0 | 20.0 | 26.7 | 19.0 | 13.5 |
| AA-Omniscience Index (public set) | -32.8 | -64.2 | -62.8 | -45.7 | -46.2 | -12.5 | -42.3 | -47.3 | -35.2 | -24.0 | -28.8 | -36.5 | -5.8 | β |
| MMLU-Pro CoT (EN) | 80.0 | 70.1 | 76.5 | 78.3 | 84.6 | 84.3 | 81.7 | 84.5 | 80.8 | 80.4 | 80.9 | 82.7 | 85.0 | 43.0 |
| MMLU-ProX CoT (DE) | 75.5 | 65.7 | 70.7 | 61.0 | 81.7 | 81.9 | 79.4 | 81.1 | 77.2 | 70.7 | 74.9 | 79.7 | 82.4 | 37.3 |
| Math | ||||||||||||||
| Average (EN) | 96.5 | 81.7 | 88.8 | 88.8 | 90.1 | 87.8 | 86.3 | 87.4 | 90.7 | 81.4 | 82.8 | 91.1 | 97.8 | 0.6 |
| Average (DE) | 88.8 | 74.3 | 45.2 | 84.3 | 79.6 | 83.7 | 83.7 | 88.1 | 90.8 | 75.4 | 80.6 | 86.5 | 96.7 | 0.1 |
| AIME 2025 (EN) | 96.9 | 81.9 | 89.4 | 89.6 | 88.1 | 84.6 | 84.2 | 87.3 | 90.6 | 79.8 | 81.9 | 91.7 | 97.9 | 0.6 |
| AIME 2025 (DE) | 87.5 | 73.5 | 43.8 | 84.4 | 76.7 | 82.9 | 80.6 | 88.1 | 90.6 | 72.3 | 80.6 | 85.6 | 96.5 | 0.2 |
| AIME 2026 (EN) | 96.0 | 81.5 | 88.3 | 87.9 | 92.1 | 91.0 | 88.5 | 87.5 | 90.8 | 83.1 | 83.8 | 90.4 | 97.7 | 0.6 |
| AIME 2026 (DE) | 90.0 | 75.2 | 46.7 | 84.2 | 82.5 | 84.4 | 86.7 | 88.1 | 91.0 | 78.5 | 80.6 | 87.5 | 96.9 | 0.0 |
| Agentic | ||||||||||||||
| Average (EN) | 63.4 | 41.6 | 58.9 | 46.4 | 63.4 | 62.1 | 46.3 | 54.6 | 54.0 | 40.7 | 53.5 | 54.9 | 66.7 | β |
| TerminalBench 2.1 | 27.7 | β | 20.2 | 9.7 | 39.7 | β | 8.6 | β | 29.2 | 21.0 | β | 39.7 | 76.8 | β |
| Tau2-Bench (Telecom) | 94.7 | 67.5 | 95.9 | 45.9 | 97.7 | 99.1 | 43.9 | 45.3 | 73.1 | 41.5 | 53.8 | 68.1 | 82.5 | 10.8 |
| Tau2-Bench (Retail) | 69.9 | 58.5 | 57.9 | 64.9 | 70.8 | 71.6 | 60.8 | 71.3 | 60.5 | 62.9 | 61.4 | 67.5 | 68.7 | 9.6 |
| Tau2-Bench (Airline) | 76.7 | 58.7 | 68.7 | 52.7 | 76.0 | 70.7 | 65.3 | 73.3 | 72.7 | 40.0 | 70.7 | 72.7 | 83.3 | 40.0 |
| Tau3-Bench (Banking) | 38.1 | 5.7 | 7.2 | 5.7 | 11.3 | 10.6 | 5.4 | 16.0 | 14.7 | 5.7 | 6.4 | 15.5 | 50.0 | 2.1 |
| BFCL v3 (multi-turn) | 39.8 | 22.8 | 58.2 | 47.9 | 54.0 | 53.5 | 51.4 | 53.4 | 45.6 | 36.2 | 61.6 | 44.6 | 42.5 | 0.6 |
| BFCL v4 (overall) | 61.4 | 36.4 | 65.4 | 61.5 | 70.5 | 67.2 | 51.0 | 68.2 | 57.3 | 58.0 | 67.2 | 61.0 | 73.2 | β |
| BFCL v4 (non-live AST) | 79.1 | 78.1 | 83.3 | 85.0 | 85.8 | 88.2 | 83.6 | 83.7 | 35.8 | 83.6 | 85.5 | 45.0 | 85.3 | β |
| BFCL v4 (live) | 78.9 | 73.7 | 78.3 | 78.8 | 80.2 | 81.4 | 82.5 | 80.2 | 70.4 | 78.4 | 78.2 | 77.6 | 79.9 | β |
| BFCL v4 (multi-turn) | 47.5 | 27.5 | 62.7 | 53.5 | 59.9 | 58.1 | 56.0 | 61.4 | 55.4 | 40.4 | 65.2 | 51.7 | 55.5 | β |
| BFCL v4 (memory) | 62.8 | 19.4 | 41.5 | 39.1 | 62.6 | 53.8 | 35.3 | 52.9 | 50.7 | 39.1 | 43.4 | 59.6 | 79.6 | β |
| BFCL v4 (web search) | 62.5 | 10.5 | 69.0 | 66.0 | 75.0 | 68.5 | 12.5 | 75.0 | 57.0 | 69.0 | 71.0 | 71.5 | 82.0 | β |
| BrowseComp | 29.4 | 4.4 | β | 14.5 | 36.5 | 26.9 | 2.8 | 25.5 | 31.2 | β | β | 29.1 | 46.4 | β |
| Code | ||||||||||||||
| Average (EN) | 89.3 | 68.0 | 67.8 | 81.8 | 85.0 | 87.7 | 83.6 | 89.0 | 90.8 | 82.0 | 79.8 | 88.3 | 94.2 | 25.1 |
| LiveCodeBench v6 | 85.9 | 59.2 | 46.5 | 71.3 | 77.8 | 82.5 | 73.9 | 82.3 | 87.5 | 71.2 | 67.8 | 82.0 | 93.8 | 8.7 |
| HumanEval+ | 92.7 | 76.8 | 89.0 | 92.4 | 92.2 | 92.8 | 93.3 | 95.7 | 94.1 | 92.8 | 91.8 | 94.7 | 94.7 | 41.6 |
| SWE-Bench Verified | 66.4 | β | 51.0 | 38.6 | 71.6 | 73.8 | β | 57.8 | β | 60.8 | 11.6 | 60.2 | 72.6 | β |
| Instruction Following | ||||||||||||||
| Average (EN) | 78.1 | 62.5 | 64.5 | 73.2 | 72.7 | 66.1 | 60.7 | 79.9 | 71.1 | 49.8 | 38.2 | 73.7 | 81.9 | 25.5 |
| IFBench (loose-prompt) | 78.1 | 62.5 | 64.5 | 73.2 | 72.7 | 66.1 | 60.7 | 79.9 | 71.1 | 49.8 | 38.2 | 73.7 | 81.9 | 25.5 |
| Grounding / Hallucinations | ||||||||||||||
| Average (EN) | 59.4 | 42.7 | 57.6 | 55.9 | 67.5 | 68.3 | 61.0 | 62.5 | 62.6 | 61.5 | 63.8 | 64.4 | 67.5 | 31.5 |
| SQuAD (M/A Grounding Score) | 23.4 | 0.0 | 1.5 | 0.0 | 23.0 | 12.1 | 0.0 | 0.0 | 0.0 | 6.3 | 0.0 | 0.0 | 8.3 | 0.0 |
| SQuAD (Utility Accuracy) | 83.4 | 77.4 | 82.7 | 75.0 | 89.8 | 88.9 | 73.5 | 88.4 | 74.2 | 76.8 | 81.9 | 80.8 | 85.6 | 1.6 |
| RGB Closed-Book | 51.0 | 52.0 | 78.0 | 80.0 | 81.0 | 79.0 | 89.0 | 79.0 | 85.0 | 86.0 | 92.0 | 93.0 | 73.0 | 85.0 |
| RGB Negative (Abstention) | 85.6 | 73.9 | 79.6 | 76.9 | 86.6 | 79.6 | 81.3 | 86.0 | 79.6 | 82.3 | 89.0 | 74.6 | 70.6 | 55.9 |
| AA-Omniscience Non-Hallucination Rate (1 β Hallucination Rate, public set) | 44.0 | 15.0 | 3.8 | 19.0 | 11.1 | 56.7 | 12.3 | 14.3 | 23.7 | 34.7 | 39.0 | 13.9 | 67.3 | 16.4 |
| RGB Fact-Check (Error Correction) | 34.0 | 14.0 | 68.0 | 58.0 | 74.0 | 74.0 | 85.0 | 61.0 | 77.0 | 60.0 | 74.0 | 90.0 | 53.0 | 14.0 |
| FRAMES (<24k) | 71.2 | 65.7 | 67.8 | 70.9 | 75.9 | 74.7 | 68.9 | 69.5 | 78.3 | 71.9 | 70.4 | 74.9 | 78.6 | 58.9 |
| FRAMES (>24k) | 73.0 | β | 73.9 | 74.3 | 81.5 | 78.8 | 68.5 | 76.6 | 78.4 | 76.6 | β | 78.8 | 83.3 | β |
| SealQA (no distractors, <24k) | 80.7 | 56.6 | 71.7 | 69.0 | 88.3 | 82.8 | 82.1 | 82.1 | 80.7 | 73.1 | 67.6 | 82.8 | 86.9 | 35.9 |
| SealQA (12 distractors, <24k) | 61.0 | 30.0 | 65.0 | 54.0 | 78.0 | 67.0 | 57.0 | 82.0 | 65.0 | 62.0 | 60.0 | 70.0 | 84.0 | 16.0 |
| SealQA (no distractors, >24k) | 100.0 | β | 61.1 | 66.7 | 100.0 | 88.9 | 83.3 | 77.8 | 83.3 | 77.8 | β | 94.4 | 94.4 | β |
| SealQA (12 distractors, >24k) | 65.1 | β | 57.1 | 39.7 | 73.0 | 71.4 | 50.8 | 68.3 | 63.5 | 49.2 | β | 68.3 | 82.5 | β |
| Agentic Retrieval | ||||||||||||||
| Average (EN) | 77.3 | 42.7 | 58.8 | 71.5 | 78.9 | 61.2 | 50.5 | 71.2 | 76.1 | 66.8 | 69.3 | 79.1 | 83.7 | 14.0 |
| Average (DE) | 69.4 | 23.7 | 49.0 | 65.4 | 68.4 | 58.8 | 61.6 | 66.2 | 67.7 | 65.4 | 62.9 | 67.9 | 73.5 | 26.5 |
| MuSiQue (EN) | 77.3 | 42.7 | 58.8 | 71.5 | 78.9 | 61.2 | 50.5 | 71.2 | 76.1 | 66.8 | 69.3 | 79.1 | 83.7 | 14.0 |
| Honeypot | 80.8 | 25.3 | 69.4 | 60.2 | 77.5 | 74.3 | 13.5 | 75.4 | 66.6 | 68.1 | β | 68.8 | 85.0 | β |
| Agentic Wiki QA (DE) | 69.4 | 23.7 | 49.0 | 65.4 | 68.4 | 58.8 | 61.6 | 66.2 | 67.7 | 65.4 | 62.9 | 67.9 | 73.5 | 26.5 |
| Industry RAG | ||||||||||||||
| Average (EN) | 89.7 | 53.8 | 75.9 | 61.3 | 87.0 | 86.0 | 62.7 | 79.4 | 82.8 | 74.9 | 82.3 | 80.3 | 93.3 | 23.9 |
| Average (DE) | 67.5 | 42.7 | 60.9 | 46.2 | 70.0 | 65.8 | 31.1 | 49.4 | 64.2 | 53.4 | 63.4 | 57.6 | 80.2 | 14.4 |
| Semiconductors | 80.4 | 35.3 | 64.7 | 39.2 | 79.4 | 79.4 | 41.2 | 63.7 | 70.6 | 62.7 | 72.5 | 69.6 | 89.2 | 11.8 |
| German Public Sector | 75.0 | 54.0 | 69.5 | 65.5 | 80.0 | 72.0 | 29.5 | 66.5 | 77.0 | 50.0 | 69.5 | 78.0 | 89.0 | 8.0 |
| Aerospace | 58.9 | 14.1 | 50.4 | 44.1 | 62.3 | 59.0 | 48.1 | 58.1 | 48.1 | 47.0 | β | 54.9 | 73.8 | β |
| Automotive Supplier | 99.0 | 72.4 | 87.2 | 83.3 | 94.6 | 92.6 | 84.2 | 95.1 | 95.0 | 87.1 | 92.0 | 91.0 | 97.3 | 35.9 |
| Industrial Drive Technology | 60.0 | 31.4 | 52.3 | 26.8 | 60.0 | 59.5 | 32.7 | 32.3 | 51.4 | 56.8 | 57.3 | 37.3 | 71.4 | 20.9 |
| Long Context | ||||||||||||||
| LongBench Pro | 64.5 | β | β | 53.2 | 70.3 | 70.8 | 63.8 | 64.3 | β | 56.4 | β | 62.9 | 76.9 | β |
| AA-LCR | 68.3 | β | 49.7 | 42.3 | 66.3 | 69.7 | 46.3 | 68.3 | β | 52.3 | β | 67.0 | 81.3 | β |
All models use the same evaluation setup: eval-framework for most benchmarks and Harbor for TerminalBench and SWE-Bench. Each model uses its documented context window and sampling parameters, with Kolibri at reasoning effort
high. A model reports no score (β) when it cannot call tools or when a prompt or agent trajectory exceeds its window; long-context benchmarks instead score an overlength prompt as 0.Category averages are unweighted means over rows scored by every compared model except Apertus. They also exclude AA-Omniscience Index, Honeypot and the individual BFCL v4 splits. Excluded scores are greyed out. Each Overall is the unweighted mean of that language's category averages.
All models in the table below are pre-trained base models. Kolibri Base is the checkpoint Kolibri's post-training starts from.
| Type | MoE | Dense | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Active parameters | 3-4B | 12B | 7B | 32B | 70B | |||||
| Ours | Baseline models | |||||||||
| Eval | Kolibri Base | Kolibri Origin Base | Gemma 4 26B-A4B Base | Nemotron 3 Nano 30B-A3B Base | Qwen3.5 35B-A3B Base | GLM-4.5 Air 106B-A12B Base | Nemotron 3 Super 120B-A12B Base | OLMo 3 7B Base | OLMo 3 32B Base | Apertus 70B Base |
| Overall (EN) | 81.1 | 58.4 | 58.1 | 77.5 | 73.8 | 77.1 | 83.1 | 57.6 | 67.9 | 48.6 |
| Overall (DE) | 81.5 | 61.1 | 61.1 | 76.0 | 76.6 | 79.4 | 85.0 | 48.0 | 65.2 | 50.5 |
| General Knowledge | ||||||||||
| Average (EN) | 81.3 | 70.8 | 77.2 | 80.5 | 82.7 | 82.5 | 87.2 | 67.8 | 77.8 | 73.8 |
| Average (DE) | 82.7 | 70.7 | 80.5 | 80.0 | 84.9 | 81.7 | 87.5 | 51.8 | 68.7 | 74.7 |
| MMLU (EN) | 81.0 | 68.9 | 78.2 | 79.0 | 84.7 | 82.8 | 86.8 | 67.1 | 76.3 | 69.3 |
| Global MMLU (DE) | 77.5 | 65.1 | 75.1 | 74.3 | 81.5 | 78.2 | 84.7 | 51.6 | 65.1 | 64.9 |
| ARC (EN) | 95.7 | 88.9 | 95.2 | 94.3 | 97.0 | 96.3 | 97.5 | 88.6 | 94.5 | 90.7 |
| ARC (DE) | 96.4 | 88.5 | 95.5 | 94.2 | 97.6 | 95.9 | 97.8 | 67.9 | 88.1 | 89.9 |
| PIQA (EN) | 89.5 | 76.9 | 88.0 | 89.4 | 91.9 | 89.8 | 95.1 | 77.7 | 86.8 | 80.1 |
| PIQA (DE) | 96.8 | 87.9 | 96.8 | 96.5 | 98.5 | 97.2 | 99.3 | 75.0 | 90.0 | 94.6 |
| HellaSwag (EN) | 83.1 | 80.7 | 84.8 | 85.4 | 85.3 | 87.2 | 88.8 | 76.1 | 83.5 | 84.4 |
| HellaSwag (DE) | 86.9 | 75.3 | 88.6 | 85.8 | 88.7 | 86.5 | 93.0 | 36.8 | 61.0 | 88.7 |
| MMLU-Pro (EN) | 61.1 | 39.3 | 51.2 | 55.6 | 62.6 | 55.3 | 68.9 | 38.0 | 50.9 | 40.6 |
| MMLU-ProX (DE) | 55.7 | 36.5 | 46.7 | 49.0 | 58.4 | 50.8 | 62.8 | 27.9 | 39.5 | 35.6 |
| TriviaQA (EN) | 77.2 | 69.9 | 65.6 | 79.5 | 74.6 | 83.8 | 86.4 | 59.0 | 74.7 | 77.4 |
| Wahl-O-Mat (DE) | 16.7 | 52.3 | 57.1 | 55.9 | 43.2 | 56.7 | 61.1 | 49.6 | 56.3 | 57.0 |
| Math | ||||||||||
| Average (EN) | 84.9 | 54.3 | 45.9 | 83.0 | 73.5 | 64.7 | 82.7 | 57.7 | 63.3 | 39.6 |
| Average (DE) | 76.8 | 55.8 | 47.0 | 72.5 | 77.5 | 69.8 | 81.2 | 43.8 | 62.9 | 39.0 |
| GSM8K (EN) | 89.8 | 71.6 | 65.7 | 86.3 | 89.5 | 82.8 | 87.5 | 75.3 | 81.1 | 62.3 |
| GSM8K Platinum (DE) | 90.5 | 71.7 | 64.1 | 86.3 | 88.0 | 86.5 | 93.1 | 58.7 | 80.8 | 60.2 |
| MATH Minerva (EN) | 80.1 | 37.0 | 26.1 | 79.8 | 57.6 | 46.5 | 77.8 | 40.1 | 45.6 | 16.9 |
| MATH Minerva (DE) | 63.2 | 39.9 | 29.9 | 58.8 | 67.0 | 53.2 | 69.3 | 28.9 | 45.0 | 17.7 |
| Code | ||||||||||
| Average (EN) | 77.0 | 50.1 | 51.1 | 69.0 | 65.2 | 84.1 | 79.3 | 47.4 | 62.5 | 32.5 |
| Average (DE) | 85.1 | 57.0 | 55.7 | 75.5 | 67.4 | 86.7 | 86.2 | 48.3 | 64.0 | 37.9 |
| HumanEval (EN) | 86.7 | 50.8 | 52.2 | 73.5 | 67.2 | 96.3 | 83.2 | 46.8 | 63.9 | 28.3 |
| HumanEval (DE) | 88.4 | 49.0 | 51.7 | 73.5 | 59.2 | 94.0 | 85.0 | 39.5 | 56.2 | 28.3 |
| MBPP (EN) | 67.3 | 49.5 | 50.0 | 64.5 | 63.1 | 71.9 | 75.4 | 48.0 | 61.0 | 36.8 |
| MBPP (DE) | 81.7 | 64.9 | 59.7 | 77.4 | 75.6 | 79.4 | 87.5 | 57.1 | 71.7 | 47.5 |
All results in the table above are produced with the same evaluation setup for every model, based on our eval-framework; this includes identical prompts, few-shot configurations and task settings. All models use the sampling parameters
temperature = 0.6,top_p = 0.6,max_tokens = 1024and a maximum context length of 65,536 tokens, except Gemma 4 26B-A4B Base, which uses Google's recommendedtemperature = 1.0,top_p = 0.95. Each model is served with its long-context extension on: Qwen3.5 35B-A3B Base with static YaRN (factor 4).Each group score is the unweighted mean of the evals in that group: Math (EN), for example, is the mean of GSM8K (EN) and MATH Minerva (EN). The Overall score is the unweighted mean of the three group scores, so each capability (General Knowledge, Math, Code) contributes equally regardless of how many evals it contains. The aggregates are meant for comparing models within one capability and language, not a model's English against its German scores: an eval is not necessarily equally difficult in both languages, and the groups are not composed identically. General Knowledge (EN) includes MMLU-Pro and TriviaQA while General Knowledge (DE) includes neither, so it and, by extension, Overall (EN) average over two additional and comparatively hard evals.
| Type | MoE | Dense | ||||||
|---|---|---|---|---|---|---|---|---|
| Active parameters | 3-4B | 7B | 32B | 70B | ||||
| Ours | Baseline models | |||||||
| Eval | Kolibri Base | Kolibri Origin Base | Gemma 4 26B-A4B Base | Nemotron 3 Nano 30B-A3B Base | Qwen3.5 35B-A3B Base | OLMo 3 7B Base | OLMo 3 32B Base | Apertus 70B Base |
| RULER | ||||||||
| 4k | 86.9 | 91.8 | 84.7 | 94.6 | 96.3 | 92.0 | 94.8 | 89.5 |
| 8k | 83.7 | 88.8 | 85.0 | 93.4 | 95.3 | 80.8 | 92.6 | 77.7 |
| 16k | 80.9 | 83.7 | 87.1 | 92.0 | 95.2 | 70.8 | 88.8 | 71.3 |
| 32k | 76.3 | 72.7 | 88.7 | 86.4 | 93.7 | 64.4 | 80.9 | 70.8 |
| 64k | 72.2 | β | 86.8 | 82.6 | 91.0 | β | β | β |
| 128k | 67.9 | β | 87.8 | 81.0 | 89.9 | β | β | β |
| 256k | 69.8 | β | β | 72.1 | 80.1 | β | ||
Truncated β view the full README on Hugging Face.
Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a focus on German and English. The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency.
| Model | Kolibri 1 |
|---|---|
| Model Provider | Aleph Alpha GmbH |
| Model Developer | Aleph Alpha Research GmbH |
| Architecture | Mixture-of-Experts |
| Total parameters | 78B (78,103,074,560) |
| Active parameters / token | 3.46B (3,457,573,120) |
| Languages | German, English |
| Context length | 1,048,576 tokens; we recommend β€262,144 tokens for serving efficiency and complex tasks |
| Precision | float8_e4m3fn weights in 128Γ128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in bfloat16 |
| Reasoning mode | Yes |
| Tool calling | Yes |
| License | Apache 2.0 |
| Knowledge cutoff | EN: June 18, 2026, DE: June 18, 2026 This only affects implicit knowledge, the model may use more recent information through tool use. |
| Hardware requirements | Model memory footprint: ~78 GB (FP8 weights). Minimum: 2Γ A100 80 GB, 2Γ H100 SXM5, 1Γ H200, 1Γ B200 or 1Γ B300. Recommended: 2Γ H100 SXM5, 2Γ H200, 1Γ B200 or 1Γ B300. |
| Best for | Multi-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant |
| Release Date | 3rd of October 2026 |
| Code of Practice | Aleph Alpha is a signatory of the EU GPAI Code of Practice, see https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai. |
| Training Data | Pre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension. Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases. |
| Training method | We trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed. |
| Computing Resources | Pre-training (based on actual measurements), excluding mid-training and long-context: Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes); Parallelism: EP8 FSDP16 DP6; Time: 21 days (511h, 392k GPUh) Mid-training: 5 days, 90k GPUh (same setup as above) Long-Context: 13h, 10k GPUh (same setup as above except parallelism: FSDP128 DP6) FLOPS: 6.4e23 |
| Tech Report | https://aleph-alpha.com/downloads/tech-report.pdf |
Kolibri is intended to process text input and output in German and English and to perform a wide range of tasks beyond natural-language generation, for example multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and agentic tool calling.
Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length. Because positional encoding is applied only in the sliding-window layers, the context can be extended beyond that length without any position scaling, in principle to arbitrary lengths. We have validated quality and serving efficiency up to 1,048,576 tokens. For latency- or throughput-sensitive deployments and for complex tasks, we recommend contexts of at most 262,144 tokens. See the technical report for details.
For the kinds of systems Kolibri is meant to be integrated into, see AI system types below; for uses that we encourage users to refrain from, see Responsible Use.
Kolibri is designed to deliver strong German and English performance at low serving cost. Its mixture-of-experts architecture activates only a small fraction of its parameters for each token, keeping compute per token low while retaining the capacity of a much larger model. The trade-off is memory: the full model must be held in memory even though only part of it is active at any time. To keep long contexts affordable, most attention layers focus on nearby text, while a smaller number attend across the whole context. We also developed a tokenizer tailored to German word structure, so German text is processed efficiently without sacrificing English. Supporting two languages rather than many is a deliberate choice of depth over breadth.
Kolibri is intended for integration into conversational assistants and agentic workflows in German and English, in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed. It suits document-processing and drafting systems, question-answering systems over an organisation's own material, and internal knowledge and research tools. Its tool-calling and structured-output capabilities make it appropriate for orchestration layers that call APIs, execute code or run searches, provided the calling system validates the results. In decision-support systems it belongs on the advisory side, surfacing evidence and drafting options for a human to weigh, and it is not intended as the deciding component. More broadly, it is built for human-AI collaboration rather than unsupervised operation.
9.5Γ10Β² MWh (estimated), including node power and data-centre overhead (PUE). This includes pre-training, mid-training and long-context training. It excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models.
For pre-training, we obtained the average power usage per node and the power usage effectiveness (PUE) from the cluster provider. We multiplied these with the number of nodes and the runtime to obtain a total energy estimate.
Kolibri requires the aleph-alpha-inference package that provides the
Kolibri vLLM plugin. You can either use the provided container image
ghcr.io/aleph-alpha/aleph-alpha-inference, or install the package from
Aleph-Alpha/aleph-alpha-inference,
which also installs the vLLM version it supports:
pip install 'aleph-alpha-inference>=1'
Serve the model with reasoning and tool-calling enabled:
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.
The recommended sampling parameters for the model are temperature=1.0,
top_p=0.97 and top_k=128.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="Aleph-Alpha/Kolibri-1",
messages=[
{
"role": "user",
"content": "ErklΓ€re kurz, was ein Mixture-of-Experts-Modell ist.",
},
],
extra_body={
"chat_template_kwargs": {
"reasoning_effort": "high",
"enable_thinking": True,
}
},
)
print(response.choices[0].message.content)
Kolibri supports explicit thinking effort levels that need to be configured
through the chat template. You can pass reasoning_effort values low,
medium and high to configure the amount of effort our model puts into
finding the answer. You can disable thinking altogether by setting
reasoning_effort to none or passing enable_thinking=false. The model
will then immediately respond.
The serving command enables Hermes-style tool calling (--tool-call-parser kolibri1 --enable-auto-tool-choice). Pass your function schemas via the
standard tools field of the chat completions request; the model will emit
tool calls that the parser converts into structured output. Tool calling can
be combined with reasoning mode.
import json
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
MODEL = "Aleph-Alpha/Kolibri-1"
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, e.g. Heidelberg",
},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["city"],
},
},
}
]
def get_weather(city, unit="celsius"):
return {"city": city, "temperature": 18, "unit": unit, "conditions": "cloudy"}
messages = [{"role": "user", "content": "Wie ist das Wetter gerade in Heidelberg?"}]
first = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
msg = first.choices[0].message
messages.append(msg)
for call in msg.tool_calls or []:
args = json.loads(call.function.arguments)
result = get_weather(**args)
messages.append(
{
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
}
)
final = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
print(final.choices[0].message.content)
The best value per row is bolded. In each table, the best value of each group with more than one model is underlined, and both marks compare the MoE models only. The dense models activate several times as many parameters per token, so they are greyed out and unmarked.
| Type | MoE | Dense | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Active parameters | 3B | 4-6B | 12B | 27B | 70B | |||||||||
| Ours | Baseline models | |||||||||||||
| Eval | Kolibri | Kolibri Origin | GLM-4.7 Flash 30B-A3B | Nemotron 3 Nano 30B-A3B | Qwen3.5 35B-A3B | Qwen3.6 35B-A3B | Qwen3-Next 80B-A3B Thinking | Gemma 4 26B-A4B IT | GPT-OSS 120B | Mistral Small 4 119B-A6B | GLM-4.5 Air 106B-A12B | Nemotron 3 Super 120B-A12B | Qwen3.8 27B | Apertus 70B Instruct |
| Overall (EN) | 75.5 | 54.1 | 64.7 | 65.6 | 74.7 | 71.4 | 62.4 | 71.9 | 72.3 | 63.1 | 64.4 | 73.0 | 80.2 | β |
| Overall (DE) | 70.8 | 46.4 | 50.4 | 59.3 | 69.8 | 67.3 | 58.0 | 66.3 | 70.2 | 61.4 | 64.8 | 67.9 | 79.9 | β |
| Knowledge | ||||||||||||||
| Average (EN) | 50.1 | 39.7 | 45.5 | 46.0 | 52.7 | 52.1 | 48.4 | 51.4 | 50.0 | 47.5 | 45.7 | 52.0 | 56.8 | 22.8 |
| Average (DE) | 57.6 | 44.9 | 46.5 | 41.2 | 61.3 | 61.0 | 55.7 | 61.5 | 58.0 | 51.4 | 52.2 | 59.5 | 69.2 | 24.8 |
| GPQA Diamond (EN) | 84.3 | 68.1 | 73.1 | 73.9 | 83.8 | 83.4 | 76.1 | 81.1 | 76.4 | 74.7 | 73.2 | 78.0 | 89.2 | 29.5 |
| GPQA Diamond (DE) | 81.3 | 58.5 | 59.8 | 49.6 | 84.2 | 80.6 | 72.2 | 80.1 | 76.0 | 72.9 | 71.1 | 76.6 | 88.1 | 31.4 |
| Humanity's Last Exam (EN) | 21.5 | 9.4 | 15.4 | 12.1 | 20.4 | 21.1 | 11.6 | 19.2 | 19.4 | 9.7 | 8.7 | 20.6 | 35.6 | 5.2 |
| Humanity's Last Exam (DE) | 15.9 | 10.4 | 9.1 | 13.1 | 18.1 | 20.5 | 15.6 | 23.4 | 20.7 | 10.5 | 10.5 | 22.3 | 37.2 | 5.7 |
| AA-Omniscience Accuracy (public set) | 14.8 | 11.3 | 17.0 | 19.5 | 22.2 | 21.0 | 24.2 | 20.7 | 23.3 | 25.0 | 20.0 | 26.7 | 19.0 | 13.5 |
| AA-Omniscience Index (public set) | -32.8 | -64.2 | -62.8 | -45.7 | -46.2 | -12.5 | -42.3 | -47.3 | -35.2 | -24.0 | -28.8 | -36.5 | -5.8 | β |
| MMLU-Pro CoT (EN) | 80.0 | 70.1 | 76.5 | 78.3 | 84.6 | 84.3 | 81.7 | 84.5 | 80.8 | 80.4 | 80.9 | 82.7 | 85.0 | 43.0 |
| MMLU-ProX CoT (DE) | 75.5 | 65.7 | 70.7 | 61.0 | 81.7 | 81.9 | 79.4 | 81.1 | 77.2 | 70.7 | 74.9 | 79.7 | 82.4 | 37.3 |
| Math | ||||||||||||||
| Average (EN) | 96.5 | 81.7 | 88.8 | 88.8 | 90.1 | 87.8 | 86.3 | 87.4 | 90.7 | 81.4 | 82.8 | 91.1 | 97.8 | 0.6 |
| Average (DE) | 88.8 | 74.3 | 45.2 | 84.3 | 79.6 | 83.7 | 83.7 | 88.1 | 90.8 | 75.4 | 80.6 | 86.5 | 96.7 | 0.1 |
| AIME 2025 (EN) | 96.9 | 81.9 | 89.4 | 89.6 | 88.1 | 84.6 | 84.2 | 87.3 | 90.6 | 79.8 | 81.9 | 91.7 | 97.9 | 0.6 |
| AIME 2025 (DE) | 87.5 | 73.5 | 43.8 | 84.4 | 76.7 | 82.9 | 80.6 | 88.1 | 90.6 | 72.3 | 80.6 | 85.6 | 96.5 | 0.2 |
| AIME 2026 (EN) | 96.0 | 81.5 | 88.3 | 87.9 | 92.1 | 91.0 | 88.5 | 87.5 | 90.8 | 83.1 | 83.8 | 90.4 | 97.7 | 0.6 |
| AIME 2026 (DE) | 90.0 | 75.2 | 46.7 | 84.2 | 82.5 | 84.4 | 86.7 | 88.1 | 91.0 | 78.5 | 80.6 | 87.5 | 96.9 | 0.0 |
| Agentic | ||||||||||||||
| Average (EN) | 63.4 | 41.6 | 58.9 | 46.4 | 63.4 | 62.1 | 46.3 | 54.6 | 54.0 | 40.7 | 53.5 | 54.9 | 66.7 | β |
| TerminalBench 2.1 | 27.7 | β | 20.2 | 9.7 | 39.7 | β | 8.6 | β | 29.2 | 21.0 | β | 39.7 | 76.8 | β |
| Tau2-Bench (Telecom) | 94.7 | 67.5 | 95.9 | 45.9 | 97.7 | 99.1 | 43.9 | 45.3 | 73.1 | 41.5 | 53.8 | 68.1 | 82.5 | 10.8 |
| Tau2-Bench (Retail) | 69.9 | 58.5 | 57.9 | 64.9 | 70.8 | 71.6 | 60.8 | 71.3 | 60.5 | 62.9 | 61.4 | 67.5 | 68.7 | 9.6 |
| Tau2-Bench (Airline) | 76.7 | 58.7 | 68.7 | 52.7 | 76.0 | 70.7 | 65.3 | 73.3 | 72.7 | 40.0 | 70.7 | 72.7 | 83.3 | 40.0 |
| Tau3-Bench (Banking) | 38.1 | 5.7 | 7.2 | 5.7 | 11.3 | 10.6 | 5.4 | 16.0 | 14.7 | 5.7 | 6.4 | 15.5 | 50.0 | 2.1 |
| BFCL v3 (multi-turn) | 39.8 | 22.8 | 58.2 | 47.9 | 54.0 | 53.5 | 51.4 | 53.4 | 45.6 | 36.2 | 61.6 | 44.6 | 42.5 | 0.6 |
| BFCL v4 (overall) | 61.4 | 36.4 | 65.4 | 61.5 | 70.5 | 67.2 | 51.0 | 68.2 | 57.3 | 58.0 | 67.2 | 61.0 | 73.2 | β |
| BFCL v4 (non-live AST) | 79.1 | 78.1 | 83.3 | 85.0 | 85.8 | 88.2 | 83.6 | 83.7 | 35.8 | 83.6 | 85.5 | 45.0 | 85.3 | β |
| BFCL v4 (live) | 78.9 | 73.7 | 78.3 | 78.8 | 80.2 | 81.4 | 82.5 | 80.2 | 70.4 | 78.4 | 78.2 | 77.6 | 79.9 | β |
| BFCL v4 (multi-turn) | 47.5 | 27.5 | 62.7 | 53.5 | 59.9 | 58.1 | 56.0 | 61.4 | 55.4 | 40.4 | 65.2 | 51.7 | 55.5 | β |
| BFCL v4 (memory) | 62.8 | 19.4 | 41.5 | 39.1 | 62.6 | 53.8 | 35.3 | 52.9 | 50.7 | 39.1 | 43.4 | 59.6 | 79.6 | β |
| BFCL v4 (web search) | 62.5 | 10.5 | 69.0 | 66.0 | 75.0 | 68.5 | 12.5 | 75.0 | 57.0 | 69.0 | 71.0 | 71.5 | 82.0 | β |
| BrowseComp | 29.4 | 4.4 | β | 14.5 | 36.5 | 26.9 | 2.8 | 25.5 | 31.2 | β | β | 29.1 | 46.4 | β |
| Code | ||||||||||||||
| Average (EN) | 89.3 | 68.0 | 67.8 | 81.8 | 85.0 | 87.7 | 83.6 | 89.0 | 90.8 | 82.0 | 79.8 | 88.3 | 94.2 | 25.1 |
| LiveCodeBench v6 | 85.9 | 59.2 | 46.5 | 71.3 | 77.8 | 82.5 | 73.9 | 82.3 | 87.5 | 71.2 | 67.8 | 82.0 | 93.8 | 8.7 |
| HumanEval+ | 92.7 | 76.8 | 89.0 | 92.4 | 92.2 | 92.8 | 93.3 | 95.7 | 94.1 | 92.8 | 91.8 | 94.7 | 94.7 | 41.6 |
| SWE-Bench Verified | 66.4 | β | 51.0 | 38.6 | 71.6 | 73.8 | β | 57.8 | β | 60.8 | 11.6 | 60.2 | 72.6 | β |
| Instruction Following | ||||||||||||||
| Average (EN) | 78.1 | 62.5 | 64.5 | 73.2 | 72.7 | 66.1 | 60.7 | 79.9 | 71.1 | 49.8 | 38.2 | 73.7 | 81.9 | 25.5 |
| IFBench (loose-prompt) | 78.1 | 62.5 | 64.5 | 73.2 | 72.7 | 66.1 | 60.7 | 79.9 | 71.1 | 49.8 | 38.2 | 73.7 | 81.9 | 25.5 |
| Grounding / Hallucinations | ||||||||||||||
| Average (EN) | 59.4 | 42.7 | 57.6 | 55.9 | 67.5 | 68.3 | 61.0 | 62.5 | 62.6 | 61.5 | 63.8 | 64.4 | 67.5 | 31.5 |
| SQuAD (M/A Grounding Score) | 23.4 | 0.0 | 1.5 | 0.0 | 23.0 | 12.1 | 0.0 | 0.0 | 0.0 | 6.3 | 0.0 | 0.0 | 8.3 | 0.0 |
| SQuAD (Utility Accuracy) | 83.4 | 77.4 | 82.7 | 75.0 | 89.8 | 88.9 | 73.5 | 88.4 | 74.2 | 76.8 | 81.9 | 80.8 | 85.6 | 1.6 |
| RGB Closed-Book | 51.0 | 52.0 | 78.0 | 80.0 | 81.0 | 79.0 | 89.0 | 79.0 | 85.0 | 86.0 | 92.0 | 93.0 | 73.0 | 85.0 |
| RGB Negative (Abstention) | 85.6 | 73.9 | 79.6 | 76.9 | 86.6 | 79.6 | 81.3 | 86.0 | 79.6 | 82.3 | 89.0 | 74.6 | 70.6 | 55.9 |
| AA-Omniscience Non-Hallucination Rate (1 β Hallucination Rate, public set) | 44.0 | 15.0 | 3.8 | 19.0 | 11.1 | 56.7 | 12.3 | 14.3 | 23.7 | 34.7 | 39.0 | 13.9 | 67.3 | 16.4 |
| RGB Fact-Check (Error Correction) | 34.0 | 14.0 | 68.0 | 58.0 | 74.0 | 74.0 | 85.0 | 61.0 | 77.0 | 60.0 | 74.0 | 90.0 | 53.0 | 14.0 |
| FRAMES (<24k) | 71.2 | 65.7 | 67.8 | 70.9 | 75.9 | 74.7 | 68.9 | 69.5 | 78.3 | 71.9 | 70.4 | 74.9 | 78.6 | 58.9 |
| FRAMES (>24k) | 73.0 | β | 73.9 | 74.3 | 81.5 | 78.8 | 68.5 | 76.6 | 78.4 | 76.6 | β | 78.8 | 83.3 | β |
| SealQA (no distractors, <24k) | 80.7 | 56.6 | 71.7 | 69.0 | 88.3 | 82.8 | 82.1 | 82.1 | 80.7 | 73.1 | 67.6 | 82.8 | 86.9 | 35.9 |
| SealQA (12 distractors, <24k) | 61.0 | 30.0 | 65.0 | 54.0 | 78.0 | 67.0 | 57.0 | 82.0 | 65.0 | 62.0 | 60.0 | 70.0 | 84.0 | 16.0 |
| SealQA (no distractors, >24k) | 100.0 | β | 61.1 | 66.7 | 100.0 | 88.9 | 83.3 | 77.8 | 83.3 | 77.8 | β | 94.4 | 94.4 | β |
| SealQA (12 distractors, >24k) | 65.1 | β | 57.1 | 39.7 | 73.0 | 71.4 | 50.8 | 68.3 | 63.5 | 49.2 | β | 68.3 | 82.5 | β |
| Agentic Retrieval | ||||||||||||||
| Average (EN) | 77.3 | 42.7 | 58.8 | 71.5 | 78.9 | 61.2 | 50.5 | 71.2 | 76.1 | 66.8 | 69.3 | 79.1 | 83.7 | 14.0 |
| Average (DE) | 69.4 | 23.7 | 49.0 | 65.4 | 68.4 | 58.8 | 61.6 | 66.2 | 67.7 | 65.4 | 62.9 | 67.9 | 73.5 | 26.5 |
| MuSiQue (EN) | 77.3 | 42.7 | 58.8 | 71.5 | 78.9 | 61.2 | 50.5 | 71.2 | 76.1 | 66.8 | 69.3 | 79.1 | 83.7 | 14.0 |
| Honeypot | 80.8 | 25.3 | 69.4 | 60.2 | 77.5 | 74.3 | 13.5 | 75.4 | 66.6 | 68.1 | β | 68.8 | 85.0 | β |
| Agentic Wiki QA (DE) | 69.4 | 23.7 | 49.0 | 65.4 | 68.4 | 58.8 | 61.6 | 66.2 | 67.7 | 65.4 | 62.9 | 67.9 | 73.5 | 26.5 |
| Industry RAG | ||||||||||||||
| Average (EN) | 89.7 | 53.8 | 75.9 | 61.3 | 87.0 | 86.0 | 62.7 | 79.4 | 82.8 | 74.9 | 82.3 | 80.3 | 93.3 | 23.9 |
| Average (DE) | 67.5 | 42.7 | 60.9 | 46.2 | 70.0 | 65.8 | 31.1 | 49.4 | 64.2 | 53.4 | 63.4 | 57.6 | 80.2 | 14.4 |
| Semiconductors | 80.4 | 35.3 | 64.7 | 39.2 | 79.4 | 79.4 | 41.2 | 63.7 | 70.6 | 62.7 | 72.5 | 69.6 | 89.2 | 11.8 |
| German Public Sector | 75.0 | 54.0 | 69.5 | 65.5 | 80.0 | 72.0 | 29.5 | 66.5 | 77.0 | 50.0 | 69.5 | 78.0 | 89.0 | 8.0 |
| Aerospace | 58.9 | 14.1 | 50.4 | 44.1 | 62.3 | 59.0 | 48.1 | 58.1 | 48.1 | 47.0 | β | 54.9 | 73.8 | β |
| Automotive Supplier | 99.0 | 72.4 | 87.2 | 83.3 | 94.6 | 92.6 | 84.2 | 95.1 | 95.0 | 87.1 | 92.0 | 91.0 | 97.3 | 35.9 |
| Industrial Drive Technology | 60.0 | 31.4 | 52.3 | 26.8 | 60.0 | 59.5 | 32.7 | 32.3 | 51.4 | 56.8 | 57.3 | 37.3 | 71.4 | 20.9 |
| Long Context | ||||||||||||||
| LongBench Pro | 64.5 | β | β | 53.2 | 70.3 | 70.8 | 63.8 | 64.3 | β | 56.4 | β | 62.9 | 76.9 | β |
| AA-LCR | 68.3 | β | 49.7 | 42.3 | 66.3 | 69.7 | 46.3 | 68.3 | β | 52.3 | β | 67.0 | 81.3 | β |
All models use the same evaluation setup: eval-framework for most benchmarks and Harbor for TerminalBench and SWE-Bench. Each model uses its documented context window and sampling parameters, with Kolibri at reasoning effort
high. A model reports no score (β) when it cannot call tools or when a prompt or agent trajectory exceeds its window; long-context benchmarks instead score an overlength prompt as 0.Category averages are unweighted means over rows scored by every compared model except Apertus. They also exclude AA-Omniscience Index, Honeypot and the individual BFCL v4 splits. Excluded scores are greyed out. Each Overall is the unweighted mean of that language's category averages.
All models in the table below are pre-trained base models. Kolibri Base is the checkpoint Kolibri's post-training starts from.
| Type | MoE | Dense | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Active parameters | 3-4B | 12B | 7B | 32B | 70B | |||||
| Ours | Baseline models | |||||||||
| Eval | Kolibri Base | Kolibri Origin Base | Gemma 4 26B-A4B Base | Nemotron 3 Nano 30B-A3B Base | Qwen3.5 35B-A3B Base | GLM-4.5 Air 106B-A12B Base | Nemotron 3 Super 120B-A12B Base | OLMo 3 7B Base | OLMo 3 32B Base | Apertus 70B Base |
| Overall (EN) | 81.1 | 58.4 | 58.1 | 77.5 | 73.8 | 77.1 | 83.1 | 57.6 | 67.9 | 48.6 |
| Overall (DE) | 81.5 | 61.1 | 61.1 | 76.0 | 76.6 | 79.4 | 85.0 | 48.0 | 65.2 | 50.5 |
| General Knowledge | ||||||||||
| Average (EN) | 81.3 | 70.8 | 77.2 | 80.5 | 82.7 | 82.5 | 87.2 | 67.8 | 77.8 | 73.8 |
| Average (DE) | 82.7 | 70.7 | 80.5 | 80.0 | 84.9 | 81.7 | 87.5 | 51.8 | 68.7 | 74.7 |
| MMLU (EN) | 81.0 | 68.9 | 78.2 | 79.0 | 84.7 | 82.8 | 86.8 | 67.1 | 76.3 | 69.3 |
| Global MMLU (DE) | 77.5 | 65.1 | 75.1 | 74.3 | 81.5 | 78.2 | 84.7 | 51.6 | 65.1 | 64.9 |
| ARC (EN) | 95.7 | 88.9 | 95.2 | 94.3 | 97.0 | 96.3 | 97.5 | 88.6 | 94.5 | 90.7 |
| ARC (DE) | 96.4 | 88.5 | 95.5 | 94.2 | 97.6 | 95.9 | 97.8 | 67.9 | 88.1 | 89.9 |
| PIQA (EN) | 89.5 | 76.9 | 88.0 | 89.4 | 91.9 | 89.8 | 95.1 | 77.7 | 86.8 | 80.1 |
| PIQA (DE) | 96.8 | 87.9 | 96.8 | 96.5 | 98.5 | 97.2 | 99.3 | 75.0 | 90.0 | 94.6 |
| HellaSwag (EN) | 83.1 | 80.7 | 84.8 | 85.4 | 85.3 | 87.2 | 88.8 | 76.1 | 83.5 | 84.4 |
| HellaSwag (DE) | 86.9 | 75.3 | 88.6 | 85.8 | 88.7 | 86.5 | 93.0 | 36.8 | 61.0 | 88.7 |
| MMLU-Pro (EN) | 61.1 | 39.3 | 51.2 | 55.6 | 62.6 | 55.3 | 68.9 | 38.0 | 50.9 | 40.6 |
| MMLU-ProX (DE) | 55.7 | 36.5 | 46.7 | 49.0 | 58.4 | 50.8 | 62.8 | 27.9 | 39.5 | 35.6 |
| TriviaQA (EN) | 77.2 | 69.9 | 65.6 | 79.5 | 74.6 | 83.8 | 86.4 | 59.0 | 74.7 | 77.4 |
| Wahl-O-Mat (DE) | 16.7 | 52.3 | 57.1 | 55.9 | 43.2 | 56.7 | 61.1 | 49.6 | 56.3 | 57.0 |
| Math | ||||||||||
| Average (EN) | 84.9 | 54.3 | 45.9 | 83.0 | 73.5 | 64.7 | 82.7 | 57.7 | 63.3 | 39.6 |
| Average (DE) | 76.8 | 55.8 | 47.0 | 72.5 | 77.5 | 69.8 | 81.2 | 43.8 | 62.9 | 39.0 |
| GSM8K (EN) | 89.8 | 71.6 | 65.7 | 86.3 | 89.5 | 82.8 | 87.5 | 75.3 | 81.1 | 62.3 |
| GSM8K Platinum (DE) | 90.5 | 71.7 | 64.1 | 86.3 | 88.0 | 86.5 | 93.1 | 58.7 | 80.8 | 60.2 |
| MATH Minerva (EN) | 80.1 | 37.0 | 26.1 | 79.8 | 57.6 | 46.5 | 77.8 | 40.1 | 45.6 | 16.9 |
| MATH Minerva (DE) | 63.2 | 39.9 | 29.9 | 58.8 | 67.0 | 53.2 | 69.3 | 28.9 | 45.0 | 17.7 |
| Code | ||||||||||
| Average (EN) | 77.0 | 50.1 | 51.1 | 69.0 | 65.2 | 84.1 | 79.3 | 47.4 | 62.5 | 32.5 |
| Average (DE) | 85.1 | 57.0 | 55.7 | 75.5 | 67.4 | 86.7 | 86.2 | 48.3 | 64.0 | 37.9 |
| HumanEval (EN) | 86.7 | 50.8 | 52.2 | 73.5 | 67.2 | 96.3 | 83.2 | 46.8 | 63.9 | 28.3 |
| HumanEval (DE) | 88.4 | 49.0 | 51.7 | 73.5 | 59.2 | 94.0 | 85.0 | 39.5 | 56.2 | 28.3 |
| MBPP (EN) | 67.3 | 49.5 | 50.0 | 64.5 | 63.1 | 71.9 | 75.4 | 48.0 | 61.0 | 36.8 |
| MBPP (DE) | 81.7 | 64.9 | 59.7 | 77.4 | 75.6 | 79.4 | 87.5 | 57.1 | 71.7 | 47.5 |
All results in the table above are produced with the same evaluation setup for every model, based on our eval-framework; this includes identical prompts, few-shot configurations and task settings. All models use the sampling parameters
temperature = 0.6,top_p = 0.6,max_tokens = 1024and a maximum context length of 65,536 tokens, except Gemma 4 26B-A4B Base, which uses Google's recommendedtemperature = 1.0,top_p = 0.95. Each model is served with its long-context extension on: Qwen3.5 35B-A3B Base with static YaRN (factor 4).Each group score is the unweighted mean of the evals in that group: Math (EN), for example, is the mean of GSM8K (EN) and MATH Minerva (EN). The Overall score is the unweighted mean of the three group scores, so each capability (General Knowledge, Math, Code) contributes equally regardless of how many evals it contains. The aggregates are meant for comparing models within one capability and language, not a model's English against its German scores: an eval is not necessarily equally difficult in both languages, and the groups are not composed identically. General Knowledge (EN) includes MMLU-Pro and TriviaQA while General Knowledge (DE) includes neither, so it and, by extension, Overall (EN) average over two additional and comparatively hard evals.
| Type | MoE | Dense | ||||||
|---|---|---|---|---|---|---|---|---|
| Active parameters | 3-4B | 7B | 32B | 70B | ||||
| Ours | Baseline models | |||||||
| Eval | Kolibri Base | Kolibri Origin Base | Gemma 4 26B-A4B Base | Nemotron 3 Nano 30B-A3B Base | Qwen3.5 35B-A3B Base | OLMo 3 7B Base | OLMo 3 32B Base | Apertus 70B Base |
| RULER | ||||||||
| 4k | 86.9 | 91.8 | 84.7 | 94.6 | 96.3 | 92.0 | 94.8 | 89.5 |
| 8k | 83.7 | 88.8 | 85.0 | 93.4 | 95.3 | 80.8 | 92.6 | 77.7 |
| 16k | 80.9 | 83.7 | 87.1 | 92.0 | 95.2 | 70.8 | 88.8 | 71.3 |
| 32k | 76.3 | 72.7 | 88.7 | 86.4 | 93.7 | 64.4 | 80.9 | 70.8 |
| 64k | 72.2 | β | 86.8 | 82.6 | 91.0 | β | β | β |
| 128k | 67.9 | β | 87.8 | 81.0 | 89.9 | β | β | β |
| 256k | 69.8 | β | β | 72.1 | 80.1 | β | ||
Truncated β view the full README on Hugging Face.