Aleph-Alpha/Kolibri-1

Model

Kolibri

94

3 commits

3 linked in READMEs

updated Oct 3, 2026

See the code

README

Kolibri

Tech report | Tech blog

Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a focus on German and English. The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency.

Model overview

ModelKolibri 1
Model ProviderAleph Alpha GmbH
Model DeveloperAleph Alpha Research GmbH
ArchitectureMixture-of-Experts
Total parameters78B (78,103,074,560)
Active parameters / token3.46B (3,457,573,120)
LanguagesGerman, English
Context length1,048,576 tokens; we recommend ≀262,144 tokens for serving efficiency and complex tasks
Precisionfloat8_e4m3fn weights in 128Γ—128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in bfloat16
Reasoning modeYes
Tool callingYes
LicenseApache 2.0
Knowledge cutoffEN: June 18, 2026, DE: June 18, 2026
This only affects implicit knowledge, the model may use more recent information through tool use.
Hardware requirementsModel memory footprint: ~78 GB (FP8 weights). Minimum: 2Γ— A100 80 GB, 2Γ— H100 SXM5, 1Γ— H200, 1Γ— B200 or 1Γ— B300. Recommended: 2Γ— H100 SXM5, 2Γ— H200, 1Γ— B200 or 1Γ— B300.
Best forMulti-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant
Release Date3rd of October 2026
Code of PracticeAleph Alpha is a signatory of the EU GPAI Code of Practice, see https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai.
Training DataPre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension.
Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases.
Training methodWe trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed.
Computing ResourcesPre-training (based on actual measurements), excluding mid-training and long-context: Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes); Parallelism: EP8 FSDP16 DP6; Time: 21 days (511h, 392k GPUh)
Mid-training: 5 days, 90k GPUh (same setup as above)
Long-Context: 13h, 10k GPUh (same setup as above except parallelism: FSDP128 DP6)
FLOPS: 6.4e23
Tech Reporthttps://aleph-alpha.com/downloads/tech-report.pdf

Intended use

Kolibri is intended to process text input and output in German and English and to perform a wide range of tasks beyond natural-language generation, for example multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and agentic tool calling.

Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length. Because positional encoding is applied only in the sliding-window layers, the context can be extended beyond that length without any position scaling, in principle to arbitrary lengths. We have validated quality and serving efficiency up to 1,048,576 tokens. For latency- or throughput-sensitive deployments and for complex tasks, we recommend contexts of at most 262,144 tokens. See the technical report for details.

For the kinds of systems Kolibri is meant to be integrated into, see AI system types below; for uses that we encourage users to refrain from, see Responsible Use.

Design goals

Kolibri is designed to deliver strong German and English performance at low serving cost. Its mixture-of-experts architecture activates only a small fraction of its parameters for each token, keeping compute per token low while retaining the capacity of a much larger model. The trade-off is memory: the full model must be held in memory even though only part of it is active at any time. To keep long contexts affordable, most attention layers focus on nearby text, while a smaller number attend across the whole context. We also developed a tokenizer tailored to German word structure, so German text is processed efficiently without sacrificing English. Supporting two languages rather than many is a deliberate choice of depth over breadth.

AI system types

Kolibri is intended for integration into conversational assistants and agentic workflows in German and English, in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed. It suits document-processing and drafting systems, question-answering systems over an organisation's own material, and internal knowledge and research tools. Its tool-calling and structured-output capabilities make it appropriate for orchestration layers that call APIs, execute code or run searches, provided the calling system validates the results. In decision-support systems it belongs on the advisory side, surfacing evidence and drafting options for a human to weigh, and it is not intended as the deciding component. More broadly, it is built for human-AI collaboration rather than unsupervised operation.

Sustainability

Energy consumption

9.5Γ—10Β² MWh (estimated), including node power and data-centre overhead (PUE). This includes pre-training, mid-training and long-context training. It excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models.

Energy measurement

For pre-training, we obtained the average power usage per node and the power usage effectiveness (PUE) from the cluster provider. We multiplied these with the number of nodes and the runtime to obtain a total energy estimate.

Getting started

Kolibri requires the aleph-alpha-inference package that provides the Kolibri vLLM plugin. You can either use the provided container image ghcr.io/aleph-alpha/aleph-alpha-inference, or install the package from Aleph-Alpha/aleph-alpha-inference, which also installs the vLLM version it supports:

pip install 'aleph-alpha-inference>=1'

Serve the model with reasoning and tool-calling enabled:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.

The recommended sampling parameters for the model are temperature=1.0, top_p=0.97 and top_k=128.

Querying the server (OpenAI-compatible API)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="Aleph-Alpha/Kolibri-1",
    messages=[
        {
            "role": "user",
            "content": "ErklΓ€re kurz, was ein Mixture-of-Experts-Modell ist.",
        },
    ],
    extra_body={
        "chat_template_kwargs": {
            "reasoning_effort": "high",
            "enable_thinking": True,
        }
    },
)
print(response.choices[0].message.content)

Reasoning mode

Kolibri supports explicit thinking effort levels that need to be configured through the chat template. You can pass reasoning_effort values low, medium and high to configure the amount of effort our model puts into finding the answer. You can disable thinking altogether by setting reasoning_effort to none or passing enable_thinking=false. The model will then immediately respond.

Tool calling

The serving command enables Hermes-style tool calling (--tool-call-parser kolibri1 --enable-auto-tool-choice). Pass your function schemas via the standard tools field of the chat completions request; the model will emit tool calls that the parser converts into structured output. Tool calling can be combined with reasoning mode.

import json
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
MODEL = "Aleph-Alpha/Kolibri-1"

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {
                        "type": "string",
                        "description": "City name, e.g. Heidelberg",
                    },
                    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
                },
                "required": ["city"],
            },
        },
    }
]


def get_weather(city, unit="celsius"):
    return {"city": city, "temperature": 18, "unit": unit, "conditions": "cloudy"}


messages = [{"role": "user", "content": "Wie ist das Wetter gerade in Heidelberg?"}]

first = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
msg = first.choices[0].message
messages.append(msg)

for call in msg.tool_calls or []:
    args = json.loads(call.function.arguments)
    result = get_weather(**args)
    messages.append(
        {
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result),
        }
    )

final = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
print(final.choices[0].message.content)

Evaluation

The best value per row is bolded. In each table, the best value of each group with more than one model is underlined, and both marks compare the MoE models only. The dense models activate several times as many parameters per token, so they are greyed out and unmarked.

Post-training

TypeMoEDense
Active parameters3B4-6B12B27B70B
OursBaseline models
Eval
Kolibri
Kolibri Origin
GLM-4.7 Flash 30B-A3B
Nemotron 3 Nano 30B-A3B
Qwen3.5 35B-A3B
Qwen3.6 35B-A3B
Qwen3-Next 80B-A3B Thinking
Gemma 4 26B-A4B IT
GPT-OSS 120B
Mistral Small 4 119B-A6B
GLM-4.5 Air 106B-A12B
Nemotron 3 Super 120B-A12B
Qwen3.8 27B
Apertus 70B Instruct
Overall (EN)75.554.164.765.674.771.462.471.972.363.164.473.080.2–
Overall (DE)70.846.450.459.369.867.358.066.370.261.464.867.979.9–
Knowledge
Average (EN)50.139.745.546.052.752.148.451.450.047.545.752.056.822.8
Average (DE)57.644.946.541.261.361.055.761.558.051.452.259.569.224.8
GPQA Diamond (EN)84.368.173.173.983.883.476.181.176.474.773.278.089.229.5
GPQA Diamond (DE)81.358.559.849.684.280.672.280.176.072.971.176.688.131.4
Humanity's Last Exam (EN)21.59.415.412.120.421.111.619.219.49.78.720.635.65.2
Humanity's Last Exam (DE)15.910.49.113.118.120.515.623.420.710.510.522.337.25.7
AA-Omniscience Accuracy (public set)14.811.317.019.522.221.024.220.723.325.020.026.719.013.5
AA-Omniscience Index (public set)-32.8-64.2-62.8-45.7-46.2-12.5-42.3-47.3-35.2-24.0-28.8-36.5-5.8–
MMLU-Pro CoT (EN)80.070.176.578.384.684.381.784.580.880.480.982.785.043.0
MMLU-ProX CoT (DE)75.565.770.761.081.781.979.481.177.270.774.979.782.437.3
Math
Average (EN)96.581.788.888.890.187.886.387.490.781.482.891.197.80.6
Average (DE)88.874.345.284.379.683.783.788.190.875.480.686.596.70.1
AIME 2025 (EN)96.981.989.489.688.184.684.287.390.679.881.991.797.90.6
AIME 2025 (DE)87.573.543.884.476.782.980.688.190.672.380.685.696.50.2
AIME 2026 (EN)96.081.588.387.992.191.088.587.590.883.183.890.497.70.6
AIME 2026 (DE)90.075.246.784.282.584.486.788.191.078.580.687.596.90.0
Agentic
Average (EN)63.441.658.946.463.462.146.354.654.040.753.554.966.7–
TerminalBench 2.127.7–20.29.739.7–8.6–29.221.0–39.776.8–
Tau2-Bench (Telecom)94.767.595.945.997.799.143.945.373.141.553.868.182.510.8
Tau2-Bench (Retail)69.958.557.964.970.871.660.871.360.562.961.467.568.79.6
Tau2-Bench (Airline)76.758.768.752.776.070.765.373.372.740.070.772.783.340.0
Tau3-Bench (Banking)38.15.77.25.711.310.65.416.014.75.76.415.550.02.1
BFCL v3 (multi-turn)39.822.858.247.954.053.551.453.445.636.261.644.642.50.6
BFCL v4 (overall)61.436.465.461.570.567.251.068.257.358.067.261.073.2–
BFCL v4 (non-live AST)79.178.183.385.085.888.283.683.735.883.685.545.085.3–
BFCL v4 (live)78.973.778.378.880.281.482.580.270.478.478.277.679.9–
BFCL v4 (multi-turn)47.527.562.753.559.958.156.061.455.440.465.251.755.5–
BFCL v4 (memory)62.819.441.539.162.653.835.352.950.739.143.459.679.6–
BFCL v4 (web search)62.510.569.066.075.068.512.575.057.069.071.071.582.0–
BrowseComp29.44.4–14.536.526.92.825.531.2––29.146.4–
Code
Average (EN)89.368.067.881.885.087.783.689.090.882.079.888.394.225.1
LiveCodeBench v685.959.246.571.377.882.573.982.387.571.267.882.093.88.7
HumanEval+92.776.889.092.492.292.893.395.794.192.891.894.794.741.6
SWE-Bench Verified66.4–51.038.671.673.8–57.8–60.811.660.272.6–
Instruction Following
Average (EN)78.162.564.573.272.766.160.779.971.149.838.273.781.925.5
IFBench (loose-prompt)78.162.564.573.272.766.160.779.971.149.838.273.781.925.5
Grounding / Hallucinations
Average (EN)59.442.757.655.967.568.361.062.562.661.563.864.467.531.5
SQuAD (M/A Grounding Score)23.40.01.50.023.012.10.00.00.06.30.00.08.30.0
SQuAD (Utility Accuracy)83.477.482.775.089.888.973.588.474.276.881.980.885.61.6
RGB Closed-Book51.052.078.080.081.079.089.079.085.086.092.093.073.085.0
RGB Negative (Abstention)85.673.979.676.986.679.681.386.079.682.389.074.670.655.9
AA-Omniscience Non-Hallucination Rate (1 βˆ’ Hallucination Rate, public set)44.015.03.819.011.156.712.314.323.734.739.013.967.316.4
RGB Fact-Check (Error Correction)34.014.068.058.074.074.085.061.077.060.074.090.053.014.0
FRAMES (<24k)71.265.767.870.975.974.768.969.578.371.970.474.978.658.9
FRAMES (>24k)73.0–73.974.381.578.868.576.678.476.6–78.883.3–
SealQA (no distractors, <24k)80.756.671.769.088.382.882.182.180.773.167.682.886.935.9
SealQA (12 distractors, <24k)61.030.065.054.078.067.057.082.065.062.060.070.084.016.0
SealQA (no distractors, >24k)100.0–61.166.7100.088.983.377.883.377.8–94.494.4–
SealQA (12 distractors, >24k)65.1–57.139.773.071.450.868.363.549.2–68.382.5–
Agentic Retrieval
Average (EN)77.342.758.871.578.961.250.571.276.166.869.379.183.714.0
Average (DE)69.423.749.065.468.458.861.666.267.765.462.967.973.526.5
MuSiQue (EN)77.342.758.871.578.961.250.571.276.166.869.379.183.714.0
Honeypot80.825.369.460.277.574.313.575.466.668.1–68.885.0–
Agentic Wiki QA (DE)69.423.749.065.468.458.861.666.267.765.462.967.973.526.5
Industry RAG
Average (EN)89.753.875.961.387.086.062.779.482.874.982.380.393.323.9
Average (DE)67.542.760.946.270.065.831.149.464.253.463.457.680.214.4
Semiconductors80.435.364.739.279.479.441.263.770.662.772.569.689.211.8
German Public Sector75.054.069.565.580.072.029.566.577.050.069.578.089.08.0
Aerospace58.914.150.444.162.359.048.158.148.147.0–54.973.8–
Automotive Supplier99.072.487.283.394.692.684.295.195.087.192.091.097.335.9
Industrial Drive Technology60.031.452.326.860.059.532.732.351.456.857.337.371.420.9
Long Context
LongBench Pro64.5––53.270.370.863.864.3–56.4–62.976.9–
AA-LCR68.3–49.742.366.369.746.368.3–52.3–67.081.3–

All models use the same evaluation setup: eval-framework for most benchmarks and Harbor for TerminalBench and SWE-Bench. Each model uses its documented context window and sampling parameters, with Kolibri at reasoning effort high. A model reports no score (–) when it cannot call tools or when a prompt or agent trajectory exceeds its window; long-context benchmarks instead score an overlength prompt as 0.

Category averages are unweighted means over rows scored by every compared model except Apertus. They also exclude AA-Omniscience Index, Honeypot and the individual BFCL v4 splits. Excluded scores are greyed out. Each Overall is the unweighted mean of that language's category averages.

Pre-training

All models in the table below are pre-trained base models. Kolibri Base is the checkpoint Kolibri's post-training starts from.

TypeMoEDense
Active parameters3-4B12B7B32B70B
OursBaseline models
Eval
Kolibri Base
Kolibri Origin Base
Gemma 4 26B-A4B Base
Nemotron 3 Nano 30B-A3B Base
Qwen3.5 35B-A3B Base
GLM-4.5 Air 106B-A12B Base
Nemotron 3 Super 120B-A12B Base
OLMo 3 7B Base
OLMo 3 32B Base
Apertus 70B Base
Overall (EN)81.158.458.177.573.877.183.157.667.948.6
Overall (DE)81.561.161.176.076.679.485.048.065.250.5
General Knowledge
Average (EN)81.370.877.280.582.782.587.267.877.873.8
Average (DE)82.770.780.580.084.981.787.551.868.774.7
MMLU (EN)81.068.978.279.084.782.886.867.176.369.3
Global MMLU (DE)77.565.175.174.381.578.284.751.665.164.9
ARC (EN)95.788.995.294.397.096.397.588.694.590.7
ARC (DE)96.488.595.594.297.695.997.867.988.189.9
PIQA (EN)89.576.988.089.491.989.895.177.786.880.1
PIQA (DE)96.887.996.896.598.597.299.375.090.094.6
HellaSwag (EN)83.180.784.885.485.387.288.876.183.584.4
HellaSwag (DE)86.975.388.685.888.786.593.036.861.088.7
MMLU-Pro (EN)61.139.351.255.662.655.368.938.050.940.6
MMLU-ProX (DE)55.736.546.749.058.450.862.827.939.535.6
TriviaQA (EN)77.269.965.679.574.683.886.459.074.777.4
Wahl-O-Mat (DE)16.752.357.155.943.256.761.149.656.357.0
Math
Average (EN)84.954.345.983.073.564.782.757.763.339.6
Average (DE)76.855.847.072.577.569.881.243.862.939.0
GSM8K (EN)89.871.665.786.389.582.887.575.381.162.3
GSM8K Platinum (DE)90.571.764.186.388.086.593.158.780.860.2
MATH Minerva (EN)80.137.026.179.857.646.577.840.145.616.9
MATH Minerva (DE)63.239.929.958.867.053.269.328.945.017.7
Code
Average (EN)77.050.151.169.065.284.179.347.462.532.5
Average (DE)85.157.055.775.567.486.786.248.364.037.9
HumanEval (EN)86.750.852.273.567.296.383.246.863.928.3
HumanEval (DE)88.449.051.773.559.294.085.039.556.228.3
MBPP (EN)67.349.550.064.563.171.975.448.061.036.8
MBPP (DE)81.764.959.777.475.679.487.557.171.747.5

All results in the table above are produced with the same evaluation setup for every model, based on our eval-framework; this includes identical prompts, few-shot configurations and task settings. All models use the sampling parameters temperature = 0.6, top_p = 0.6, max_tokens = 1024 and a maximum context length of 65,536 tokens, except Gemma 4 26B-A4B Base, which uses Google's recommended temperature = 1.0, top_p = 0.95. Each model is served with its long-context extension on: Qwen3.5 35B-A3B Base with static YaRN (factor 4).

Each group score is the unweighted mean of the evals in that group: Math (EN), for example, is the mean of GSM8K (EN) and MATH Minerva (EN). The Overall score is the unweighted mean of the three group scores, so each capability (General Knowledge, Math, Code) contributes equally regardless of how many evals it contains. The aggregates are meant for comparing models within one capability and language, not a model's English against its German scores: an eval is not necessarily equally difficult in both languages, and the groups are not composed identically. General Knowledge (EN) includes MMLU-Pro and TriviaQA while General Knowledge (DE) includes neither, so it and, by extension, Overall (EN) average over two additional and comparatively hard evals.

Long context

TypeMoEDense
Active parameters3-4B7B32B70B
OursBaseline models
Eval
Kolibri Base
Kolibri Origin Base
Gemma 4 26B-A4B Base
Nemotron 3 Nano 30B-A3B Base
Qwen3.5 35B-A3B Base
OLMo 3 7B Base
OLMo 3 32B Base
Apertus 70B Base
RULER
4k86.991.884.794.696.392.094.889.5
8k83.788.885.093.495.380.892.677.7
16k80.983.787.192.095.270.888.871.3
32k76.372.788.786.493.764.480.970.8
64k72.2–86.882.691.0–––
128k67.9–87.881.089.9–––
256k69.8––72.180.1–

Truncated β€” view the full README on Hugging Face.

conversational
kolibri1
reasoning
safetensors
text-generation
vllm

Aleph-Alpha/Kolibri-1

Model

Kolibri

94

3 commits

3 linked in READMEs

updated Oct 3, 2026

See the code

README

Kolibri

Tech report | Tech blog

Kolibri is Aleph Alpha's mixture-of-experts (MoE) reasoning model, with a focus on German and English. The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency.

Model overview

ModelKolibri 1
Model ProviderAleph Alpha GmbH
Model DeveloperAleph Alpha Research GmbH
ArchitectureMixture-of-Experts
Total parameters78B (78,103,074,560)
Active parameters / token3.46B (3,457,573,120)
LanguagesGerman, English
Context length1,048,576 tokens; we recommend ≀262,144 tokens for serving efficiency and complex tasks
Precisionfloat8_e4m3fn weights in 128Γ—128 blocks with dynamically quantized activations, evaluated with an FP8 KV cache; embeddings, LM head, norms and MoE router in bfloat16
Reasoning modeYes
Tool callingYes
LicenseApache 2.0
Knowledge cutoffEN: June 18, 2026, DE: June 18, 2026
This only affects implicit knowledge, the model may use more recent information through tool use.
Hardware requirementsModel memory footprint: ~78 GB (FP8 weights). Minimum: 2Γ— A100 80 GB, 2Γ— H100 SXM5, 1Γ— H200, 1Γ— B200 or 1Γ— B300. Recommended: 2Γ— H100 SXM5, 2Γ— H200, 1Γ— B200 or 1Γ— B300.
Best forMulti-step reasoning, retrieval-augmented generation, agentic tool calling, coding, German- and English-language assistant
Release Date3rd of October 2026
Code of PracticeAleph Alpha is a signatory of the EU GPAI Code of Practice, see https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai.
Training DataPre-training: Trained on 20T tokens of a filtered, bilingual corpus (~62.5% English, ~23.9% German, ~13.6% code) combining curated web data, synthetic rephrasings and translations, and high-quality sources. Additionally trained on 3.44T in mid-training and 201B for long-context extension.
Post-training: The SFT mix contained filtered, bilingual data that combines open-source datasets and synthetically generated data. For RL we used a broad mix of environments that cover reasoning, agentic, and instruction following use-cases.
Training methodWe trained a transformer 50-layer MoE model with 4:1 SWA:GQA attention, using Muon and Exact Quantile Balancing on 384 experts per layer, with 1 shared and 6 routed.
Computing ResourcesPre-training (based on actual measurements), excluding mid-training and long-context: Hardware: 768 NVIDIA B200 (96 HGX 8xB200 nodes); Parallelism: EP8 FSDP16 DP6; Time: 21 days (511h, 392k GPUh)
Mid-training: 5 days, 90k GPUh (same setup as above)
Long-Context: 13h, 10k GPUh (same setup as above except parallelism: FSDP128 DP6)
FLOPS: 6.4e23
Tech Reporthttps://aleph-alpha.com/downloads/tech-report.pdf

Intended use

Kolibri is intended to process text input and output in German and English and to perform a wide range of tasks beyond natural-language generation, for example multi-step reasoning, coding, structured extraction, retrieval-augmented generation, long-document processing and agentic tool calling.

Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained on 65,536 and trained on 262,144 tokens in a final long-context phase, which is its native context length. Because positional encoding is applied only in the sliding-window layers, the context can be extended beyond that length without any position scaling, in principle to arbitrary lengths. We have validated quality and serving efficiency up to 1,048,576 tokens. For latency- or throughput-sensitive deployments and for complex tasks, we recommend contexts of at most 262,144 tokens. See the technical report for details.

For the kinds of systems Kolibri is meant to be integrated into, see AI system types below; for uses that we encourage users to refrain from, see Responsible Use.

Design goals

Kolibri is designed to deliver strong German and English performance at low serving cost. Its mixture-of-experts architecture activates only a small fraction of its parameters for each token, keeping compute per token low while retaining the capacity of a much larger model. The trade-off is memory: the full model must be held in memory even though only part of it is active at any time. To keep long contexts affordable, most attention layers focus on nearby text, while a smaller number attend across the whole context. We also developed a tokenizer tailored to German word structure, so German text is processed efficiently without sacrificing English. Supporting two languages rather than many is a deliberate choice of depth over breadth.

AI system types

Kolibri is intended for integration into conversational assistants and agentic workflows in German and English, in which a person reviews the model's output before it is acted on rather than autonomous systems that act unreviewed. It suits document-processing and drafting systems, question-answering systems over an organisation's own material, and internal knowledge and research tools. Its tool-calling and structured-output capabilities make it appropriate for orchestration layers that call APIs, execute code or run searches, provided the calling system validates the results. In decision-support systems it belongs on the advisory side, surfacing evidence and drafting options for a human to weigh, and it is not intended as the deciding component. More broadly, it is built for human-AI collaboration rather than unsupervised operation.

Sustainability

Energy consumption

9.5Γ—10Β² MWh (estimated), including node power and data-centre overhead (PUE). This includes pre-training, mid-training and long-context training. It excludes SFT and RL, peak, idle and low-load states, and proxy and ablation models.

Energy measurement

For pre-training, we obtained the average power usage per node and the power usage effectiveness (PUE) from the cluster provider. We multiplied these with the number of nodes and the runtime to obtain a total energy estimate.

Getting started

Kolibri requires the aleph-alpha-inference package that provides the Kolibri vLLM plugin. You can either use the provided container image ghcr.io/aleph-alpha/aleph-alpha-inference, or install the package from Aleph-Alpha/aleph-alpha-inference, which also installs the vLLM version it supports:

pip install 'aleph-alpha-inference>=1'

Serve the model with reasoning and tool-calling enabled:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'.

The recommended sampling parameters for the model are temperature=1.0, top_p=0.97 and top_k=128.

Querying the server (OpenAI-compatible API)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="Aleph-Alpha/Kolibri-1",
    messages=[
        {
            "role": "user",
            "content": "ErklΓ€re kurz, was ein Mixture-of-Experts-Modell ist.",
        },
    ],
    extra_body={
        "chat_template_kwargs": {
            "reasoning_effort": "high",
            "enable_thinking": True,
        }
    },
)
print(response.choices[0].message.content)

Reasoning mode

Kolibri supports explicit thinking effort levels that need to be configured through the chat template. You can pass reasoning_effort values low, medium and high to configure the amount of effort our model puts into finding the answer. You can disable thinking altogether by setting reasoning_effort to none or passing enable_thinking=false. The model will then immediately respond.

Tool calling

The serving command enables Hermes-style tool calling (--tool-call-parser kolibri1 --enable-auto-tool-choice). Pass your function schemas via the standard tools field of the chat completions request; the model will emit tool calls that the parser converts into structured output. Tool calling can be combined with reasoning mode.

import json
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
MODEL = "Aleph-Alpha/Kolibri-1"

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {
                        "type": "string",
                        "description": "City name, e.g. Heidelberg",
                    },
                    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
                },
                "required": ["city"],
            },
        },
    }
]


def get_weather(city, unit="celsius"):
    return {"city": city, "temperature": 18, "unit": unit, "conditions": "cloudy"}


messages = [{"role": "user", "content": "Wie ist das Wetter gerade in Heidelberg?"}]

first = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
msg = first.choices[0].message
messages.append(msg)

for call in msg.tool_calls or []:
    args = json.loads(call.function.arguments)
    result = get_weather(**args)
    messages.append(
        {
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result),
        }
    )

final = client.chat.completions.create(model=MODEL, messages=messages, tools=tools)
print(final.choices[0].message.content)

Evaluation

The best value per row is bolded. In each table, the best value of each group with more than one model is underlined, and both marks compare the MoE models only. The dense models activate several times as many parameters per token, so they are greyed out and unmarked.

Post-training

TypeMoEDense
Active parameters3B4-6B12B27B70B
OursBaseline models
Eval
Kolibri
Kolibri Origin
GLM-4.7 Flash 30B-A3B
Nemotron 3 Nano 30B-A3B
Qwen3.5 35B-A3B
Qwen3.6 35B-A3B
Qwen3-Next 80B-A3B Thinking
Gemma 4 26B-A4B IT
GPT-OSS 120B
Mistral Small 4 119B-A6B
GLM-4.5 Air 106B-A12B
Nemotron 3 Super 120B-A12B
Qwen3.8 27B
Apertus 70B Instruct
Overall (EN)75.554.164.765.674.771.462.471.972.363.164.473.080.2–
Overall (DE)70.846.450.459.369.867.358.066.370.261.464.867.979.9–
Knowledge
Average (EN)50.139.745.546.052.752.148.451.450.047.545.752.056.822.8
Average (DE)57.644.946.541.261.361.055.761.558.051.452.259.569.224.8
GPQA Diamond (EN)84.368.173.173.983.883.476.181.176.474.773.278.089.229.5
GPQA Diamond (DE)81.358.559.849.684.280.672.280.176.072.971.176.688.131.4
Humanity's Last Exam (EN)21.59.415.412.120.421.111.619.219.49.78.720.635.65.2
Humanity's Last Exam (DE)15.910.49.113.118.120.515.623.420.710.510.522.337.25.7
AA-Omniscience Accuracy (public set)14.811.317.019.522.221.024.220.723.325.020.026.719.013.5
AA-Omniscience Index (public set)-32.8-64.2-62.8-45.7-46.2-12.5-42.3-47.3-35.2-24.0-28.8-36.5-5.8–
MMLU-Pro CoT (EN)80.070.176.578.384.684.381.784.580.880.480.982.785.043.0
MMLU-ProX CoT (DE)75.565.770.761.081.781.979.481.177.270.774.979.782.437.3
Math
Average (EN)96.581.788.888.890.187.886.387.490.781.482.891.197.80.6
Average (DE)88.874.345.284.379.683.783.788.190.875.480.686.596.70.1
AIME 2025 (EN)96.981.989.489.688.184.684.287.390.679.881.991.797.90.6
AIME 2025 (DE)87.573.543.884.476.782.980.688.190.672.380.685.696.50.2
AIME 2026 (EN)96.081.588.387.992.191.088.587.590.883.183.890.497.70.6
AIME 2026 (DE)90.075.246.784.282.584.486.788.191.078.580.687.596.90.0
Agentic
Average (EN)63.441.658.946.463.462.146.354.654.040.753.554.966.7–
TerminalBench 2.127.7–20.29.739.7–8.6–29.221.0–39.776.8–
Tau2-Bench (Telecom)94.767.595.945.997.799.143.945.373.141.553.868.182.510.8
Tau2-Bench (Retail)69.958.557.964.970.871.660.871.360.562.961.467.568.79.6
Tau2-Bench (Airline)76.758.768.752.776.070.765.373.372.740.070.772.783.340.0
Tau3-Bench (Banking)38.15.77.25.711.310.65.416.014.75.76.415.550.02.1
BFCL v3 (multi-turn)39.822.858.247.954.053.551.453.445.636.261.644.642.50.6
BFCL v4 (overall)61.436.465.461.570.567.251.068.257.358.067.261.073.2–
BFCL v4 (non-live AST)79.178.183.385.085.888.283.683.735.883.685.545.085.3–
BFCL v4 (live)78.973.778.378.880.281.482.580.270.478.478.277.679.9–
BFCL v4 (multi-turn)47.527.562.753.559.958.156.061.455.440.465.251.755.5–
BFCL v4 (memory)62.819.441.539.162.653.835.352.950.739.143.459.679.6–
BFCL v4 (web search)62.510.569.066.075.068.512.575.057.069.071.071.582.0–
BrowseComp29.44.4–14.536.526.92.825.531.2––29.146.4–
Code
Average (EN)89.368.067.881.885.087.783.689.090.882.079.888.394.225.1
LiveCodeBench v685.959.246.571.377.882.573.982.387.571.267.882.093.88.7
HumanEval+92.776.889.092.492.292.893.395.794.192.891.894.794.741.6
SWE-Bench Verified66.4–51.038.671.673.8–57.8–60.811.660.272.6–
Instruction Following
Average (EN)78.162.564.573.272.766.160.779.971.149.838.273.781.925.5
IFBench (loose-prompt)78.162.564.573.272.766.160.779.971.149.838.273.781.925.5
Grounding / Hallucinations
Average (EN)59.442.757.655.967.568.361.062.562.661.563.864.467.531.5
SQuAD (M/A Grounding Score)23.40.01.50.023.012.10.00.00.06.30.00.08.30.0
SQuAD (Utility Accuracy)83.477.482.775.089.888.973.588.474.276.881.980.885.61.6
RGB Closed-Book51.052.078.080.081.079.089.079.085.086.092.093.073.085.0
RGB Negative (Abstention)85.673.979.676.986.679.681.386.079.682.389.074.670.655.9
AA-Omniscience Non-Hallucination Rate (1 βˆ’ Hallucination Rate, public set)44.015.03.819.011.156.712.314.323.734.739.013.967.316.4
RGB Fact-Check (Error Correction)34.014.068.058.074.074.085.061.077.060.074.090.053.014.0
FRAMES (<24k)71.265.767.870.975.974.768.969.578.371.970.474.978.658.9
FRAMES (>24k)73.0–73.974.381.578.868.576.678.476.6–78.883.3–
SealQA (no distractors, <24k)80.756.671.769.088.382.882.182.180.773.167.682.886.935.9
SealQA (12 distractors, <24k)61.030.065.054.078.067.057.082.065.062.060.070.084.016.0
SealQA (no distractors, >24k)100.0–61.166.7100.088.983.377.883.377.8–94.494.4–
SealQA (12 distractors, >24k)65.1–57.139.773.071.450.868.363.549.2–68.382.5–
Agentic Retrieval
Average (EN)77.342.758.871.578.961.250.571.276.166.869.379.183.714.0
Average (DE)69.423.749.065.468.458.861.666.267.765.462.967.973.526.5
MuSiQue (EN)77.342.758.871.578.961.250.571.276.166.869.379.183.714.0
Honeypot80.825.369.460.277.574.313.575.466.668.1–68.885.0–
Agentic Wiki QA (DE)69.423.749.065.468.458.861.666.267.765.462.967.973.526.5
Industry RAG
Average (EN)89.753.875.961.387.086.062.779.482.874.982.380.393.323.9
Average (DE)67.542.760.946.270.065.831.149.464.253.463.457.680.214.4
Semiconductors80.435.364.739.279.479.441.263.770.662.772.569.689.211.8
German Public Sector75.054.069.565.580.072.029.566.577.050.069.578.089.08.0
Aerospace58.914.150.444.162.359.048.158.148.147.0–54.973.8–
Automotive Supplier99.072.487.283.394.692.684.295.195.087.192.091.097.335.9
Industrial Drive Technology60.031.452.326.860.059.532.732.351.456.857.337.371.420.9
Long Context
LongBench Pro64.5––53.270.370.863.864.3–56.4–62.976.9–
AA-LCR68.3–49.742.366.369.746.368.3–52.3–67.081.3–

All models use the same evaluation setup: eval-framework for most benchmarks and Harbor for TerminalBench and SWE-Bench. Each model uses its documented context window and sampling parameters, with Kolibri at reasoning effort high. A model reports no score (–) when it cannot call tools or when a prompt or agent trajectory exceeds its window; long-context benchmarks instead score an overlength prompt as 0.

Category averages are unweighted means over rows scored by every compared model except Apertus. They also exclude AA-Omniscience Index, Honeypot and the individual BFCL v4 splits. Excluded scores are greyed out. Each Overall is the unweighted mean of that language's category averages.

Pre-training

All models in the table below are pre-trained base models. Kolibri Base is the checkpoint Kolibri's post-training starts from.

TypeMoEDense
Active parameters3-4B12B7B32B70B
OursBaseline models
Eval
Kolibri Base
Kolibri Origin Base
Gemma 4 26B-A4B Base
Nemotron 3 Nano 30B-A3B Base
Qwen3.5 35B-A3B Base
GLM-4.5 Air 106B-A12B Base
Nemotron 3 Super 120B-A12B Base
OLMo 3 7B Base
OLMo 3 32B Base
Apertus 70B Base
Overall (EN)81.158.458.177.573.877.183.157.667.948.6
Overall (DE)81.561.161.176.076.679.485.048.065.250.5
General Knowledge
Average (EN)81.370.877.280.582.782.587.267.877.873.8
Average (DE)82.770.780.580.084.981.787.551.868.774.7
MMLU (EN)81.068.978.279.084.782.886.867.176.369.3
Global MMLU (DE)77.565.175.174.381.578.284.751.665.164.9
ARC (EN)95.788.995.294.397.096.397.588.694.590.7
ARC (DE)96.488.595.594.297.695.997.867.988.189.9
PIQA (EN)89.576.988.089.491.989.895.177.786.880.1
PIQA (DE)96.887.996.896.598.597.299.375.090.094.6
HellaSwag (EN)83.180.784.885.485.387.288.876.183.584.4
HellaSwag (DE)86.975.388.685.888.786.593.036.861.088.7
MMLU-Pro (EN)61.139.351.255.662.655.368.938.050.940.6
MMLU-ProX (DE)55.736.546.749.058.450.862.827.939.535.6
TriviaQA (EN)77.269.965.679.574.683.886.459.074.777.4
Wahl-O-Mat (DE)16.752.357.155.943.256.761.149.656.357.0
Math
Average (EN)84.954.345.983.073.564.782.757.763.339.6
Average (DE)76.855.847.072.577.569.881.243.862.939.0
GSM8K (EN)89.871.665.786.389.582.887.575.381.162.3
GSM8K Platinum (DE)90.571.764.186.388.086.593.158.780.860.2
MATH Minerva (EN)80.137.026.179.857.646.577.840.145.616.9
MATH Minerva (DE)63.239.929.958.867.053.269.328.945.017.7
Code
Average (EN)77.050.151.169.065.284.179.347.462.532.5
Average (DE)85.157.055.775.567.486.786.248.364.037.9
HumanEval (EN)86.750.852.273.567.296.383.246.863.928.3
HumanEval (DE)88.449.051.773.559.294.085.039.556.228.3
MBPP (EN)67.349.550.064.563.171.975.448.061.036.8
MBPP (DE)81.764.959.777.475.679.487.557.171.747.5

All results in the table above are produced with the same evaluation setup for every model, based on our eval-framework; this includes identical prompts, few-shot configurations and task settings. All models use the sampling parameters temperature = 0.6, top_p = 0.6, max_tokens = 1024 and a maximum context length of 65,536 tokens, except Gemma 4 26B-A4B Base, which uses Google's recommended temperature = 1.0, top_p = 0.95. Each model is served with its long-context extension on: Qwen3.5 35B-A3B Base with static YaRN (factor 4).

Each group score is the unweighted mean of the evals in that group: Math (EN), for example, is the mean of GSM8K (EN) and MATH Minerva (EN). The Overall score is the unweighted mean of the three group scores, so each capability (General Knowledge, Math, Code) contributes equally regardless of how many evals it contains. The aggregates are meant for comparing models within one capability and language, not a model's English against its German scores: an eval is not necessarily equally difficult in both languages, and the groups are not composed identically. General Knowledge (EN) includes MMLU-Pro and TriviaQA while General Knowledge (DE) includes neither, so it and, by extension, Overall (EN) average over two additional and comparatively hard evals.

Long context

TypeMoEDense
Active parameters3-4B7B32B70B
OursBaseline models
Eval
Kolibri Base
Kolibri Origin Base
Gemma 4 26B-A4B Base
Nemotron 3 Nano 30B-A3B Base
Qwen3.5 35B-A3B Base
OLMo 3 7B Base
OLMo 3 32B Base
Apertus 70B Base
RULER
4k86.991.884.794.696.392.094.889.5
8k83.788.885.093.495.380.892.677.7
16k80.983.787.192.095.270.888.871.3
32k76.372.788.786.493.764.480.970.8
64k72.2–86.882.691.0–––
128k67.9–87.881.089.9–––
256k69.8––72.180.1–

Truncated β€” view the full README on Hugging Face.

conversational
kolibri1
reasoning
safetensors
text-generation
vllm