Aleph-Alpha/aleph-alpha-inference

Python

6

9 commits

updated Oct 3, 2026

See the code

README

aleph-alpha-inference

A vLLM plugin that serves Aleph Alpha models. It adds:

ModelArchitectureReasoning parserTool call parser
Kolibri 1Kolibri1ForCausalLMkolibri1kolibri1

vLLM finds the plugin through its vllm.general_plugins entry point, so installing the package is all the setup there is.

Install

Each release supports one vLLM minor version, currently vLLM 0.29.

pip install aleph-alpha-inference

This installs the supported vLLM as well. To add the plugin to an existing vLLM 0.29 installation, such as the vllm/vllm-openai:v0.29.0 image, without touching its vLLM, torch or transformers:

pip install --no-deps aleph-alpha-inference

Serve

vllm serve Aleph-Alpha/Kolibri-1 \
  --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

For the BF16 weights, serve Aleph-Alpha/Kolibri-1-BF16 and drop --kv-cache-dtype fp8.

Thinking is on by default. Turn it off for a request with reasoning_effort: "none" or enable_thinking: false:

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Aleph-Alpha/Kolibri-1",
    "messages": [{"role": "user", "content": "What is the capital of France?"}],
    "chat_template_kwargs": {"enable_thinking": false}
  }'

The kolibri1 reasoning parser reads the same switch as the chat template, so the response's reasoning and content are split correctly whichever way thinking is set.

Development

uv run pytest -m "not gpu"   # runs anywhere
uv run pytest -m gpu         # needs a CUDA GPU supported by vLLM

License

Apache-2.0

Aleph-Alpha/aleph-alpha-inference

Python

6

9 commits

updated Oct 3, 2026

See the code

README

aleph-alpha-inference

A vLLM plugin that serves Aleph Alpha models. It adds:

ModelArchitectureReasoning parserTool call parser
Kolibri 1Kolibri1ForCausalLMkolibri1kolibri1

vLLM finds the plugin through its vllm.general_plugins entry point, so installing the package is all the setup there is.

Install

Each release supports one vLLM minor version, currently vLLM 0.29.

pip install aleph-alpha-inference

This installs the supported vLLM as well. To add the plugin to an existing vLLM 0.29 installation, such as the vllm/vllm-openai:v0.29.0 image, without touching its vLLM, torch or transformers:

pip install --no-deps aleph-alpha-inference

Serve

vllm serve Aleph-Alpha/Kolibri-1 \
  --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

For the BF16 weights, serve Aleph-Alpha/Kolibri-1-BF16 and drop --kv-cache-dtype fp8.

Thinking is on by default. Turn it off for a request with reasoning_effort: "none" or enable_thinking: false:

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Aleph-Alpha/Kolibri-1",
    "messages": [{"role": "user", "content": "What is the capital of France?"}],
    "chat_template_kwargs": {"enable_thinking": false}
  }'

The kolibri1 reasoning parser reads the same switch as the chat template, so the response's reasoning and content are split correctly whichever way thinking is set.

Development

uv run pytest -m "not gpu"   # runs anywhere
uv run pytest -m gpu         # needs a CUDA GPU supported by vLLM

License

Apache-2.0

Languages

Python

85.9%

Jinja

12.9%

Dockerfile

1.3%