A LoRA adapter and 22 trained token embeddings that turn Gemma 4 12B QAT (int4 weights) into a full-duplex turn-taking model. It reads the user's words as speech recognition writes them, plus cues from a turn-taking head, and after every input event it answers with one decision token: speak, keep listening, interrupt, stay quiet, or say a few words without taking the turn (interjection). When it decides to speak, it writes the spoken reply, with tool calls when needed. The spoken reply is intended to be spoken out using a streaming TTS.
It is the brain of Speakrail, a full-duplex voice assistant that runs on a single RTX 4090 next to its speech recognizer (Voxtral Mini 4B Realtime + turn head) and its TTS.
The model is unlikely to work without the Speakrail harness or with a different STT, because it was tuned for Voxtral-specific misses and the model expects to see tokens that ordinary text-based use doesn't have. The model protocol is described below.
The context is a token stream. The harness appends events as they happen, and the model is asked for a decision at each one:
<system prompt, tools, date>
... user words as they arrive ... <complete> <- the turn head says the user finished
β decision: <turn|> (speak) β spoken reply, streamed to TTS
... user words while the assistant speaks ... <- an overlap
β decision: <continue> | <yield> | <listen>
<sil:4s> <- nobody has spoken for 4 s
β decision: <listen>
Decisions are single tokens, constrained to the allowed set, so a decision costs one forward pass (about 30-60 ms on an RTX 4090).
The 22 new tokens reuse Gemma's unused slots (ids 6-27), so the vocabulary size is unchanged. "Speak" is Gemma's own
end-of-turn token <turn|> (id 106).
| id | token | kind | meaning |
|---|---|---|---|
| 6 | <complete> | input | the user finished their turn (turn head or silence fallback) |
| 7-15 | <sil:1s> ... <sil:256s> | input | silence ladder: 1, 2, 4 ... 256 s with nobody speaking |
| 25 | <user_bc> | input | the user made a backchannel ("mm-hmm", "yeah") |
| 16 | <listen> | decision | stay quiet, keep listening |
| 17 | <listen_muted> | decision | stay silent: the user asked the assistant to wait |
| 18 | <interrupt> | decision | cut in while the user is still talking |
| 19 | <continue> | decision (overlap) | that was only a backchannel, keep talking |
| 20 | <yield> | decision (overlap) | stop talking, the user is taking the floor |
| 26 | <interject> | decision | say 1-4 words without taking the turn (e.g. counting reps) |
| 27 | </interject> | reply | closes the interjected words |
| 21-24 | <bc>, <bc_mm>, <bc_yeah>, <bc_laugh> | reserved | not trained, never used |
Decision sets:
<turn|>, <interrupt>, <listen>, <listen_muted>, <interject>.<continue>, <yield>, <listen>.| file | what |
|---|---|
adapter_model.safetensors, adapter_config.json | PEFT LoRA (rank 32, alpha 64, all attention and MLP projections), 251 MB |
token_rows.safetensors | the 22 trained embedding rows (ids 6-27, 3840 wide) |
tokenizer/ | tokenizer v1.2 (Gemma 4 + the 22 tokens) and the chat template |
Google's base weights are not uploaded here, only the adapter.
git clone https://github.com/speakrail/speakrail && cd speakrail
cp .env.example .env
docker compose up -d
The model-init step downloads the base model, writes the token rows into it and starts vLLM with the adapter.
These measure the whole Speakrail system on one RTX 4090: Voxtral Mini 4B Realtime + turn head, this model with listening notes (the base model takes notes while the user speaks), and Breeze TTS 2.
Full-Duplex-Bench v3: tool use on real, disfluent speech. The benchmark's 100 recordings are played live into the system and scored with its own evaluators.
| Speakrail | TML-Interaction-Small | GPT-Realtime 1.5 | GPT-Live-1 + gpt-6-astra | Gemini Live 3.1 | NVIDIA frontend-backend (Qwen3-235B, ext. ASR) | Cascaded Whisper + GPT-4o + TTS | NVIDIA frontend-backend (Qwen3-30B, ext. ASR) | Ultravox Realtime v0.7 | Gander 9B (Tencent) | NemotronLabs VoiceChat 11B | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Open | Open weights + code | No | No | No | No | No (paper only) | No | No (paper only) | Open weights | Open weights + code | Open weights |
| Pass@1 | .520 | .680 | .600 | .550 | .540 | .480 | .450 | .440 | .410 | .400 | .330 |
| Tool F1 | .915 | β | .876 | .835 | .817 | .717 | .803 | .746 | .794 | .759 | .825 |
| Arg acc | .633 | β | .680 | .643 | .588 | .552 | .562 | .528 | .513 | .503 | .422 |
| Reply | .755 | .828 | .792 | .717 | .718 | .670 | .600 | .540 | .510 | .490 | β |
| Turn-take | 98% | β | 96% | 99% | 78% | 100% | 100% | 100% | 96% | 100% | β |
| Interrupt β | 22.4% | β | 13.5% | 14.1% | 19.2% | 51.0% | 33.0% | 54.0% | 47.9% | 8.0% | β |
| Filler | 100% | β | 16.9% | 97.6% | 31.7% | 83.3% | 26.9% | 84.8% | 88.0% | 51.6% | β |
| 1st word β | 1.82 s | β | 6.36 s | 2.85 s | 3.95 s | β | 8.78 s | β | 3.88 s | β | β |
| Tool call β | 2.39 s | β | 3.89 s | 4.12 s | 2.21 s | β | 3.15 s | β | 6.01 s | β | β |
| Done β | 3.92 s | β | 6.89 s | 8.47 s | 4.25 s | β | 10.12 s | β | 8.40 s | β | β |
| Src | ours | e | a | ours | a | b | a | b | a | c | d |
Sources:
How to read the Speakrail column (bold = best value per metric):
Training for short, plain spoken replies did cost some text instruction following and intelligence. Concrete degradation benchmarks will be released later.
google/gemma-4-12B-it-qat-w4a16-ct, trained directly on the QAT int4 weights.The training data is a mix of scripts, generated by different models (GLM-5.3 and its variants, for example).
Apache 2.0, same as Gemma 4 (Gemma 4 license).
@misc{speakrail2026gemma,
title = {Speakrail Gemma 4 12B: a full-duplex turn-taking model},
author = {Speakrail},
year = {2026},
url = {https://huggingface.co/speakrail/gemma-4-12B-it-qat-Speakrail}
}
A LoRA adapter and 22 trained token embeddings that turn Gemma 4 12B QAT (int4 weights) into a full-duplex turn-taking model. It reads the user's words as speech recognition writes them, plus cues from a turn-taking head, and after every input event it answers with one decision token: speak, keep listening, interrupt, stay quiet, or say a few words without taking the turn (interjection). When it decides to speak, it writes the spoken reply, with tool calls when needed. The spoken reply is intended to be spoken out using a streaming TTS.
It is the brain of Speakrail, a full-duplex voice assistant that runs on a single RTX 4090 next to its speech recognizer (Voxtral Mini 4B Realtime + turn head) and its TTS.
The model is unlikely to work without the Speakrail harness or with a different STT, because it was tuned for Voxtral-specific misses and the model expects to see tokens that ordinary text-based use doesn't have. The model protocol is described below.
The context is a token stream. The harness appends events as they happen, and the model is asked for a decision at each one:
<system prompt, tools, date>
... user words as they arrive ... <complete> <- the turn head says the user finished
β decision: <turn|> (speak) β spoken reply, streamed to TTS
... user words while the assistant speaks ... <- an overlap
β decision: <continue> | <yield> | <listen>
<sil:4s> <- nobody has spoken for 4 s
β decision: <listen>
Decisions are single tokens, constrained to the allowed set, so a decision costs one forward pass (about 30-60 ms on an RTX 4090).
The 22 new tokens reuse Gemma's unused slots (ids 6-27), so the vocabulary size is unchanged. "Speak" is Gemma's own
end-of-turn token <turn|> (id 106).
| id | token | kind | meaning |
|---|---|---|---|
| 6 | <complete> | input | the user finished their turn (turn head or silence fallback) |
| 7-15 | <sil:1s> ... <sil:256s> | input | silence ladder: 1, 2, 4 ... 256 s with nobody speaking |
| 25 | <user_bc> | input | the user made a backchannel ("mm-hmm", "yeah") |
| 16 | <listen> | decision | stay quiet, keep listening |
| 17 | <listen_muted> | decision | stay silent: the user asked the assistant to wait |
| 18 | <interrupt> | decision | cut in while the user is still talking |
| 19 | <continue> | decision (overlap) | that was only a backchannel, keep talking |
| 20 | <yield> | decision (overlap) | stop talking, the user is taking the floor |
| 26 | <interject> | decision | say 1-4 words without taking the turn (e.g. counting reps) |
| 27 | </interject> | reply | closes the interjected words |
| 21-24 | <bc>, <bc_mm>, <bc_yeah>, <bc_laugh> | reserved | not trained, never used |
Decision sets:
<turn|>, <interrupt>, <listen>, <listen_muted>, <interject>.<continue>, <yield>, <listen>.| file | what |
|---|---|
adapter_model.safetensors, adapter_config.json | PEFT LoRA (rank 32, alpha 64, all attention and MLP projections), 251 MB |
token_rows.safetensors | the 22 trained embedding rows (ids 6-27, 3840 wide) |
tokenizer/ | tokenizer v1.2 (Gemma 4 + the 22 tokens) and the chat template |
Google's base weights are not uploaded here, only the adapter.
git clone https://github.com/speakrail/speakrail && cd speakrail
cp .env.example .env
docker compose up -d
The model-init step downloads the base model, writes the token rows into it and starts vLLM with the adapter.
These measure the whole Speakrail system on one RTX 4090: Voxtral Mini 4B Realtime + turn head, this model with listening notes (the base model takes notes while the user speaks), and Breeze TTS 2.
Full-Duplex-Bench v3: tool use on real, disfluent speech. The benchmark's 100 recordings are played live into the system and scored with its own evaluators.
| Speakrail | TML-Interaction-Small | GPT-Realtime 1.5 | GPT-Live-1 + gpt-6-astra | Gemini Live 3.1 | NVIDIA frontend-backend (Qwen3-235B, ext. ASR) | Cascaded Whisper + GPT-4o + TTS | NVIDIA frontend-backend (Qwen3-30B, ext. ASR) | Ultravox Realtime v0.7 | Gander 9B (Tencent) | NemotronLabs VoiceChat 11B | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Open | Open weights + code | No | No | No | No | No (paper only) | No | No (paper only) | Open weights | Open weights + code | Open weights |
| Pass@1 | .520 | .680 | .600 | .550 | .540 | .480 | .450 | .440 | .410 | .400 | .330 |
| Tool F1 | .915 | β | .876 | .835 | .817 | .717 | .803 | .746 | .794 | .759 | .825 |
| Arg acc | .633 | β | .680 | .643 | .588 | .552 | .562 | .528 | .513 | .503 | .422 |
| Reply | .755 | .828 | .792 | .717 | .718 | .670 | .600 | .540 | .510 | .490 | β |
| Turn-take | 98% | β | 96% | 99% | 78% | 100% | 100% | 100% | 96% | 100% | β |
| Interrupt β | 22.4% | β | 13.5% | 14.1% | 19.2% | 51.0% | 33.0% | 54.0% | 47.9% | 8.0% | β |
| Filler | 100% | β | 16.9% | 97.6% | 31.7% | 83.3% | 26.9% | 84.8% | 88.0% | 51.6% | β |
| 1st word β | 1.82 s | β | 6.36 s | 2.85 s | 3.95 s | β | 8.78 s | β | 3.88 s | β | β |
| Tool call β | 2.39 s | β | 3.89 s | 4.12 s | 2.21 s | β | 3.15 s | β | 6.01 s | β | β |
| Done β | 3.92 s | β | 6.89 s | 8.47 s | 4.25 s | β | 10.12 s | β | 8.40 s | β | β |
| Src | ours | e | a | ours | a | b | a | b | a | c | d |
Sources:
How to read the Speakrail column (bold = best value per metric):
Training for short, plain spoken replies did cost some text instruction following and intelligence. Concrete degradation benchmarks will be released later.
google/gemma-4-12B-it-qat-w4a16-ct, trained directly on the QAT int4 weights.The training data is a mix of scripts, generated by different models (GLM-5.3 and its variants, for example).
Apache 2.0, same as Gemma 4 (Gemma 4 license).
@misc{speakrail2026gemma,
title = {Speakrail Gemma 4 12B: a full-duplex turn-taking model},
author = {Speakrail},
year = {2026},
url = {https://huggingface.co/speakrail/gemma-4-12B-it-qat-Speakrail}
}