speakrail/gemma-4-12B-it-qat-Speakrail

Model

Speakrail Gemma 4 12B

0

7 commits

2 linked in READMEs

updated Oct 5, 2026

See the code

README

Speakrail Gemma 4 12B

A LoRA adapter and 22 trained token embeddings that turn Gemma 4 12B QAT (int4 weights) into a full-duplex turn-taking model. It reads the user's words as speech recognition writes them, plus cues from a turn-taking head, and after every input event it answers with one decision token: speak, keep listening, interrupt, stay quiet, or say a few words without taking the turn (interjection). When it decides to speak, it writes the spoken reply, with tool calls when needed. The spoken reply is intended to be spoken out using a streaming TTS.

It is the brain of Speakrail, a full-duplex voice assistant that runs on a single RTX 4090 next to its speech recognizer (Voxtral Mini 4B Realtime + turn head) and its TTS.

The model is unlikely to work without the Speakrail harness or with a different STT, because it was tuned for Voxtral-specific misses and the model expects to see tokens that ordinary text-based use doesn't have. The model protocol is described below.

How it works

The context is a token stream. The harness appends events as they happen, and the model is asked for a decision at each one:

<system prompt, tools, date>
... user words as they arrive ... <complete>           <- the turn head says the user finished
β†’ decision: <turn|> (speak)  β†’ spoken reply, streamed to TTS
... user words while the assistant speaks ...          <- an overlap
β†’ decision: <continue> | <yield> | <listen>
<sil:4s>                                               <- nobody has spoken for 4 s
β†’ decision: <listen>

Decisions are single tokens, constrained to the allowed set, so a decision costs one forward pass (about 30-60 ms on an RTX 4090).

Tokens

The 22 new tokens reuse Gemma's unused slots (ids 6-27), so the vocabulary size is unchanged. "Speak" is Gemma's own end-of-turn token <turn|> (id 106).

idtokenkindmeaning
6<complete>inputthe user finished their turn (turn head or silence fallback)
7-15<sil:1s> ... <sil:256s>inputsilence ladder: 1, 2, 4 ... 256 s with nobody speaking
25<user_bc>inputthe user made a backchannel ("mm-hmm", "yeah")
16<listen>decisionstay quiet, keep listening
17<listen_muted>decisionstay silent: the user asked the assistant to wait
18<interrupt>decisioncut in while the user is still talking
19<continue>decision (overlap)that was only a backchannel, keep talking
20<yield>decision (overlap)stop talking, the user is taking the floor
26<interject>decisionsay 1-4 words without taking the turn (e.g. counting reps)
27</interject>replycloses the interjected words
21-24<bc>, <bc_mm>, <bc_yeah>, <bc_laugh>reservednot trained, never used

Decision sets:

  • Assistant silent: <turn|>, <interrupt>, <listen>, <listen_muted>, <interject>.
  • Assistant speaking: <continue>, <yield>, <listen>.

Files

filewhat
adapter_model.safetensors, adapter_config.jsonPEFT LoRA (rank 32, alpha 64, all attention and MLP projections), 251 MB
token_rows.safetensorsthe 22 trained embedding rows (ids 6-27, 3840 wide)
tokenizer/tokenizer v1.2 (Gemma 4 + the 22 tokens) and the chat template

Google's base weights are not uploaded here, only the adapter.

Usage

git clone https://github.com/speakrail/speakrail && cd speakrail
cp .env.example .env
docker compose up -d

The model-init step downloads the base model, writes the token rows into it and starts vLLM with the adapter.

Results

These measure the whole Speakrail system on one RTX 4090: Voxtral Mini 4B Realtime + turn head, this model with listening notes (the base model takes notes while the user speaks), and Breeze TTS 2.

Full-Duplex-Bench v3: tool use on real, disfluent speech. The benchmark's 100 recordings are played live into the system and scored with its own evaluators.

SpeakrailTML-Interaction-SmallGPT-Realtime 1.5GPT-Live-1 + gpt-6-astraGemini Live 3.1NVIDIA frontend-backend (Qwen3-235B, ext. ASR)Cascaded Whisper + GPT-4o + TTSNVIDIA frontend-backend (Qwen3-30B, ext. ASR)Ultravox Realtime v0.7Gander 9B (Tencent)NemotronLabs VoiceChat 11B
OpenOpen weights + codeNoNoNoNoNo (paper only)NoNo (paper only)Open weightsOpen weights + codeOpen weights
Pass@1.520.680.600.550.540.480.450.440.410.400.330
Tool F1.915–.876.835.817.717.803.746.794.759.825
Arg acc.633–.680.643.588.552.562.528.513.503.422
Reply.755.828.792.717.718.670.600.540.510.490–
Turn-take98%–96%99%78%100%100%100%96%100%–
Interrupt ↓22.4%–13.5%14.1%19.2%51.0%33.0%54.0%47.9%8.0%–
Filler100%–16.9%97.6%31.7%83.3%26.9%84.8%88.0%51.6%–
1st word ↓1.82 s–6.36 s2.85 s3.95 s–8.78 s–3.88 s––
Tool call ↓2.39 s–3.89 s4.12 s2.21 s–3.15 s–6.01 s––
Done ↓3.92 s–6.89 s8.47 s4.25 s–10.12 s–8.40 s––
Srcourseaoursababacd

Sources:

  • a: the Full-Duplex-Bench v3 paper.
  • b: NVIDIA's frontend-backend paper. It scores the text output, and its Tool F1 column is tool accuracy.
  • c: Tencent's Gander paper, using the benchmark's scripts on the ASR of its synthesized speech.
  • d: NVIDIA's model card, self-reported.
  • e: the TML blog, self-reported, not reproduced under these conditions.
  • ours: our runs on the same 100 recordings with the benchmark's evaluators.

How to read the Speakrail column (bold = best value per metric):

  • Pass@1 by difficulty: easy 0.58, medium 0.53, hard 0.43. By number of tools needed: 1 tool 0.59, 2 tools 0.44, 3 tools 0.31.
  • Latencies are means, as in the paper. The medians are lower: first word 0.8 s, tool call 1.2 s, task done 2.8 s.
  • Speakrail runs locally; the cloud systems' latencies include network round trips.
  • Filler: the first word is usually a short acknowledgement ("Checking on that.", "On it.") before a tool result, so the filler rate is 100 % by design.
  • Turn-based text models that don't listen live score higher on the same scenarios (for example Gemini 3.5 Flash with thinking, 0.65).

Training for short, plain spoken replies did cost some text instruction following and intelligence. Concrete degradation benchmarks will be released later.

Training

  • Base: google/gemma-4-12B-it-qat-w4a16-ct, trained directly on the QAT int4 weights.
  • Trainable parameters: LoRA rank 32 / alpha 64 on q, k, v, o, gate, up and down projections, plus the 22 new embedding rows with their own learning rate (1e-3; LoRA 1e-4).
  • Loss: cross-entropy on the decision tokens and on the reply tokens, plus a KL term to the base model's outputs (Ξ² 0.5) to keep its general behaviour.
  • Compute: one epoch on a single H100, 7-8 hours.

The training data is a mix of scripts, generated by different models (GLM-5.3 and its variants, for example).

Limitations

  • English only.
  • Needs the harness. Without the protocol and the decision constraints it is a weaker general chat model than base Gemma.
  • Weaker text instruction following than base Gemma: it prefers short spoken answers without markdown, lists or symbols.
  • Facts: it can be confidently wrong like any 12B model; Speakrail adds web search and an optional larger model for hard questions.

License

Apache 2.0, same as Gemma 4 (Gemma 4 license).

Citation

@misc{speakrail2026gemma,
  title  = {Speakrail Gemma 4 12B: a full-duplex turn-taking model},
  author = {Speakrail},
  year   = {2026},
  url    = {https://huggingface.co/speakrail/gemma-4-12B-it-qat-Speakrail}
}
full-duplex
gemma4
lora
peft
safetensors
speech-to-speech
text-generation
turn-taking
voice-assistant

speakrail/gemma-4-12B-it-qat-Speakrail

Model

Speakrail Gemma 4 12B

0

7 commits

2 linked in READMEs

updated Oct 5, 2026

See the code

README

Speakrail Gemma 4 12B

A LoRA adapter and 22 trained token embeddings that turn Gemma 4 12B QAT (int4 weights) into a full-duplex turn-taking model. It reads the user's words as speech recognition writes them, plus cues from a turn-taking head, and after every input event it answers with one decision token: speak, keep listening, interrupt, stay quiet, or say a few words without taking the turn (interjection). When it decides to speak, it writes the spoken reply, with tool calls when needed. The spoken reply is intended to be spoken out using a streaming TTS.

It is the brain of Speakrail, a full-duplex voice assistant that runs on a single RTX 4090 next to its speech recognizer (Voxtral Mini 4B Realtime + turn head) and its TTS.

The model is unlikely to work without the Speakrail harness or with a different STT, because it was tuned for Voxtral-specific misses and the model expects to see tokens that ordinary text-based use doesn't have. The model protocol is described below.

How it works

The context is a token stream. The harness appends events as they happen, and the model is asked for a decision at each one:

<system prompt, tools, date>
... user words as they arrive ... <complete>           <- the turn head says the user finished
β†’ decision: <turn|> (speak)  β†’ spoken reply, streamed to TTS
... user words while the assistant speaks ...          <- an overlap
β†’ decision: <continue> | <yield> | <listen>
<sil:4s>                                               <- nobody has spoken for 4 s
β†’ decision: <listen>

Decisions are single tokens, constrained to the allowed set, so a decision costs one forward pass (about 30-60 ms on an RTX 4090).

Tokens

The 22 new tokens reuse Gemma's unused slots (ids 6-27), so the vocabulary size is unchanged. "Speak" is Gemma's own end-of-turn token <turn|> (id 106).

idtokenkindmeaning
6<complete>inputthe user finished their turn (turn head or silence fallback)
7-15<sil:1s> ... <sil:256s>inputsilence ladder: 1, 2, 4 ... 256 s with nobody speaking
25<user_bc>inputthe user made a backchannel ("mm-hmm", "yeah")
16<listen>decisionstay quiet, keep listening
17<listen_muted>decisionstay silent: the user asked the assistant to wait
18<interrupt>decisioncut in while the user is still talking
19<continue>decision (overlap)that was only a backchannel, keep talking
20<yield>decision (overlap)stop talking, the user is taking the floor
26<interject>decisionsay 1-4 words without taking the turn (e.g. counting reps)
27</interject>replycloses the interjected words
21-24<bc>, <bc_mm>, <bc_yeah>, <bc_laugh>reservednot trained, never used

Decision sets:

  • Assistant silent: <turn|>, <interrupt>, <listen>, <listen_muted>, <interject>.
  • Assistant speaking: <continue>, <yield>, <listen>.

Files

filewhat
adapter_model.safetensors, adapter_config.jsonPEFT LoRA (rank 32, alpha 64, all attention and MLP projections), 251 MB
token_rows.safetensorsthe 22 trained embedding rows (ids 6-27, 3840 wide)
tokenizer/tokenizer v1.2 (Gemma 4 + the 22 tokens) and the chat template

Google's base weights are not uploaded here, only the adapter.

Usage

git clone https://github.com/speakrail/speakrail && cd speakrail
cp .env.example .env
docker compose up -d

The model-init step downloads the base model, writes the token rows into it and starts vLLM with the adapter.

Results

These measure the whole Speakrail system on one RTX 4090: Voxtral Mini 4B Realtime + turn head, this model with listening notes (the base model takes notes while the user speaks), and Breeze TTS 2.

Full-Duplex-Bench v3: tool use on real, disfluent speech. The benchmark's 100 recordings are played live into the system and scored with its own evaluators.

SpeakrailTML-Interaction-SmallGPT-Realtime 1.5GPT-Live-1 + gpt-6-astraGemini Live 3.1NVIDIA frontend-backend (Qwen3-235B, ext. ASR)Cascaded Whisper + GPT-4o + TTSNVIDIA frontend-backend (Qwen3-30B, ext. ASR)Ultravox Realtime v0.7Gander 9B (Tencent)NemotronLabs VoiceChat 11B
OpenOpen weights + codeNoNoNoNoNo (paper only)NoNo (paper only)Open weightsOpen weights + codeOpen weights
Pass@1.520.680.600.550.540.480.450.440.410.400.330
Tool F1.915–.876.835.817.717.803.746.794.759.825
Arg acc.633–.680.643.588.552.562.528.513.503.422
Reply.755.828.792.717.718.670.600.540.510.490–
Turn-take98%–96%99%78%100%100%100%96%100%–
Interrupt ↓22.4%–13.5%14.1%19.2%51.0%33.0%54.0%47.9%8.0%–
Filler100%–16.9%97.6%31.7%83.3%26.9%84.8%88.0%51.6%–
1st word ↓1.82 s–6.36 s2.85 s3.95 s–8.78 s–3.88 s––
Tool call ↓2.39 s–3.89 s4.12 s2.21 s–3.15 s–6.01 s––
Done ↓3.92 s–6.89 s8.47 s4.25 s–10.12 s–8.40 s––
Srcourseaoursababacd

Sources:

  • a: the Full-Duplex-Bench v3 paper.
  • b: NVIDIA's frontend-backend paper. It scores the text output, and its Tool F1 column is tool accuracy.
  • c: Tencent's Gander paper, using the benchmark's scripts on the ASR of its synthesized speech.
  • d: NVIDIA's model card, self-reported.
  • e: the TML blog, self-reported, not reproduced under these conditions.
  • ours: our runs on the same 100 recordings with the benchmark's evaluators.

How to read the Speakrail column (bold = best value per metric):

  • Pass@1 by difficulty: easy 0.58, medium 0.53, hard 0.43. By number of tools needed: 1 tool 0.59, 2 tools 0.44, 3 tools 0.31.
  • Latencies are means, as in the paper. The medians are lower: first word 0.8 s, tool call 1.2 s, task done 2.8 s.
  • Speakrail runs locally; the cloud systems' latencies include network round trips.
  • Filler: the first word is usually a short acknowledgement ("Checking on that.", "On it.") before a tool result, so the filler rate is 100 % by design.
  • Turn-based text models that don't listen live score higher on the same scenarios (for example Gemini 3.5 Flash with thinking, 0.65).

Training for short, plain spoken replies did cost some text instruction following and intelligence. Concrete degradation benchmarks will be released later.

Training

  • Base: google/gemma-4-12B-it-qat-w4a16-ct, trained directly on the QAT int4 weights.
  • Trainable parameters: LoRA rank 32 / alpha 64 on q, k, v, o, gate, up and down projections, plus the 22 new embedding rows with their own learning rate (1e-3; LoRA 1e-4).
  • Loss: cross-entropy on the decision tokens and on the reply tokens, plus a KL term to the base model's outputs (Ξ² 0.5) to keep its general behaviour.
  • Compute: one epoch on a single H100, 7-8 hours.

The training data is a mix of scripts, generated by different models (GLM-5.3 and its variants, for example).

Limitations

  • English only.
  • Needs the harness. Without the protocol and the decision constraints it is a weaker general chat model than base Gemma.
  • Weaker text instruction following than base Gemma: it prefers short spoken answers without markdown, lists or symbols.
  • Facts: it can be confidently wrong like any 12B model; Speakrail adds web search and an optional larger model for hard questions.

License

Apache 2.0, same as Gemma 4 (Gemma 4 license).

Citation

@misc{speakrail2026gemma,
  title  = {Speakrail Gemma 4 12B: a full-duplex turn-taking model},
  author = {Speakrail},
  year   = {2026},
  url    = {https://huggingface.co/speakrail/gemma-4-12B-it-qat-Speakrail}
}
full-duplex
gemma4
lora
peft
safetensors
speech-to-speech
text-generation
turn-taking
voice-assistant