Speakrail Voxtral Mini 4B Realtime Turn Head
0
4 commits
7 linked in READMEs
updated Oct 5, 2026
A turn-taking head for Voxtral Mini 4B Realtime. It reads one hidden layer of the streaming speech recognizer and, every 80 ms, says whether the user is still speaking, has finished their turn, has paused mid-thought, or just made a backchannel ("mm-hmm", "yeah").
The head is a 2.4 M-parameter MLP that runs inside the recognizer's decode step, so the transcript and the turn state come out of the same forward pass at no extra cost (0.02 ms per step on an RTX 4090), and the transcript is bit-identical with and without the head (the model's weights were completely frozen during the training).
The head doesn't take turns by itself. It gives four streaming signals, roughly "is the user talking", "did they finish", "is this a pause mid-thought" and "was that a backchannel", for a dialogue model or a rule set to act on. In Speakrail, a full-duplex voice assistant that runs on a single RTX 4090, the head turns these signals into cue tokens for the LLM, and the LLM makes every turn-taking decision. It can still reply "listen" to a cue.
A VAD / silence timer is not suitable for full-duplex interactions, and on my tests I found the Smart Turn models unreliable. I noticed that the Kyutai STT adds additional heads to the STT to help guide the LLM (the "semantic VAD" of kyutai/stt-1b-en_fr, used in Unmute), and I recently heard a speech about attaching a turn taking head to the GigaAM model, and that motivated me to attach a similar head to Voxtral, since the model is quite large, knows how to punctuate, and, therefore, should be able to understand when to take turns. Similar work was done by X2-Turn (model), but they fully fine-tuned the model (encoder and decoder) on Chinese and English turn data to allow it to turn-take, which I thought was unnecessary, and keeping the weights frozen kept the WER exactly as it is and gave a good-enough quality for me to consider the pipeline.
| file | what |
|---|---|
turn_head.vxth | the head for audio.cpp (format below) |
turn_head.vxth.json | metadata: tap layer, sizes, class names, training hyper-parameters |
The head runs in Speakrail's fork of audio.cpp (branch turn-head). The fork
adds the head and peek decoding to the Voxtral realtime model. Use it with the unmodified q8_0 GGUF of Voxtral
Mini 4B Realtime from audio-cpp/audio.cpp-gguf.
Server config (audiocpp_server --config server.json):
{
"port": 8090,
"backend": "cuda",
"models": [
{
"id": "voxtral-rt",
"family": "voxtral_realtime",
"path": "models/voxtral-mini-4b-realtime-2602-q8_0.gguf",
"task": "asr",
"mode": "streaming",
"session_options": {
"voxtral_realtime.turn_head": "models/turn_head.vxth"
}
}
]
}
Stream 16 kHz mono s16le PCM to POST /v1/audio/transcriptions/live?model=voxtral-rt&sample_rate=16000&channels=1&sample_format=s16le
(chunked upload, SSE response). Every decoded 80 ms step carries the head's probabilities:
{"type": "transcript.steps", "steps": [{"frame": 412, "token": 1032, "text": " today", "turn": [0.02, 0.95, 0.02, 0.01, 0.0]}]}
turn = softmax over [speaking, complete, incomplete, backchannel, wait].frame is the token frame. The head has heard the audio up to frame frame + 6, because of the 480 ms delay line. Add 6 to align with the audio clock.Peek: With &live_id=<id> on the live request, POST /v1/audio/transcriptions/live/peek?live_id=<id>&frames=8, you can use the peek functionality. The peek functionality is a latency-saving measure, whenever the head fires <complete>, it creates 6 frames (480ms) of empty chunks in the "future" and starts computing the arrived chunks as fast as possible. In practice, a peek takes approximately 60-70ms on a 4090, therefore saving 410ms on a turn fire. Inspired by Unmute's flush trick (described on the Kyutai STT page).
The probabilities flicker frame to frame, so Speakrail turns them into events with streak rules. The events go into the LLM's context as tokens, and the LLM decides what to do with them.
| signal | rule | what happens |
|---|---|---|
| user finished (assistant silent) | P(complete) + P(backchannel) >= 0.9 on 3 consecutive frames (240 ms), armed only by a new word | peek for the last words, then send <complete>. The LLM picks speak / listen / interrupt. A backchannel counts as "your move" here, because the user handed the floor back. |
| user finished (assistant talking) | P(complete) >= 0.9 on 3 consecutive frames; never if the overlapping segment was a backchannel | <complete> for the overlap. The LLM picks continue / yield / listen. |
| backchannel | P(backchannel) rises above 0.95 | <user_bc>. Sent over the assistant's speech, or once inside a user turn that already has words. A backchannel over the assistant's speech never ends its turn. |
| user speaking | P(speaking) >= 0.5 | voice activity: lower the assistant's volume while the user talks over it (the LLM still decides whether to yield), cancel a speculative reply, restart the silence clock. |
| likely finished | P(complete) + P(backchannel) >= 0.5 on the first silent frame | start a speculative reply on a copy of the context. It is held until <complete> fires and thrown away if the user keeps talking. |
| silence fallback | 1 s of silence after words with no head fire | <complete> anyway, for the trailing-off endings the head misses. The LLM still decides. |
P(incomplete) is not read directly. It matters by keeping the other classes below their thresholds during mid-turn pauses.
All numbers come from the head on Voxtral's bf16 hidden states at 480 ms delay (Hugging Face transformers), except the last table, which checks the audio.cpp q8_0 port.
2,329 clips held out from the training distribution (split by clip id; same sources as training, see Training). Clips include natural trailing silence, room reverb and noise.
Frame-level:
| class | F1 | precision | recall |
|---|---|---|---|
| speaking | 0.984 | 0.989 | 0.979 |
| complete | 0.939 | 0.933 | 0.945 |
| incomplete | 0.932 | 0.936 | 0.928 |
| backchannel | 0.906 | 0.876 | 0.937 |
End of turn: fire on the first frame with P(complete) >= tau after the speech stops.
These numbers apply to two situations:
<complete> cue.| endpoint | turn ends detected | latency median / p90 | fires on unfinished utterances | fires on backchannels |
|---|---|---|---|---|
| head, tau = 0.5 | 99.1 % | 80 / 160 ms | 15.5 % | 10.9 % |
| head, tau = 0.9 | 97.7 % | 80 / 240 ms | 9.3 % | 3.6 % |
| head, tau = 0.95 | 97.1 % | 80 / 240 ms | 7.6 % | 1.6 % |
| Voxtral punctuation (. ? !) | 92.1 % | 560 / 640 ms | 33.5 % | 71.4 % |
These are event-level metrics from the head's output on the user channel of Full-Duplex-Bench v1.0/v1.5 clips. They score the head's cues alone, not a full system: in Speakrail the LLM can still answer "listen" to a wrong cue. Rows by situation:
<complete> cue.These columns score P(complete) alone, so they match the assistant-talking rule. The assistant-silent rule also counts backchannels.
| subset | metric | head (0.9 x 3 frames) | head (0.9 x 3) + 1.5 s fallback | 1.5 s silence |
|---|---|---|---|---|
| CANDOR pause handling | cuts in during a pause ↓ | 11.2 % | 13.4 % | 2.9 % |
| synthetic pause handling | cuts in during a pause ↓ | 1.5 % | 4.4 % | 3.6 % |
| user backchannel | takes the floor on a backchannel ↓ | 2.0 % | 4.1 % | 100 % |
| user interruption | interruption detected / median latency | 53.5 % / 310 ms | 98.0 % / 1.05 s | 100 % / 1.53 s |
| CANDOR turn taking | turn ends detected / median latency | 38.7 % / 440 ms | 98.3 % / 1.43 s | 99.2 % / 1.46 s |
The head almost never mistakes a pause or a backchannel for a turn end. Its weakness is recall: it catches only about 40 % of the casual, trailing-off turn ends in real conversation (CANDOR). The rest need the silence fallback.
| P(speaking) correlation | P(complete) correlation | argmax agreement | |
|---|---|---|---|
| range over the 3 recordings | 0.994-0.998 | 0.89-0.96 | 91-99 % |
mistralai/Voxtral-Mini-4B-Realtime-2602, frozen. It runs free (its own generated tokens) at the default 480 ms delay, exactly as when serving. The head reads the residual stream after decoder layer 11 (d = 3072). A layer sweep showed mid-layers (7-11) beat the final layer by about 3 F1 points.speaking class down-weighted to 0.3.wait class: reserved and untrained. Ignore it.X2-Turn adds a turn-state output to the same Voxtral backbone with full fine-tuning. This head keeps the backbone frozen and asks how much of the turn state is already in its representation.
.vxth formatLittle-endian:
VXTHEAD1.tap, d_in, hidden, classes.w1[hidden][d_in], b1[hidden], w2[classes][hidden], b2[classes].p = softmax(w2 · gelu(w1 · h + b1) + b2), where h is the raw layer-tap residual.
Apache 2.0, same as Voxtral Mini 4B Realtime.
@misc{speakrail2026turnhead,
title = {Voxtral Mini 4B Realtime Turn Head},
author = {Speakrail},
year = {2026},
url = {https://huggingface.co/speakrail/Voxtral-Mini-4B-Realtime-2602-TurnHead}
}
Speakrail Voxtral Mini 4B Realtime Turn Head
0
4 commits
7 linked in READMEs
updated Oct 5, 2026
A turn-taking head for Voxtral Mini 4B Realtime. It reads one hidden layer of the streaming speech recognizer and, every 80 ms, says whether the user is still speaking, has finished their turn, has paused mid-thought, or just made a backchannel ("mm-hmm", "yeah").
The head is a 2.4 M-parameter MLP that runs inside the recognizer's decode step, so the transcript and the turn state come out of the same forward pass at no extra cost (0.02 ms per step on an RTX 4090), and the transcript is bit-identical with and without the head (the model's weights were completely frozen during the training).
The head doesn't take turns by itself. It gives four streaming signals, roughly "is the user talking", "did they finish", "is this a pause mid-thought" and "was that a backchannel", for a dialogue model or a rule set to act on. In Speakrail, a full-duplex voice assistant that runs on a single RTX 4090, the head turns these signals into cue tokens for the LLM, and the LLM makes every turn-taking decision. It can still reply "listen" to a cue.
A VAD / silence timer is not suitable for full-duplex interactions, and on my tests I found the Smart Turn models unreliable. I noticed that the Kyutai STT adds additional heads to the STT to help guide the LLM (the "semantic VAD" of kyutai/stt-1b-en_fr, used in Unmute), and I recently heard a speech about attaching a turn taking head to the GigaAM model, and that motivated me to attach a similar head to Voxtral, since the model is quite large, knows how to punctuate, and, therefore, should be able to understand when to take turns. Similar work was done by X2-Turn (model), but they fully fine-tuned the model (encoder and decoder) on Chinese and English turn data to allow it to turn-take, which I thought was unnecessary, and keeping the weights frozen kept the WER exactly as it is and gave a good-enough quality for me to consider the pipeline.
| file | what |
|---|---|
turn_head.vxth | the head for audio.cpp (format below) |
turn_head.vxth.json | metadata: tap layer, sizes, class names, training hyper-parameters |
The head runs in Speakrail's fork of audio.cpp (branch turn-head). The fork
adds the head and peek decoding to the Voxtral realtime model. Use it with the unmodified q8_0 GGUF of Voxtral
Mini 4B Realtime from audio-cpp/audio.cpp-gguf.
Server config (audiocpp_server --config server.json):
{
"port": 8090,
"backend": "cuda",
"models": [
{
"id": "voxtral-rt",
"family": "voxtral_realtime",
"path": "models/voxtral-mini-4b-realtime-2602-q8_0.gguf",
"task": "asr",
"mode": "streaming",
"session_options": {
"voxtral_realtime.turn_head": "models/turn_head.vxth"
}
}
]
}
Stream 16 kHz mono s16le PCM to POST /v1/audio/transcriptions/live?model=voxtral-rt&sample_rate=16000&channels=1&sample_format=s16le
(chunked upload, SSE response). Every decoded 80 ms step carries the head's probabilities:
{"type": "transcript.steps", "steps": [{"frame": 412, "token": 1032, "text": " today", "turn": [0.02, 0.95, 0.02, 0.01, 0.0]}]}
turn = softmax over [speaking, complete, incomplete, backchannel, wait].frame is the token frame. The head has heard the audio up to frame frame + 6, because of the 480 ms delay line. Add 6 to align with the audio clock.Peek: With &live_id=<id> on the live request, POST /v1/audio/transcriptions/live/peek?live_id=<id>&frames=8, you can use the peek functionality. The peek functionality is a latency-saving measure, whenever the head fires <complete>, it creates 6 frames (480ms) of empty chunks in the "future" and starts computing the arrived chunks as fast as possible. In practice, a peek takes approximately 60-70ms on a 4090, therefore saving 410ms on a turn fire. Inspired by Unmute's flush trick (described on the Kyutai STT page).
The probabilities flicker frame to frame, so Speakrail turns them into events with streak rules. The events go into the LLM's context as tokens, and the LLM decides what to do with them.
| signal | rule | what happens |
|---|---|---|
| user finished (assistant silent) | P(complete) + P(backchannel) >= 0.9 on 3 consecutive frames (240 ms), armed only by a new word | peek for the last words, then send <complete>. The LLM picks speak / listen / interrupt. A backchannel counts as "your move" here, because the user handed the floor back. |
| user finished (assistant talking) | P(complete) >= 0.9 on 3 consecutive frames; never if the overlapping segment was a backchannel | <complete> for the overlap. The LLM picks continue / yield / listen. |
| backchannel | P(backchannel) rises above 0.95 | <user_bc>. Sent over the assistant's speech, or once inside a user turn that already has words. A backchannel over the assistant's speech never ends its turn. |
| user speaking | P(speaking) >= 0.5 | voice activity: lower the assistant's volume while the user talks over it (the LLM still decides whether to yield), cancel a speculative reply, restart the silence clock. |
| likely finished | P(complete) + P(backchannel) >= 0.5 on the first silent frame | start a speculative reply on a copy of the context. It is held until <complete> fires and thrown away if the user keeps talking. |
| silence fallback | 1 s of silence after words with no head fire | <complete> anyway, for the trailing-off endings the head misses. The LLM still decides. |
P(incomplete) is not read directly. It matters by keeping the other classes below their thresholds during mid-turn pauses.
All numbers come from the head on Voxtral's bf16 hidden states at 480 ms delay (Hugging Face transformers), except the last table, which checks the audio.cpp q8_0 port.
2,329 clips held out from the training distribution (split by clip id; same sources as training, see Training). Clips include natural trailing silence, room reverb and noise.
Frame-level:
| class | F1 | precision | recall |
|---|---|---|---|
| speaking | 0.984 | 0.989 | 0.979 |
| complete | 0.939 | 0.933 | 0.945 |
| incomplete | 0.932 | 0.936 | 0.928 |
| backchannel | 0.906 | 0.876 | 0.937 |
End of turn: fire on the first frame with P(complete) >= tau after the speech stops.
These numbers apply to two situations:
<complete> cue.| endpoint | turn ends detected | latency median / p90 | fires on unfinished utterances | fires on backchannels |
|---|---|---|---|---|
| head, tau = 0.5 | 99.1 % | 80 / 160 ms | 15.5 % | 10.9 % |
| head, tau = 0.9 | 97.7 % | 80 / 240 ms | 9.3 % | 3.6 % |
| head, tau = 0.95 | 97.1 % | 80 / 240 ms | 7.6 % | 1.6 % |
| Voxtral punctuation (. ? !) | 92.1 % | 560 / 640 ms | 33.5 % | 71.4 % |
These are event-level metrics from the head's output on the user channel of Full-Duplex-Bench v1.0/v1.5 clips. They score the head's cues alone, not a full system: in Speakrail the LLM can still answer "listen" to a wrong cue. Rows by situation:
<complete> cue.These columns score P(complete) alone, so they match the assistant-talking rule. The assistant-silent rule also counts backchannels.
| subset | metric | head (0.9 x 3 frames) | head (0.9 x 3) + 1.5 s fallback | 1.5 s silence |
|---|---|---|---|---|
| CANDOR pause handling | cuts in during a pause ↓ | 11.2 % | 13.4 % | 2.9 % |
| synthetic pause handling | cuts in during a pause ↓ | 1.5 % | 4.4 % | 3.6 % |
| user backchannel | takes the floor on a backchannel ↓ | 2.0 % | 4.1 % | 100 % |
| user interruption | interruption detected / median latency | 53.5 % / 310 ms | 98.0 % / 1.05 s | 100 % / 1.53 s |
| CANDOR turn taking | turn ends detected / median latency | 38.7 % / 440 ms | 98.3 % / 1.43 s | 99.2 % / 1.46 s |
The head almost never mistakes a pause or a backchannel for a turn end. Its weakness is recall: it catches only about 40 % of the casual, trailing-off turn ends in real conversation (CANDOR). The rest need the silence fallback.
| P(speaking) correlation | P(complete) correlation | argmax agreement | |
|---|---|---|---|
| range over the 3 recordings | 0.994-0.998 | 0.89-0.96 | 91-99 % |
mistralai/Voxtral-Mini-4B-Realtime-2602, frozen. It runs free (its own generated tokens) at the default 480 ms delay, exactly as when serving. The head reads the residual stream after decoder layer 11 (d = 3072). A layer sweep showed mid-layers (7-11) beat the final layer by about 3 F1 points.speaking class down-weighted to 0.3.wait class: reserved and untrained. Ignore it.X2-Turn adds a turn-state output to the same Voxtral backbone with full fine-tuning. This head keeps the backbone frozen and asks how much of the turn state is already in its representation.
.vxth formatLittle-endian:
VXTHEAD1.tap, d_in, hidden, classes.w1[hidden][d_in], b1[hidden], w2[classes][hidden], b2[classes].p = softmax(w2 · gelu(w1 · h + b1) + b2), where h is the raw layer-tap residual.
Apache 2.0, same as Voxtral Mini 4B Realtime.
@misc{speakrail2026turnhead,
title = {Voxtral Mini 4B Realtime Turn Head},
author = {Speakrail},
year = {2026},
url = {https://huggingface.co/speakrail/Voxtral-Mini-4B-Realtime-2602-TurnHead}
}