speakrail/Voxtral-Mini-4B-Realtime-2602-TurnHead

Model

Speakrail Voxtral Mini 4B Realtime Turn Head

0

4 commits

7 linked in READMEs

updated Oct 5, 2026

See the code

README

Speakrail Voxtral Mini 4B Realtime Turn Head

A turn-taking head for Voxtral Mini 4B Realtime. It reads one hidden layer of the streaming speech recognizer and, every 80 ms, says whether the user is still speaking, has finished their turn, has paused mid-thought, or just made a backchannel ("mm-hmm", "yeah").

The head is a 2.4 M-parameter MLP that runs inside the recognizer's decode step, so the transcript and the turn state come out of the same forward pass at no extra cost (0.02 ms per step on an RTX 4090), and the transcript is bit-identical with and without the head (the model's weights were completely frozen during the training).

The head doesn't take turns by itself. It gives four streaming signals, roughly "is the user talking", "did they finish", "is this a pause mid-thought" and "was that a backchannel", for a dialogue model or a rule set to act on. In Speakrail, a full-duplex voice assistant that runs on a single RTX 4090, the head turns these signals into cue tokens for the LLM, and the LLM makes every turn-taking decision. It can still reply "listen" to a cue.

Why

A VAD / silence timer is not suitable for full-duplex interactions, and on my tests I found the Smart Turn models unreliable. I noticed that the Kyutai STT adds additional heads to the STT to help guide the LLM (the "semantic VAD" of kyutai/stt-1b-en_fr, used in Unmute), and I recently heard a speech about attaching a turn taking head to the GigaAM model, and that motivated me to attach a similar head to Voxtral, since the model is quite large, knows how to punctuate, and, therefore, should be able to understand when to take turns. Similar work was done by X2-Turn (model), but they fully fine-tuned the model (encoder and decoder) on Chinese and English turn data to allow it to turn-take, which I thought was unnecessary, and keeping the weights frozen kept the WER exactly as it is and gave a good-enough quality for me to consider the pipeline.

Files

filewhat
turn_head.vxththe head for audio.cpp (format below)
turn_head.vxth.jsonmetadata: tap layer, sizes, class names, training hyper-parameters

Usage (audio.cpp)

The head runs in Speakrail's fork of audio.cpp (branch turn-head). The fork adds the head and peek decoding to the Voxtral realtime model. Use it with the unmodified q8_0 GGUF of Voxtral Mini 4B Realtime from audio-cpp/audio.cpp-gguf.

Server config (audiocpp_server --config server.json):

{
  "port": 8090,
  "backend": "cuda",
  "models": [
    {
      "id": "voxtral-rt",
      "family": "voxtral_realtime",
      "path": "models/voxtral-mini-4b-realtime-2602-q8_0.gguf",
      "task": "asr",
      "mode": "streaming",
      "session_options": {
        "voxtral_realtime.turn_head": "models/turn_head.vxth"
      }
    }
  ]
}

Stream 16 kHz mono s16le PCM to POST /v1/audio/transcriptions/live?model=voxtral-rt&sample_rate=16000&channels=1&sample_format=s16le (chunked upload, SSE response). Every decoded 80 ms step carries the head's probabilities:

{"type": "transcript.steps", "steps": [{"frame": 412, "token": 1032, "text": " today", "turn": [0.02, 0.95, 0.02, 0.01, 0.0]}]}
  • turn = softmax over [speaking, complete, incomplete, backchannel, wait].
  • frame is the token frame. The head has heard the audio up to frame frame + 6, because of the 480 ms delay line. Add 6 to align with the audio clock.
  • Use Voxtral's default 480 ms transcription delay. That is what the head was trained on.

Peek: With &live_id=<id> on the live request, POST /v1/audio/transcriptions/live/peek?live_id=<id>&frames=8, you can use the peek functionality. The peek functionality is a latency-saving measure, whenever the head fires <complete>, it creates 6 frames (480ms) of empty chunks in the "future" and starts computing the arrived chunks as fast as possible. In practice, a peek takes approximately 60-70ms on a 4090, therefore saving 410ms on a turn fire. Inspired by Unmute's flush trick (described on the Kyutai STT page).

How Speakrail uses it

The probabilities flicker frame to frame, so Speakrail turns them into events with streak rules. The events go into the LLM's context as tokens, and the LLM decides what to do with them.

signalrulewhat happens
user finished (assistant silent)P(complete) + P(backchannel) >= 0.9 on 3 consecutive frames (240 ms), armed only by a new wordpeek for the last words, then send <complete>. The LLM picks speak / listen / interrupt. A backchannel counts as "your move" here, because the user handed the floor back.
user finished (assistant talking)P(complete) >= 0.9 on 3 consecutive frames; never if the overlapping segment was a backchannel<complete> for the overlap. The LLM picks continue / yield / listen.
backchannelP(backchannel) rises above 0.95<user_bc>. Sent over the assistant's speech, or once inside a user turn that already has words. A backchannel over the assistant's speech never ends its turn.
user speakingP(speaking) >= 0.5voice activity: lower the assistant's volume while the user talks over it (the LLM still decides whether to yield), cancel a speculative reply, restart the silence clock.
likely finishedP(complete) + P(backchannel) >= 0.5 on the first silent framestart a speculative reply on a copy of the context. It is held until <complete> fires and thrown away if the user keeps talking.
silence fallback1 s of silence after words with no head fire<complete> anyway, for the trailing-off endings the head misses. The LLM still decides.

P(incomplete) is not read directly. It matters by keeping the other classes below their thresholds during mid-turn pauses.

Evaluation

All numbers come from the head on Voxtral's bf16 hidden states at 480 ms delay (Hugging Face transformers), except the last table, which checks the audio.cpp q8_0 port.

Held-out test clips

2,329 clips held out from the training distribution (split by clip id; same sources as training, see Training). Clips include natural trailing silence, room reverb and noise.

Frame-level:

classF1precisionrecall
speaking0.9840.9890.979
complete0.9390.9330.945
incomplete0.9320.9360.928
backchannel0.9060.8760.937

End of turn: fire on the first frame with P(complete) >= tau after the speech stops.

These numbers apply to two situations:

  • Assistant silent: turn ends detected, latency and fires on unfinished utterances measure the <complete> cue.
  • Assistant talking: fires on backchannels measures how often a backchannel would wrongly end the assistant's turn. While the assistant is silent, Speakrail deliberately counts a backchannel as a turn end, so this column doesn't apply there.
endpointturn ends detectedlatency median / p90fires on unfinished utterancesfires on backchannels
head, tau = 0.599.1 %80 / 160 ms15.5 %10.9 %
head, tau = 0.997.7 %80 / 240 ms9.3 %3.6 %
head, tau = 0.9597.1 %80 / 240 ms7.6 %1.6 %
Voxtral punctuation (. ? !)92.1 %560 / 640 ms33.5 %71.4 %

Full-Duplex-Bench (never seen in training)

These are event-level metrics from the head's output on the user channel of Full-Duplex-Bench v1.0/v1.5 clips. They score the head's cues alone, not a full system: in Speakrail the LLM can still answer "listen" to a wrong cue. Rows by situation:

  • Assistant silent: pause handling and CANDOR turn taking measure the <complete> cue.
  • Assistant talking: user backchannel and user interruption measure what happens while it speaks.

These columns score P(complete) alone, so they match the assistant-talking rule. The assistant-silent rule also counts backchannels.

subsetmetrichead (0.9 x 3 frames)head (0.9 x 3) + 1.5 s fallback1.5 s silence
CANDOR pause handlingcuts in during a pause ↓11.2 %13.4 %2.9 %
synthetic pause handlingcuts in during a pause ↓1.5 %4.4 %3.6 %
user backchanneltakes the floor on a backchannel ↓2.0 %4.1 %100 %
user interruptioninterruption detected / median latency53.5 % / 310 ms98.0 % / 1.05 s100 % / 1.53 s
CANDOR turn takingturn ends detected / median latency38.7 % / 440 ms98.3 % / 1.43 s99.2 % / 1.46 s

The head almost never mistakes a pause or a backchannel for a turn end. Its weakness is recall: it catches only about 40 % of the casual, trailing-off turn ends in real conversation (CANDOR). The rest need the silence fallback.

audio.cpp q8_0 port

  • Head export: the exported head matches the PyTorch head to a maximum |Δp| of 1e-8.
  • Trunk: the q8_0 GGUF trunk was compared, frame by frame, against an 8-bit PyTorch reference trunk with the same head, on 3 recordings (165 s).
P(speaking) correlationP(complete) correlationargmax agreement
range over the 3 recordings0.994-0.9980.89-0.9691-99 %

Training

  • Trunk: mistralai/Voxtral-Mini-4B-Realtime-2602, frozen. It runs free (its own generated tokens) at the default 480 ms delay, exactly as when serving. The head reads the residual stream after decoder layer 11 (d = 3072). A layer sweep showed mid-layers (7-11) beat the final layer by about 3 F1 points.
  • Head: MLP, 3072 → 768 (GELU, dropout 0.1) → 5, with the input standardisation folded into the first layer.
    • Training: 8 epochs, AdamW with a one-cycle schedule (peak lr 1e-3), weight decay 0.01, batch 64.
    • Loss: class-balanced cross-entropy, with the speaking class down-weighted to 0.3.
  • Data: 23,757 English clips (19,132 train / 2,296 val / 2,329 test).
    • 19,740 clips from pipecat-ai/smart-turn-data-v3.1-train (CC BY 4.0), human and synthetic speech, labelled complete / incomplete.
    • 4,017 synthetic clips (backchannels and short complete / incomplete phrases) in 134 speaker voices.
  • Frame labels: from Silero VAD.
    • Frames with speech: speaking.
    • Frames after the final speech offset: the clip's label (complete / incomplete / backchannel).
    • Silent frames before the final offset: incomplete, so mid-turn pauses are supervised.
    • Clips are never cut at the offset; the natural tail is kept.
  • Augmentation: room impulse responses (OpenSLR-28) and noise (MUSAN and others) at 5-40 dB SNR, plus random gain.
  • wait class: reserved and untrained. Ignore it.

Limitations

  • English only. Trained and evaluated on English. Voxtral is multilingual, but the head has not been tested on other languages.
  • Fixed delay. It only works at 480 ms. Other transcription delays shift the hidden states; there was a separate head for 80 ms, which is not released.
  • No anticipation. It fires when the audio goes quiet, not before. The completeness information is there about 240 ms earlier, but this head was not trained to use it.
  • Casual turn ends. It catches only about 40 % of trailing-off turn ends in real conversation, you should pair it with a silence fallback. This happens because the training data was focused on assistant speech and is probably overtuned to questions in general, rather than chit-chatting.
  • Speaker-blind. It only hears the user channel. Background speech, or the user talking to someone else, looks like a turn.
  • No public benchmark yet. I will release a better turn taking version later, along with proper benchmarking, since this turn head was good enough for the prototype, and I just used it.

X2-Turn adds a turn-state output to the same Voxtral backbone with full fine-tuning. This head keeps the backbone frozen and asks how much of the turn state is already in its representation.

.vxth format

Little-endian:

  • 8-byte magic VXTHEAD1.
  • int32 tap, d_in, hidden, classes.
  • f32 w1[hidden][d_in], b1[hidden], w2[classes][hidden], b2[classes].
  • The class names, one per line.

p = softmax(w2 · gelu(w1 · h + b1) + b2), where h is the raw layer-tap residual.

License

Apache 2.0, same as Voxtral Mini 4B Realtime.

Citation

@misc{speakrail2026turnhead,
  title  = {Voxtral Mini 4B Realtime Turn Head},
  author = {Speakrail},
  year   = {2026},
  url    = {https://huggingface.co/speakrail/Voxtral-Mini-4B-Realtime-2602-TurnHead}
}
audio-classification
audio.cpp
backchannel
end-of-turn
endpointing
full-duplex
turn-taking
voice-assistant
voxtral

speakrail/Voxtral-Mini-4B-Realtime-2602-TurnHead

Model

Speakrail Voxtral Mini 4B Realtime Turn Head

0

4 commits

7 linked in READMEs

updated Oct 5, 2026

See the code

README

Speakrail Voxtral Mini 4B Realtime Turn Head

A turn-taking head for Voxtral Mini 4B Realtime. It reads one hidden layer of the streaming speech recognizer and, every 80 ms, says whether the user is still speaking, has finished their turn, has paused mid-thought, or just made a backchannel ("mm-hmm", "yeah").

The head is a 2.4 M-parameter MLP that runs inside the recognizer's decode step, so the transcript and the turn state come out of the same forward pass at no extra cost (0.02 ms per step on an RTX 4090), and the transcript is bit-identical with and without the head (the model's weights were completely frozen during the training).

The head doesn't take turns by itself. It gives four streaming signals, roughly "is the user talking", "did they finish", "is this a pause mid-thought" and "was that a backchannel", for a dialogue model or a rule set to act on. In Speakrail, a full-duplex voice assistant that runs on a single RTX 4090, the head turns these signals into cue tokens for the LLM, and the LLM makes every turn-taking decision. It can still reply "listen" to a cue.

Why

A VAD / silence timer is not suitable for full-duplex interactions, and on my tests I found the Smart Turn models unreliable. I noticed that the Kyutai STT adds additional heads to the STT to help guide the LLM (the "semantic VAD" of kyutai/stt-1b-en_fr, used in Unmute), and I recently heard a speech about attaching a turn taking head to the GigaAM model, and that motivated me to attach a similar head to Voxtral, since the model is quite large, knows how to punctuate, and, therefore, should be able to understand when to take turns. Similar work was done by X2-Turn (model), but they fully fine-tuned the model (encoder and decoder) on Chinese and English turn data to allow it to turn-take, which I thought was unnecessary, and keeping the weights frozen kept the WER exactly as it is and gave a good-enough quality for me to consider the pipeline.

Files

filewhat
turn_head.vxththe head for audio.cpp (format below)
turn_head.vxth.jsonmetadata: tap layer, sizes, class names, training hyper-parameters

Usage (audio.cpp)

The head runs in Speakrail's fork of audio.cpp (branch turn-head). The fork adds the head and peek decoding to the Voxtral realtime model. Use it with the unmodified q8_0 GGUF of Voxtral Mini 4B Realtime from audio-cpp/audio.cpp-gguf.

Server config (audiocpp_server --config server.json):

{
  "port": 8090,
  "backend": "cuda",
  "models": [
    {
      "id": "voxtral-rt",
      "family": "voxtral_realtime",
      "path": "models/voxtral-mini-4b-realtime-2602-q8_0.gguf",
      "task": "asr",
      "mode": "streaming",
      "session_options": {
        "voxtral_realtime.turn_head": "models/turn_head.vxth"
      }
    }
  ]
}

Stream 16 kHz mono s16le PCM to POST /v1/audio/transcriptions/live?model=voxtral-rt&sample_rate=16000&channels=1&sample_format=s16le (chunked upload, SSE response). Every decoded 80 ms step carries the head's probabilities:

{"type": "transcript.steps", "steps": [{"frame": 412, "token": 1032, "text": " today", "turn": [0.02, 0.95, 0.02, 0.01, 0.0]}]}
  • turn = softmax over [speaking, complete, incomplete, backchannel, wait].
  • frame is the token frame. The head has heard the audio up to frame frame + 6, because of the 480 ms delay line. Add 6 to align with the audio clock.
  • Use Voxtral's default 480 ms transcription delay. That is what the head was trained on.

Peek: With &live_id=<id> on the live request, POST /v1/audio/transcriptions/live/peek?live_id=<id>&frames=8, you can use the peek functionality. The peek functionality is a latency-saving measure, whenever the head fires <complete>, it creates 6 frames (480ms) of empty chunks in the "future" and starts computing the arrived chunks as fast as possible. In practice, a peek takes approximately 60-70ms on a 4090, therefore saving 410ms on a turn fire. Inspired by Unmute's flush trick (described on the Kyutai STT page).

How Speakrail uses it

The probabilities flicker frame to frame, so Speakrail turns them into events with streak rules. The events go into the LLM's context as tokens, and the LLM decides what to do with them.

signalrulewhat happens
user finished (assistant silent)P(complete) + P(backchannel) >= 0.9 on 3 consecutive frames (240 ms), armed only by a new wordpeek for the last words, then send <complete>. The LLM picks speak / listen / interrupt. A backchannel counts as "your move" here, because the user handed the floor back.
user finished (assistant talking)P(complete) >= 0.9 on 3 consecutive frames; never if the overlapping segment was a backchannel<complete> for the overlap. The LLM picks continue / yield / listen.
backchannelP(backchannel) rises above 0.95<user_bc>. Sent over the assistant's speech, or once inside a user turn that already has words. A backchannel over the assistant's speech never ends its turn.
user speakingP(speaking) >= 0.5voice activity: lower the assistant's volume while the user talks over it (the LLM still decides whether to yield), cancel a speculative reply, restart the silence clock.
likely finishedP(complete) + P(backchannel) >= 0.5 on the first silent framestart a speculative reply on a copy of the context. It is held until <complete> fires and thrown away if the user keeps talking.
silence fallback1 s of silence after words with no head fire<complete> anyway, for the trailing-off endings the head misses. The LLM still decides.

P(incomplete) is not read directly. It matters by keeping the other classes below their thresholds during mid-turn pauses.

Evaluation

All numbers come from the head on Voxtral's bf16 hidden states at 480 ms delay (Hugging Face transformers), except the last table, which checks the audio.cpp q8_0 port.

Held-out test clips

2,329 clips held out from the training distribution (split by clip id; same sources as training, see Training). Clips include natural trailing silence, room reverb and noise.

Frame-level:

classF1precisionrecall
speaking0.9840.9890.979
complete0.9390.9330.945
incomplete0.9320.9360.928
backchannel0.9060.8760.937

End of turn: fire on the first frame with P(complete) >= tau after the speech stops.

These numbers apply to two situations:

  • Assistant silent: turn ends detected, latency and fires on unfinished utterances measure the <complete> cue.
  • Assistant talking: fires on backchannels measures how often a backchannel would wrongly end the assistant's turn. While the assistant is silent, Speakrail deliberately counts a backchannel as a turn end, so this column doesn't apply there.
endpointturn ends detectedlatency median / p90fires on unfinished utterancesfires on backchannels
head, tau = 0.599.1 %80 / 160 ms15.5 %10.9 %
head, tau = 0.997.7 %80 / 240 ms9.3 %3.6 %
head, tau = 0.9597.1 %80 / 240 ms7.6 %1.6 %
Voxtral punctuation (. ? !)92.1 %560 / 640 ms33.5 %71.4 %

Full-Duplex-Bench (never seen in training)

These are event-level metrics from the head's output on the user channel of Full-Duplex-Bench v1.0/v1.5 clips. They score the head's cues alone, not a full system: in Speakrail the LLM can still answer "listen" to a wrong cue. Rows by situation:

  • Assistant silent: pause handling and CANDOR turn taking measure the <complete> cue.
  • Assistant talking: user backchannel and user interruption measure what happens while it speaks.

These columns score P(complete) alone, so they match the assistant-talking rule. The assistant-silent rule also counts backchannels.

subsetmetrichead (0.9 x 3 frames)head (0.9 x 3) + 1.5 s fallback1.5 s silence
CANDOR pause handlingcuts in during a pause ↓11.2 %13.4 %2.9 %
synthetic pause handlingcuts in during a pause ↓1.5 %4.4 %3.6 %
user backchanneltakes the floor on a backchannel ↓2.0 %4.1 %100 %
user interruptioninterruption detected / median latency53.5 % / 310 ms98.0 % / 1.05 s100 % / 1.53 s
CANDOR turn takingturn ends detected / median latency38.7 % / 440 ms98.3 % / 1.43 s99.2 % / 1.46 s

The head almost never mistakes a pause or a backchannel for a turn end. Its weakness is recall: it catches only about 40 % of the casual, trailing-off turn ends in real conversation (CANDOR). The rest need the silence fallback.

audio.cpp q8_0 port

  • Head export: the exported head matches the PyTorch head to a maximum |Δp| of 1e-8.
  • Trunk: the q8_0 GGUF trunk was compared, frame by frame, against an 8-bit PyTorch reference trunk with the same head, on 3 recordings (165 s).
P(speaking) correlationP(complete) correlationargmax agreement
range over the 3 recordings0.994-0.9980.89-0.9691-99 %

Training

  • Trunk: mistralai/Voxtral-Mini-4B-Realtime-2602, frozen. It runs free (its own generated tokens) at the default 480 ms delay, exactly as when serving. The head reads the residual stream after decoder layer 11 (d = 3072). A layer sweep showed mid-layers (7-11) beat the final layer by about 3 F1 points.
  • Head: MLP, 3072 → 768 (GELU, dropout 0.1) → 5, with the input standardisation folded into the first layer.
    • Training: 8 epochs, AdamW with a one-cycle schedule (peak lr 1e-3), weight decay 0.01, batch 64.
    • Loss: class-balanced cross-entropy, with the speaking class down-weighted to 0.3.
  • Data: 23,757 English clips (19,132 train / 2,296 val / 2,329 test).
    • 19,740 clips from pipecat-ai/smart-turn-data-v3.1-train (CC BY 4.0), human and synthetic speech, labelled complete / incomplete.
    • 4,017 synthetic clips (backchannels and short complete / incomplete phrases) in 134 speaker voices.
  • Frame labels: from Silero VAD.
    • Frames with speech: speaking.
    • Frames after the final speech offset: the clip's label (complete / incomplete / backchannel).
    • Silent frames before the final offset: incomplete, so mid-turn pauses are supervised.
    • Clips are never cut at the offset; the natural tail is kept.
  • Augmentation: room impulse responses (OpenSLR-28) and noise (MUSAN and others) at 5-40 dB SNR, plus random gain.
  • wait class: reserved and untrained. Ignore it.

Limitations

  • English only. Trained and evaluated on English. Voxtral is multilingual, but the head has not been tested on other languages.
  • Fixed delay. It only works at 480 ms. Other transcription delays shift the hidden states; there was a separate head for 80 ms, which is not released.
  • No anticipation. It fires when the audio goes quiet, not before. The completeness information is there about 240 ms earlier, but this head was not trained to use it.
  • Casual turn ends. It catches only about 40 % of trailing-off turn ends in real conversation, you should pair it with a silence fallback. This happens because the training data was focused on assistant speech and is probably overtuned to questions in general, rather than chit-chatting.
  • Speaker-blind. It only hears the user channel. Background speech, or the user talking to someone else, looks like a turn.
  • No public benchmark yet. I will release a better turn taking version later, along with proper benchmarking, since this turn head was good enough for the prototype, and I just used it.

X2-Turn adds a turn-state output to the same Voxtral backbone with full fine-tuning. This head keeps the backbone frozen and asks how much of the turn state is already in its representation.

.vxth format

Little-endian:

  • 8-byte magic VXTHEAD1.
  • int32 tap, d_in, hidden, classes.
  • f32 w1[hidden][d_in], b1[hidden], w2[classes][hidden], b2[classes].
  • The class names, one per line.

p = softmax(w2 · gelu(w1 · h + b1) + b2), where h is the raw layer-tap residual.

License

Apache 2.0, same as Voxtral Mini 4B Realtime.

Citation

@misc{speakrail2026turnhead,
  title  = {Voxtral Mini 4B Realtime Turn Head},
  author = {Speakrail},
  year   = {2026},
  url    = {https://huggingface.co/speakrail/Voxtral-Mini-4B-Realtime-2602-TurnHead}
}
audio-classification
audio.cpp
backchannel
end-of-turn
endpointing
full-duplex
turn-taking
voice-assistant
voxtral