TedZhangHao/turn-taking-naturalness

4

stars

5

commits

Python

primary language

Jul 15, 2026

updated

README

Turn-Taking Naturalness

TurnNat is an automatic evaluation framework for conversational turn-taking naturalness. It focuses on whether a dialogue sounds temporally plausible around speaker transitions, holds, interruptions, response delays, and backchannels, rather than on lexical content alone. The core idea is to score a natural conversation and a matched timing-perturbed version with a future-VAD (FVAD) model, then use the change in dialogue-level NLL as a proxy for how unnatural the timing perturbation is.

This repository contains two connected pieces:

  • TurnNat Perturbation Benchmark: a fixed 1,000-pair benchmark with five conversational timing perturbation types and 200 paired examples per type.
  • TurnNat FVAD scorer: training and evaluation code for FVAD-based naturalness scoring, including a released DualTurn FVAD-256 checkpoint.

The public pipeline keeps perturbation generation and model evaluation separate: perturbed samples are produced from data-only timing and VAD rules, while model scoring is run as a standalone post-training evaluation step. This separation is intentional, so the benchmark examples are not generated by the released checkpoint and the checkpoint can be evaluated on the fixed benchmark.

Setup

There are two requirements files because data generation and model training have very different dependency footprints:

  • requirements.txt: lightweight dependencies for perturbation generation.
  • requirements-train.txt: optional model stack for FVAD training and checkpoint evaluation.

For data generation only:

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For model training or checkpoint evaluation:

pip install -r requirements-train.txt
pip install -e VAP-main

Checkpoint and data locations:

  • TurnNat artifact package: https://figshare.com/s/65e5bf5290085220ee88. The package contains both the 1,000-pair TurnNat Perturbation Benchmark and the released TurnNat FVAD checkpoint.
  • Benchmark placement: after downloading, place or symlink the benchmark manifest at data/benchmark/manifests/test2.csv.
  • Checkpoint placement: place the downloaded checkpoint at checkpoints/turnnat-dualturn-fvad256.pt, or pass the downloaded path with --checkpoint. Expected SHA256: 96472553406762c662adae1197941d42347b6ada93bb8bccb80479a72b66ba97.
  • Checkpoint configuration/provenance: see configs/released_turnnat_dualturn_fvad256/. This directory contains the released checkpoint's training override YAML plus sanitized training-time config snapshots.
  • Official DualTurn backbone/checkpoint: anyreach-ai/dualturn-qwen2.5-mimi-0.5B. The DualTurn scripts load this model from Hugging Face by default. For offline runs, pre-cache it with Hugging Face/Transformers and use --local-files-only.
  • Official VAP code/checkpoint example: ErikEkstedt/VAP. The VAP scorer/trainer expects the raw state dict at VAP-main/example/checkpoints/VAP_state_dict.pt. Download it with:
mkdir -p VAP-main/example/checkpoints
curl -L \
  https://github.com/ErikEkstedt/VAP/raw/main/example/checkpoints/VAP_state_dict.pt \
  -o VAP-main/example/checkpoints/VAP_state_dict.pt

You can also pass a different VAP state-dict path with --checkpoint or --vap-ckpt.

Data Preparation

The fixed TurnNat Perturbation Benchmark contains 1,000 paired natural/perturbed examples:

  • early_entry: 200 pairs
  • late_response: 200 pairs
  • shift_instead_of_hold: 200 pairs
  • hold_instead_of_shift: 200 pairs
  • excessive_backchannel: 200 pairs

The benchmark and released FVAD checkpoint are distributed in one Figshare artifact package:

https://figshare.com/s/65e5bf5290085220ee88

See data/benchmark/README.md for the benchmark layout and checkpoints/README.md for the checkpoint filename and SHA256. If you use the released benchmark, point the evaluation commands below to the downloaded manifest.

To generate perturbations from your own conversations, prepare CSV manifests in the format described in data/manifests/README.md:

  • train.csv, dev.csv, and test.csv for FVAD training/evaluation.
  • test_natural.csv for perturbation generation from natural conversations.

Each natural-conversation row should point to the two participant audio files and transcript/metadata paths when available.

Generate Perturbed Data

Generate all five perturbation types:

python turnnat/scripts/build_perturbations.py make-unnatural \
  --natural-csv data/manifests/test_natural.csv \
  --out-root data/generated \
  --split test \
  --per-type 200 \
  --types early_entry,late_response,shift_instead_of_hold,hold_instead_of_shift,excessive_backchannel \
  --short-context

Outputs are written under data/generated/:

  • manifests/test.csv: paired natural/perturbed manifest
  • test/audio/: perturbed audio
  • test/json/: perturbation metadata (edit_meta in JSON for compatibility)
  • test/natural_audio/ and test/natural_json/: matched natural references
  • generated_test_rows.csv, test_generation_log.csv, and test_failures.csv: generation bookkeeping and failure reasons

Default Perturbation Logic

Common defaults:

  • turn source: silero
  • one perturbation per generated clip
  • short-context crop target: 20-25s
  • crop guard keeps the perturbed boundary and nearby turn-taking context inside the crop
  • generated WAVs are normalized to -20 dBFS RMS and -1 dBFS peak
  • splice boundaries use short fades/crossfades where audio is shifted, inserted, or removed
  • inserted silence and removed speech regions are filled with nearby room tone where synthetic background is needed

Per-type defaults:

  • early_entry: move the responder earlier by 1.2-2.5s at a clean A-B turn transition.
  • late_response: delay the responder by 1.2-2.0s; require adjacent speaker and responder turns of at least 1s, with at most 200ms overlap.
  • shift_instead_of_hold: insert a short shift turn into an original same- speaker hold gap of 0.3-1.0s; inserted turn duration is 1-4s.
  • hold_instead_of_shift: remove a responder turn of 1.2-8.0s, then compact the resulting hold gap to about 0.5-1.5s.
  • excessive_backchannel: insert two short, distinct backchannels by default, with 800ms target spacing and no volume change. To make three-backchannel examples, rerun with --bc-insert-count 3.

These rules are deliberately not overly restrictive. For cleaner public datasets, curate the input natural manifest to avoid clips where one channel is mostly silent, transcripts are missing, or the two speakers are poorly aligned.

Useful generation variants:

# One perturbation type only
python turnnat/scripts/build_perturbations.py make-unnatural \
  --natural-csv data/manifests/test_natural.csv \
  --out-root data/generated_late \
  --split test \
  --per-type 200 \
  --types late_response \
  --short-context

# Three inserted backchannels instead of the default two
python turnnat/scripts/build_perturbations.py make-unnatural \
  --natural-csv data/manifests/test_natural.csv \
  --out-root data/generated_bc3 \
  --split test \
  --per-type 100 \
  --types excessive_backchannel \
  --bc-insert-count 3 \
  --short-context

# Sharded generation; give each worker a separate output directory
python turnnat/scripts/build_perturbations.py make-unnatural \
  --natural-csv data/manifests/test_natural.csv \
  --out-root data/generated_late_shard00 \
  --split test \
  --per-type 200 \
  --types late_response \
  --num-shards 16 \
  --shard-index 0 \
  --short-context

Try Your Own Audio

turnnat/scripts/try_audio.py is a lightweight entry point for quick demos. It expects two-channel audio where channel 0 and channel 1 are the two speakers. Mono files are duplicated to two channels, but stereo audio is recommended.

Score one input recording directly with the official DualTurn FVAD head:

python turnnat/scripts/try_audio.py score \
  --audio path/to/conversation.wav \
  --score-backend official-dualturn \
  --output-dir outputs/try_audio_dualturn

Score one recording with the released FVAD checkpoint downloaded from Figshare:

python turnnat/scripts/try_audio.py score \
  --audio path/to/conversation.wav \
  --score-backend fvad-checkpoint \
  --checkpoint checkpoints/turnnat-dualturn-fvad256.pt \
  --experiment auto \
  --output-dir outputs/try_audio_fvad

Compare a natural/perturbed pair. Positive delta_nll means the perturbed file has higher DialogNLL than the natural reference under the selected model:

python turnnat/scripts/try_audio.py score-pair \
  --natural-audio path/to/natural.wav \
  --perturbed-audio path/to/perturbed.wav \
  --score-backend official-vap \
  --checkpoint VAP-main/example/checkpoints/VAP_state_dict.pt \
  --perturbation-type late_response \
  --output-dir outputs/try_audio_pair_vap

For high-quality perturbation generation, use a natural manifest rather than a standalone wav. Transcript/metadata sidecars let the generator avoid weak candidates and produce cleaner timing perturbations. The same script can run one generation pass and immediately score the generated pairs:

python turnnat/scripts/try_audio.py perturb-and-score \
  --natural-csv data/manifests/test_natural.csv \
  --perturbation-type late_response \
  --per-type 1 \
  --score-backend fvad-checkpoint \
  --checkpoint checkpoints/turnnat-dualturn-fvad256.pt \
  --experiment auto \
  --output-dir outputs/try_late_response

Outputs include metrics.json, segment_scores.csv, units.csv, and, for pairs, pair_scores.csv. Use --score-backend official-dualturn, official-vap, or fvad-checkpoint depending on which checkpoint you want to try.

Train FVAD Models

Create data/manifests/train.csv, dev.csv, and test.csv using the format in data/manifests/README.md. The included configs train with standard FVAD train/validation loss and frame accuracy. best.pt is selected by validation loss.

The training label path is shared across backbones: load two-channel audio, build per-channel 50 Hz Silero VAD labels, then adapt that same VAD stream to VAP 256-state targets or DualTurn native/all-six targets. The chunk dataset only loads audio and frame masks; it does not derive RMS VAD or turn-action labels.

For DualTurn all-six training, first cache VAD and signal labels:

python turnnat/scripts/build_silero_vad_cache.py \
  --manifest data/manifests/train.csv --output-dir data/cache/vad
python turnnat/scripts/build_silero_vad_cache.py \
  --manifest data/manifests/dev.csv --output-dir data/cache/vad --skip-existing
python turnnat/scripts/build_silero_vad_cache.py \
  --manifest data/manifests/test.csv --output-dir data/cache/vad --skip-existing

python turnnat/scripts/build_dualturn_signal_cache.py \
  --manifest data/manifests/train.csv \
  --manifest data/manifests/dev.csv \
  --manifest data/manifests/test.csv \
  --vad-cache-dir data/cache/vad --output-dir data/cache/signals

Run an experiment:

python turnnat/scripts/run_fvad_experiment.py configs/train_vap_full.yaml
python turnnat/scripts/run_fvad_experiment.py configs/train_dualturn_native_all6.yaml
python turnnat/scripts/run_fvad_experiment.py configs/train_dualturn_fvad256_all6.yaml

Training checkpoints are written to outputs/<experiment>/checkpoints/ as last.pt and best.pt.

Evaluate Checkpoints

Score a trained FVAD checkpoint on a paired natural/perturbed manifest:

python turnnat/scripts/score_fvad_checkpoint.py \
  --checkpoint outputs/dualturn_fvad256_all6/checkpoints/best.pt \
  --experiment auto \
  --manifest data/generated/manifests/test.csv \
  --output-dir outputs/eval_dualturn_fvad256 \
  --device cuda \
  --batch-size 16

For the shared/released checkpoint, download it separately to checkpoints/turnnat-dualturn-fvad256.pt or pass its path with --checkpoint. The checkpoint was evaluated with the NaturalnessFiveTypeEvaluator implementation embedded in turnnat/scripts/train_fvad_head.py, matching the original experiment runner. You can list supported compatibility profile names with:

python turnnat/scripts/score_fvad_checkpoint.py --list-experiments

Evaluate the released 1,000-pair benchmark with the released checkpoint:

python turnnat/scripts/score_fvad_checkpoint.py \
  --checkpoint checkpoints/turnnat-dualturn-fvad256.pt \
  --experiment auto \
  --manifest data/benchmark/manifests/test2.csv \
  --output-dir outputs/eval_turnnat_dualturn_fvad256_benchmark \
  --device cuda \
  --batch-size 16 \
  --local-files-only

Score an upstream VAP checkpoint directly:

python turnnat/scripts/score_vap_nll_naturalness.py \
  --unnatural-manifest data/generated/manifests/test.csv \
  --checkpoint VAP-main/example/checkpoints/VAP_state_dict.pt \
  --output-dir outputs/eval_vap \
  --device cuda

Metrics

The scorer computes frame-level future-VAD NLL, averages it inside pre-boundary utterance units, and reports:

MetricDefinitionBetter
MeanNLLMean NLL over all utterance unitsLower for a natural recording
TailNLLMean of the worst 25% unit NLL valuesLower
DialogNLL0.5 * MeanNLL + 0.5 * TailNLLLower
NatScore / nat_scoreDialogue-level turn-taking naturalness score, defined as -DialogNLLHigher
DeltaNLLperturbed DialogNLL - natural DialogNLLPositive/larger
Pairwise AccuracyFraction of matched pairs with DeltaNLL > 0Higher
C-indexFraction of all perturbed-vs-natural dialogue-NLL comparisons correctly ordered; ties excludedHigher

Results are reported overall and separately for all five perturbation types. Output is written under OUTPUT_DIR/step_<global_step>/:

  • metrics.json and metrics.csv: aggregate metrics with variance and 95% CI
  • pair_scores.csv: natural/perturbed scores and DeltaNLL for each pair
  • segment_scores.csv: recording-level MeanNLL, TailNLL, DialogNLL, and nat_score
  • units.csv: utterance-boundary unit NLL values
  • inference_config.json: resolved experiment profile and checkpoint metadata

These metrics are for standalone evaluation and reporting.

Samples

samples/manifest.csv contains one natural/perturbed pair for each type: early_entry, late_response, shift_instead_of_hold, hold_instead_of_shift, and excessive_backchannel.

Score the sample pairs with VAP:

python turnnat/scripts/score_vap_nll_naturalness.py \
  --unnatural-manifest samples/manifest.csv \
  --checkpoint VAP-main/example/checkpoints/VAP_state_dict.pt \
  --output-dir outputs/sample_vap \
  --device cuda

Score the official DualTurn native 8-bit head:

python turnnat/scripts/score_dualturn_fvad_nll_naturalness.py \
  --unnatural-manifest samples/manifest.csv \
  --output-dir outputs/sample_dualturn \
  --device cuda

The TurnNat artifact package, containing the 1,000-pair perturbation benchmark and released FVAD checkpoint, is distributed through Figshare:

https://figshare.com/s/65e5bf5290085220ee88

See data/benchmark/README.md and checkpoints/README.md for the expected layout, filename, and SHA256.

Contributors

TedZhangHao/turn-taking-naturalness

4

stars

5

commits

Python

primary language

Jul 15, 2026

updated

README

Turn-Taking Naturalness

TurnNat is an automatic evaluation framework for conversational turn-taking naturalness. It focuses on whether a dialogue sounds temporally plausible around speaker transitions, holds, interruptions, response delays, and backchannels, rather than on lexical content alone. The core idea is to score a natural conversation and a matched timing-perturbed version with a future-VAD (FVAD) model, then use the change in dialogue-level NLL as a proxy for how unnatural the timing perturbation is.

This repository contains two connected pieces:

  • TurnNat Perturbation Benchmark: a fixed 1,000-pair benchmark with five conversational timing perturbation types and 200 paired examples per type.
  • TurnNat FVAD scorer: training and evaluation code for FVAD-based naturalness scoring, including a released DualTurn FVAD-256 checkpoint.

The public pipeline keeps perturbation generation and model evaluation separate: perturbed samples are produced from data-only timing and VAD rules, while model scoring is run as a standalone post-training evaluation step. This separation is intentional, so the benchmark examples are not generated by the released checkpoint and the checkpoint can be evaluated on the fixed benchmark.

Setup

There are two requirements files because data generation and model training have very different dependency footprints:

  • requirements.txt: lightweight dependencies for perturbation generation.
  • requirements-train.txt: optional model stack for FVAD training and checkpoint evaluation.

For data generation only:

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For model training or checkpoint evaluation:

pip install -r requirements-train.txt
pip install -e VAP-main

Checkpoint and data locations:

  • TurnNat artifact package: https://figshare.com/s/65e5bf5290085220ee88. The package contains both the 1,000-pair TurnNat Perturbation Benchmark and the released TurnNat FVAD checkpoint.
  • Benchmark placement: after downloading, place or symlink the benchmark manifest at data/benchmark/manifests/test2.csv.
  • Checkpoint placement: place the downloaded checkpoint at checkpoints/turnnat-dualturn-fvad256.pt, or pass the downloaded path with --checkpoint. Expected SHA256: 96472553406762c662adae1197941d42347b6ada93bb8bccb80479a72b66ba97.
  • Checkpoint configuration/provenance: see configs/released_turnnat_dualturn_fvad256/. This directory contains the released checkpoint's training override YAML plus sanitized training-time config snapshots.
  • Official DualTurn backbone/checkpoint: anyreach-ai/dualturn-qwen2.5-mimi-0.5B. The DualTurn scripts load this model from Hugging Face by default. For offline runs, pre-cache it with Hugging Face/Transformers and use --local-files-only.
  • Official VAP code/checkpoint example: ErikEkstedt/VAP. The VAP scorer/trainer expects the raw state dict at VAP-main/example/checkpoints/VAP_state_dict.pt. Download it with:
mkdir -p VAP-main/example/checkpoints
curl -L \
  https://github.com/ErikEkstedt/VAP/raw/main/example/checkpoints/VAP_state_dict.pt \
  -o VAP-main/example/checkpoints/VAP_state_dict.pt

You can also pass a different VAP state-dict path with --checkpoint or --vap-ckpt.

Data Preparation

The fixed TurnNat Perturbation Benchmark contains 1,000 paired natural/perturbed examples:

  • early_entry: 200 pairs
  • late_response: 200 pairs
  • shift_instead_of_hold: 200 pairs
  • hold_instead_of_shift: 200 pairs
  • excessive_backchannel: 200 pairs

The benchmark and released FVAD checkpoint are distributed in one Figshare artifact package:

https://figshare.com/s/65e5bf5290085220ee88

See data/benchmark/README.md for the benchmark layout and checkpoints/README.md for the checkpoint filename and SHA256. If you use the released benchmark, point the evaluation commands below to the downloaded manifest.

To generate perturbations from your own conversations, prepare CSV manifests in the format described in data/manifests/README.md:

  • train.csv, dev.csv, and test.csv for FVAD training/evaluation.
  • test_natural.csv for perturbation generation from natural conversations.

Each natural-conversation row should point to the two participant audio files and transcript/metadata paths when available.

Generate Perturbed Data

Generate all five perturbation types:

python turnnat/scripts/build_perturbations.py make-unnatural \
  --natural-csv data/manifests/test_natural.csv \
  --out-root data/generated \
  --split test \
  --per-type 200 \
  --types early_entry,late_response,shift_instead_of_hold,hold_instead_of_shift,excessive_backchannel \
  --short-context

Outputs are written under data/generated/:

  • manifests/test.csv: paired natural/perturbed manifest
  • test/audio/: perturbed audio
  • test/json/: perturbation metadata (edit_meta in JSON for compatibility)
  • test/natural_audio/ and test/natural_json/: matched natural references
  • generated_test_rows.csv, test_generation_log.csv, and test_failures.csv: generation bookkeeping and failure reasons

Default Perturbation Logic

Common defaults:

  • turn source: silero
  • one perturbation per generated clip
  • short-context crop target: 20-25s
  • crop guard keeps the perturbed boundary and nearby turn-taking context inside the crop
  • generated WAVs are normalized to -20 dBFS RMS and -1 dBFS peak
  • splice boundaries use short fades/crossfades where audio is shifted, inserted, or removed
  • inserted silence and removed speech regions are filled with nearby room tone where synthetic background is needed

Per-type defaults:

  • early_entry: move the responder earlier by 1.2-2.5s at a clean A-B turn transition.
  • late_response: delay the responder by 1.2-2.0s; require adjacent speaker and responder turns of at least 1s, with at most 200ms overlap.
  • shift_instead_of_hold: insert a short shift turn into an original same- speaker hold gap of 0.3-1.0s; inserted turn duration is 1-4s.
  • hold_instead_of_shift: remove a responder turn of 1.2-8.0s, then compact the resulting hold gap to about 0.5-1.5s.
  • excessive_backchannel: insert two short, distinct backchannels by default, with 800ms target spacing and no volume change. To make three-backchannel examples, rerun with --bc-insert-count 3.

These rules are deliberately not overly restrictive. For cleaner public datasets, curate the input natural manifest to avoid clips where one channel is mostly silent, transcripts are missing, or the two speakers are poorly aligned.

Useful generation variants:

# One perturbation type only
python turnnat/scripts/build_perturbations.py make-unnatural \
  --natural-csv data/manifests/test_natural.csv \
  --out-root data/generated_late \
  --split test \
  --per-type 200 \
  --types late_response \
  --short-context

# Three inserted backchannels instead of the default two
python turnnat/scripts/build_perturbations.py make-unnatural \
  --natural-csv data/manifests/test_natural.csv \
  --out-root data/generated_bc3 \
  --split test \
  --per-type 100 \
  --types excessive_backchannel \
  --bc-insert-count 3 \
  --short-context

# Sharded generation; give each worker a separate output directory
python turnnat/scripts/build_perturbations.py make-unnatural \
  --natural-csv data/manifests/test_natural.csv \
  --out-root data/generated_late_shard00 \
  --split test \
  --per-type 200 \
  --types late_response \
  --num-shards 16 \
  --shard-index 0 \
  --short-context

Try Your Own Audio

turnnat/scripts/try_audio.py is a lightweight entry point for quick demos. It expects two-channel audio where channel 0 and channel 1 are the two speakers. Mono files are duplicated to two channels, but stereo audio is recommended.

Score one input recording directly with the official DualTurn FVAD head:

python turnnat/scripts/try_audio.py score \
  --audio path/to/conversation.wav \
  --score-backend official-dualturn \
  --output-dir outputs/try_audio_dualturn

Score one recording with the released FVAD checkpoint downloaded from Figshare:

python turnnat/scripts/try_audio.py score \
  --audio path/to/conversation.wav \
  --score-backend fvad-checkpoint \
  --checkpoint checkpoints/turnnat-dualturn-fvad256.pt \
  --experiment auto \
  --output-dir outputs/try_audio_fvad

Compare a natural/perturbed pair. Positive delta_nll means the perturbed file has higher DialogNLL than the natural reference under the selected model:

python turnnat/scripts/try_audio.py score-pair \
  --natural-audio path/to/natural.wav \
  --perturbed-audio path/to/perturbed.wav \
  --score-backend official-vap \
  --checkpoint VAP-main/example/checkpoints/VAP_state_dict.pt \
  --perturbation-type late_response \
  --output-dir outputs/try_audio_pair_vap

For high-quality perturbation generation, use a natural manifest rather than a standalone wav. Transcript/metadata sidecars let the generator avoid weak candidates and produce cleaner timing perturbations. The same script can run one generation pass and immediately score the generated pairs:

python turnnat/scripts/try_audio.py perturb-and-score \
  --natural-csv data/manifests/test_natural.csv \
  --perturbation-type late_response \
  --per-type 1 \
  --score-backend fvad-checkpoint \
  --checkpoint checkpoints/turnnat-dualturn-fvad256.pt \
  --experiment auto \
  --output-dir outputs/try_late_response

Outputs include metrics.json, segment_scores.csv, units.csv, and, for pairs, pair_scores.csv. Use --score-backend official-dualturn, official-vap, or fvad-checkpoint depending on which checkpoint you want to try.

Train FVAD Models

Create data/manifests/train.csv, dev.csv, and test.csv using the format in data/manifests/README.md. The included configs train with standard FVAD train/validation loss and frame accuracy. best.pt is selected by validation loss.

The training label path is shared across backbones: load two-channel audio, build per-channel 50 Hz Silero VAD labels, then adapt that same VAD stream to VAP 256-state targets or DualTurn native/all-six targets. The chunk dataset only loads audio and frame masks; it does not derive RMS VAD or turn-action labels.

For DualTurn all-six training, first cache VAD and signal labels:

python turnnat/scripts/build_silero_vad_cache.py \
  --manifest data/manifests/train.csv --output-dir data/cache/vad
python turnnat/scripts/build_silero_vad_cache.py \
  --manifest data/manifests/dev.csv --output-dir data/cache/vad --skip-existing
python turnnat/scripts/build_silero_vad_cache.py \
  --manifest data/manifests/test.csv --output-dir data/cache/vad --skip-existing

python turnnat/scripts/build_dualturn_signal_cache.py \
  --manifest data/manifests/train.csv \
  --manifest data/manifests/dev.csv \
  --manifest data/manifests/test.csv \
  --vad-cache-dir data/cache/vad --output-dir data/cache/signals

Run an experiment:

python turnnat/scripts/run_fvad_experiment.py configs/train_vap_full.yaml
python turnnat/scripts/run_fvad_experiment.py configs/train_dualturn_native_all6.yaml
python turnnat/scripts/run_fvad_experiment.py configs/train_dualturn_fvad256_all6.yaml

Training checkpoints are written to outputs/<experiment>/checkpoints/ as last.pt and best.pt.

Evaluate Checkpoints

Score a trained FVAD checkpoint on a paired natural/perturbed manifest:

python turnnat/scripts/score_fvad_checkpoint.py \
  --checkpoint outputs/dualturn_fvad256_all6/checkpoints/best.pt \
  --experiment auto \
  --manifest data/generated/manifests/test.csv \
  --output-dir outputs/eval_dualturn_fvad256 \
  --device cuda \
  --batch-size 16

For the shared/released checkpoint, download it separately to checkpoints/turnnat-dualturn-fvad256.pt or pass its path with --checkpoint. The checkpoint was evaluated with the NaturalnessFiveTypeEvaluator implementation embedded in turnnat/scripts/train_fvad_head.py, matching the original experiment runner. You can list supported compatibility profile names with:

python turnnat/scripts/score_fvad_checkpoint.py --list-experiments

Evaluate the released 1,000-pair benchmark with the released checkpoint:

python turnnat/scripts/score_fvad_checkpoint.py \
  --checkpoint checkpoints/turnnat-dualturn-fvad256.pt \
  --experiment auto \
  --manifest data/benchmark/manifests/test2.csv \
  --output-dir outputs/eval_turnnat_dualturn_fvad256_benchmark \
  --device cuda \
  --batch-size 16 \
  --local-files-only

Score an upstream VAP checkpoint directly:

python turnnat/scripts/score_vap_nll_naturalness.py \
  --unnatural-manifest data/generated/manifests/test.csv \
  --checkpoint VAP-main/example/checkpoints/VAP_state_dict.pt \
  --output-dir outputs/eval_vap \
  --device cuda

Metrics

The scorer computes frame-level future-VAD NLL, averages it inside pre-boundary utterance units, and reports:

MetricDefinitionBetter
MeanNLLMean NLL over all utterance unitsLower for a natural recording
TailNLLMean of the worst 25% unit NLL valuesLower
DialogNLL0.5 * MeanNLL + 0.5 * TailNLLLower
NatScore / nat_scoreDialogue-level turn-taking naturalness score, defined as -DialogNLLHigher
DeltaNLLperturbed DialogNLL - natural DialogNLLPositive/larger
Pairwise AccuracyFraction of matched pairs with DeltaNLL > 0Higher
C-indexFraction of all perturbed-vs-natural dialogue-NLL comparisons correctly ordered; ties excludedHigher

Results are reported overall and separately for all five perturbation types. Output is written under OUTPUT_DIR/step_<global_step>/:

  • metrics.json and metrics.csv: aggregate metrics with variance and 95% CI
  • pair_scores.csv: natural/perturbed scores and DeltaNLL for each pair
  • segment_scores.csv: recording-level MeanNLL, TailNLL, DialogNLL, and nat_score
  • units.csv: utterance-boundary unit NLL values
  • inference_config.json: resolved experiment profile and checkpoint metadata

These metrics are for standalone evaluation and reporting.

Samples

samples/manifest.csv contains one natural/perturbed pair for each type: early_entry, late_response, shift_instead_of_hold, hold_instead_of_shift, and excessive_backchannel.

Score the sample pairs with VAP:

python turnnat/scripts/score_vap_nll_naturalness.py \
  --unnatural-manifest samples/manifest.csv \
  --checkpoint VAP-main/example/checkpoints/VAP_state_dict.pt \
  --output-dir outputs/sample_vap \
  --device cuda

Score the official DualTurn native 8-bit head:

python turnnat/scripts/score_dualturn_fvad_nll_naturalness.py \
  --unnatural-manifest samples/manifest.csv \
  --output-dir outputs/sample_dualturn \
  --device cuda

The TurnNat artifact package, containing the 1,000-pair perturbation benchmark and released FVAD checkpoint, is distributed through Figshare:

https://figshare.com/s/65e5bf5290085220ee88

See data/benchmark/README.md and checkpoints/README.md for the expected layout, filename, and SHA256.

Contributors

Languages

Python

99.9%