Ruiqi-Yan/Awesome-Full-Duplex-SDM

A curated list of full-duplex spoken dialogue models & benchmarks

247

49 commits

updated Sep 21, 2026

See the code

README

Awesome-FullDuplexSDM Awesome

A curated list of full-duplex spoken dialogue models.
Welcome to PR if you want to add some resources.

Legend

Type — what the entry actually gives you:

TypeMeaning
End-to-endA standalone full-duplex model: audio in, audio out, duplex behaviour learned inside the model. (omni) means it also takes vision.
CascadedA complete system assembled from separate parts (streaming ASR + LLM + TTS + a duplex controller).
ComponentOne piece of the puzzle — turn detection, semantic VAD, streaming TTS, retrieval, dialogue management. Needs a host system.
MethodA training recipe, reward model, or data scheme rather than a deployable system.
FrameworkEngineering scaffolding for building/serving duplex systems.

Open — what has been released:

OpenMeaning
Code + weightsPublic repo and released checkpoints.
CodePublic repo, no checkpoints released (or none linked).
API / closedUsable as a product or API; nothing released.
Paper or tech report only.
?Could not verify.

Year is the year of first public release (arXiv v1, blog post, or repo). Entries without a dated paper are marked with the year they appeared publicly.

Datasets

TitleYearOpenRelevant Resources
ConversationalVoice: Full-Duplex Speech Data from Real Conversations2026CodeGithub/Demo
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues2026Data + codearXiv/Github/Huggingface/Demo
DuplexChat2026Data + codeGithub/Huggingface
SOMMELIER: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models2026Code (pipeline)arXiv/Demo/Github/Blog
SmoothConv & DuplexConv: Large-Scale Chinese Full-Duplex Speech Datasets for Conversational AI2026Data + codeGithub/Demo/Huggingface
TURNS-2K2025DataGithub/Huggingface

Models

TitleYearTypeOpenRelevant Resources
Realtime-Venus: A full-duplex interaction system with asynchronous delegation2026End-to-endCode + weightsarXiv/Github/Huggingface/Demo
SteerDuplex: Steerable Duplex Speech Dialogue Models2026End-to-endTo be releasedarXiv/Github
Omni Interaction Agent Technical Report2026CascadedCode + weightsarXiv/Github/Huggingface/Demo
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction2026Component — semantic VADCode + weightsarXiv/Github/Huggingface
JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents2026CascadedarXiv
Qwen Audio Agent2026CascadedCodeGithub
Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models2026MethodCodearXiv/Github
DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction2026End-to-end (omni)CodearXiv/Github
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs2026End-to-endCode + weightsarXiv/Github/Huggingface/Demo
BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM2026End-to-endCode + weightsarXiv/Github/Huggingface
TurnSense and TurnSense 1.1: Three-Class Semantic Turn Detection for Chinese and English Speech Interaction2026Component — turn detectionCode + weightsGithub/Huggingface
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action2026End-to-endCodeGithub/Demo
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System2025CascadedCodearXiv/Github/Demo
Unmute2025CascadedCode + weightsarXiv/Github/Demo
Reinforcement Learning Enhanced Full-Duplex Spoken Dialogue Language Models for Conversational Interactions2026Methodpaper
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models2026Component — retrievalCode + weightsarXiv/Github/Huggingface
Seeduplex: Native Full-Duplex Speech LLM2026End-to-endAPI / closedRelease/Blog
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection2026Component — turn detectionarXiv
JAL-Turn: Joint Acoustic-Linguistic Modeling for Real-Time and Robust Turn-Taking Detection in Full-Duplex Spoken Dialogue Systems2026Component — turn detectionarXiv
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving2025MethodCodearXiv/Github/Demo
Qwen3.5-Omni2026End-to-end (omni)API / closedOfficial Blog
Covo-Audio2026End-to-endCode + weightsarXiv/Github/Huggingface
SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation2026Component — semantic VADCode + weightsarXiv/Github/Huggingface/Demo
PHOENIX-VAD: STREAMING SEMANTIC ENDPOINT DETECTION FOR FULL-DUPLEX SPEECH INTERACTION2025Component — VAD/endpointingarXiv
Turnsense: A Lightweight End-of-Utterance Detection Model2025Component — turn detectionCode + weightsGithub/Huggingface
EASY TURN: INTEGRATING ACOUSTIC AND LINGUISTIC MODALITIES FOR ROBUST TURN-TAKING IN FULL-DUPLEX SPOKEN DIALOGUE SYSTEMS2025Component — turn detectionCodearXiv/Github/Demo
Fun-Audio-Chat2025End-to-endCodearXiv/Github/Demo
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations2025CascadedCodearXiv/Github/Demo
PERSONAPLEX: VOICE AND ROLE CONTROL FOR FULL DUPLEX CONVERSATIONAL SPEECH MODELS2026End-to-endCodearXiv/Github/Demo
VITA: Towards Open-Source Interactive Omni Multimodal LLM2024End-to-end (omni)Code + weightsarXiv/Github/Demo
A Full-duplex Speech Dialogue Scheme Based On Large Language Models2024End-to-endarXiv
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM2024End-to-endCode + weightsarXiv/Github
Moshi: a speech-text foundation model for real-time dialogue2024End-to-endCode + weightsarXiv/Github
FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems2025Component — duplex controlarXiv
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction2025End-to-endarXiv/Demo
SoulX-DuoVoice2025End-to-end?Unofficial Intro
SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation2025End-to-endCodearXiv/Github
CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction2025FrameworkCodearXiv/Github
Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities2024End-to-end (omni)Code + weightsarXiv/Github
Real-Time Textless Dialogue Generation2025End-to-endCodearXiv/Github/Demo
Language Model Can Listen While Speaking2024End-to-endarXiv/Demo
Parrot: Seamless Spoken Dialogue Interaction with Double-Channel Large Language Models2024End-to-endPaper
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents2024End-to-endarXiv/Demo
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems2025Component — dialogue mgmtarXiv
Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems2022CascadedarXiv
Generative Spoken Dialogue Language Modeling2022End-to-endCode + weightsarXiv/Github/Demo
Duplex Conversation in Outbound Agent System2021CascadedPaper
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation2024End-to-endarXiv/Demo
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities2025End-to-endCodearXiv/Github
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models2024MethodCode + dataarXiv/Github/HuggingFace
KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI2025Component — retrievalCode + weightsarXiv/Github/HuggingFace
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency2025Component — streaming TTSCodearXiv/Github/Demo
VoXtream2: Full-stream TTS with dynamic speaking rate control2026Component — streaming TTSCodearXiv/Github/Demo

Benchmark

TitleYearRelevant Resources
TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue2026arXiv/Demo/Dataset/Github
Game-Time: Evaluating Temporal Dynamics in Spoken Language Models2025arXiv/Demo/Dataset
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model2026arXiv/Github
Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge2026arXiv/Github/Dataset
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency2026arXiv/Github/Demo
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner2025arXiv/Github
FULL-DUPLEX-BENCH V1.5: Evaluating Overlap Handling for Full-Duplex Speech Models2025arXiv/Github
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities2025arXiv/Github
MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models2025arXiv
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems2025arXiv/Github/Dataset
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics2025arXiv

Survey

TitleYearRelevant Resources
A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine2026arXiv/Github
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models2025arXiv/Github
awesome-list
full-duplex
spoken-dialogue-models
spoken-dialogue-systems

Contributors

Ruiqi-Yan

45 commits

yoshikawasan

2 commits

ga642381

1 commits

mathigatti

1 commits

Ruiqi-Yan/Awesome-Full-Duplex-SDM

A curated list of full-duplex spoken dialogue models & benchmarks

247

49 commits

updated Sep 21, 2026

See the code

README

Awesome-FullDuplexSDM Awesome

A curated list of full-duplex spoken dialogue models.
Welcome to PR if you want to add some resources.

Legend

Type — what the entry actually gives you:

TypeMeaning
End-to-endA standalone full-duplex model: audio in, audio out, duplex behaviour learned inside the model. (omni) means it also takes vision.
CascadedA complete system assembled from separate parts (streaming ASR + LLM + TTS + a duplex controller).
ComponentOne piece of the puzzle — turn detection, semantic VAD, streaming TTS, retrieval, dialogue management. Needs a host system.
MethodA training recipe, reward model, or data scheme rather than a deployable system.
FrameworkEngineering scaffolding for building/serving duplex systems.

Open — what has been released:

OpenMeaning
Code + weightsPublic repo and released checkpoints.
CodePublic repo, no checkpoints released (or none linked).
API / closedUsable as a product or API; nothing released.
Paper or tech report only.
?Could not verify.

Year is the year of first public release (arXiv v1, blog post, or repo). Entries without a dated paper are marked with the year they appeared publicly.

Datasets

TitleYearOpenRelevant Resources
ConversationalVoice: Full-Duplex Speech Data from Real Conversations2026CodeGithub/Demo
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues2026Data + codearXiv/Github/Huggingface/Demo
DuplexChat2026Data + codeGithub/Huggingface
SOMMELIER: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models2026Code (pipeline)arXiv/Demo/Github/Blog
SmoothConv & DuplexConv: Large-Scale Chinese Full-Duplex Speech Datasets for Conversational AI2026Data + codeGithub/Demo/Huggingface
TURNS-2K2025DataGithub/Huggingface

Models

TitleYearTypeOpenRelevant Resources
Realtime-Venus: A full-duplex interaction system with asynchronous delegation2026End-to-endCode + weightsarXiv/Github/Huggingface/Demo
SteerDuplex: Steerable Duplex Speech Dialogue Models2026End-to-endTo be releasedarXiv/Github
Omni Interaction Agent Technical Report2026CascadedCode + weightsarXiv/Github/Huggingface/Demo
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction2026Component — semantic VADCode + weightsarXiv/Github/Huggingface
JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents2026CascadedarXiv
Qwen Audio Agent2026CascadedCodeGithub
Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models2026MethodCodearXiv/Github
DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction2026End-to-end (omni)CodearXiv/Github
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs2026End-to-endCode + weightsarXiv/Github/Huggingface/Demo
BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM2026End-to-endCode + weightsarXiv/Github/Huggingface
TurnSense and TurnSense 1.1: Three-Class Semantic Turn Detection for Chinese and English Speech Interaction2026Component — turn detectionCode + weightsGithub/Huggingface
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action2026End-to-endCodeGithub/Demo
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System2025CascadedCodearXiv/Github/Demo
Unmute2025CascadedCode + weightsarXiv/Github/Demo
Reinforcement Learning Enhanced Full-Duplex Spoken Dialogue Language Models for Conversational Interactions2026Methodpaper
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models2026Component — retrievalCode + weightsarXiv/Github/Huggingface
Seeduplex: Native Full-Duplex Speech LLM2026End-to-endAPI / closedRelease/Blog
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection2026Component — turn detectionarXiv
JAL-Turn: Joint Acoustic-Linguistic Modeling for Real-Time and Robust Turn-Taking Detection in Full-Duplex Spoken Dialogue Systems2026Component — turn detectionarXiv
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving2025MethodCodearXiv/Github/Demo
Qwen3.5-Omni2026End-to-end (omni)API / closedOfficial Blog
Covo-Audio2026End-to-endCode + weightsarXiv/Github/Huggingface
SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation2026Component — semantic VADCode + weightsarXiv/Github/Huggingface/Demo
PHOENIX-VAD: STREAMING SEMANTIC ENDPOINT DETECTION FOR FULL-DUPLEX SPEECH INTERACTION2025Component — VAD/endpointingarXiv
Turnsense: A Lightweight End-of-Utterance Detection Model2025Component — turn detectionCode + weightsGithub/Huggingface
EASY TURN: INTEGRATING ACOUSTIC AND LINGUISTIC MODALITIES FOR ROBUST TURN-TAKING IN FULL-DUPLEX SPOKEN DIALOGUE SYSTEMS2025Component — turn detectionCodearXiv/Github/Demo
Fun-Audio-Chat2025End-to-endCodearXiv/Github/Demo
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations2025CascadedCodearXiv/Github/Demo
PERSONAPLEX: VOICE AND ROLE CONTROL FOR FULL DUPLEX CONVERSATIONAL SPEECH MODELS2026End-to-endCodearXiv/Github/Demo
VITA: Towards Open-Source Interactive Omni Multimodal LLM2024End-to-end (omni)Code + weightsarXiv/Github/Demo
A Full-duplex Speech Dialogue Scheme Based On Large Language Models2024End-to-endarXiv
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM2024End-to-endCode + weightsarXiv/Github
Moshi: a speech-text foundation model for real-time dialogue2024End-to-endCode + weightsarXiv/Github
FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems2025Component — duplex controlarXiv
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction2025End-to-endarXiv/Demo
SoulX-DuoVoice2025End-to-end?Unofficial Intro
SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation2025End-to-endCodearXiv/Github
CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction2025FrameworkCodearXiv/Github
Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities2024End-to-end (omni)Code + weightsarXiv/Github
Real-Time Textless Dialogue Generation2025End-to-endCodearXiv/Github/Demo
Language Model Can Listen While Speaking2024End-to-endarXiv/Demo
Parrot: Seamless Spoken Dialogue Interaction with Double-Channel Large Language Models2024End-to-endPaper
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents2024End-to-endarXiv/Demo
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems2025Component — dialogue mgmtarXiv
Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems2022CascadedarXiv
Generative Spoken Dialogue Language Modeling2022End-to-endCode + weightsarXiv/Github/Demo
Duplex Conversation in Outbound Agent System2021CascadedPaper
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation2024End-to-endarXiv/Demo
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities2025End-to-endCodearXiv/Github
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models2024MethodCode + dataarXiv/Github/HuggingFace
KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI2025Component — retrievalCode + weightsarXiv/Github/HuggingFace
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency2025Component — streaming TTSCodearXiv/Github/Demo
VoXtream2: Full-stream TTS with dynamic speaking rate control2026Component — streaming TTSCodearXiv/Github/Demo

Benchmark

TitleYearRelevant Resources
TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue2026arXiv/Demo/Dataset/Github
Game-Time: Evaluating Temporal Dynamics in Spoken Language Models2025arXiv/Demo/Dataset
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model2026arXiv/Github
Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge2026arXiv/Github/Dataset
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency2026arXiv/Github/Demo
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner2025arXiv/Github
FULL-DUPLEX-BENCH V1.5: Evaluating Overlap Handling for Full-Duplex Speech Models2025arXiv/Github
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities2025arXiv/Github
MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models2025arXiv
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems2025arXiv/Github/Dataset
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics2025arXiv

Survey

TitleYearRelevant Resources
A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine2026arXiv/Github
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models2025arXiv/Github
awesome-list
full-duplex
spoken-dialogue-models
spoken-dialogue-systems

Contributors

Ruiqi-Yan

45 commits

yoshikawasan

2 commits

ga642381

1 commits

mathigatti

1 commits