A curated list of full-duplex spoken dialogue models & benchmarks
247
49 commits
updated Sep 21, 2026
A curated list of full-duplex spoken dialogue models.
Welcome to PR if you want to add some resources.
Type — what the entry actually gives you:
| Type | Meaning |
|---|---|
| End-to-end | A standalone full-duplex model: audio in, audio out, duplex behaviour learned inside the model. (omni) means it also takes vision. |
| Cascaded | A complete system assembled from separate parts (streaming ASR + LLM + TTS + a duplex controller). |
| Component | One piece of the puzzle — turn detection, semantic VAD, streaming TTS, retrieval, dialogue management. Needs a host system. |
| Method | A training recipe, reward model, or data scheme rather than a deployable system. |
| Framework | Engineering scaffolding for building/serving duplex systems. |
Open — what has been released:
| Open | Meaning |
|---|---|
| Code + weights | Public repo and released checkpoints. |
| Code | Public repo, no checkpoints released (or none linked). |
| API / closed | Usable as a product or API; nothing released. |
| — | Paper or tech report only. |
| ? | Could not verify. |
Year is the year of first public release (arXiv v1, blog post, or repo). Entries without a dated paper are marked with the year they appeared publicly.
| Title | Year | Open | Relevant Resources |
|---|---|---|---|
| ConversationalVoice: Full-Duplex Speech Data from Real Conversations | 2026 | Code | Github/Demo |
| DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues | 2026 | Data + code | arXiv/Github/Huggingface/Demo |
| DuplexChat | 2026 | Data + code | Github/Huggingface |
| SOMMELIER: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models | 2026 | Code (pipeline) | arXiv/Demo/Github/Blog |
| SmoothConv & DuplexConv: Large-Scale Chinese Full-Duplex Speech Datasets for Conversational AI | 2026 | Data + code | Github/Demo/Huggingface |
| TURNS-2K | 2025 | Data | Github/Huggingface |
| Title | Year | Type | Open | Relevant Resources |
|---|---|---|---|---|
| Realtime-Venus: A full-duplex interaction system with asynchronous delegation | 2026 | End-to-end | Code + weights | arXiv/Github/Huggingface/Demo |
| SteerDuplex: Steerable Duplex Speech Dialogue Models | 2026 | End-to-end | To be released | arXiv/Github |
| Omni Interaction Agent Technical Report | 2026 | Cascaded | Code + weights | arXiv/Github/Huggingface/Demo |
| X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction | 2026 | Component — semantic VAD | Code + weights | arXiv/Github/Huggingface |
| JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents | 2026 | Cascaded | — | arXiv |
| Qwen Audio Agent | 2026 | Cascaded | Code | Github |
| Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models | 2026 | Method | Code | arXiv/Github |
| DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction | 2026 | End-to-end (omni) | Code | arXiv/Github |
| Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs | 2026 | End-to-end | Code + weights | arXiv/Github/Huggingface/Demo |
| BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM | 2026 | End-to-end | Code + weights | arXiv/Github/Huggingface |
| TurnSense and TurnSense 1.1: Three-Class Semantic Turn Detection for Chinese and English Speech Interaction | 2026 | Component — turn detection | Code + weights | Github/Huggingface |
| DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action | 2026 | End-to-end | Code | Github/Demo |
| X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System | 2025 | Cascaded | Code | arXiv/Github/Demo |
| Unmute | 2025 | Cascaded | Code + weights | arXiv/Github/Demo |
| Reinforcement Learning Enhanced Full-Duplex Spoken Dialogue Language Models for Conversational Interactions | 2026 | Method | — | paper |
| MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models | 2026 | Component — retrieval | Code + weights | arXiv/Github/Huggingface |
| Seeduplex: Native Full-Duplex Speech LLM | 2026 | End-to-end | API / closed | Release/Blog |
| FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection | 2026 | Component — turn detection | — | arXiv |
| JAL-Turn: Joint Acoustic-Linguistic Modeling for Real-Time and Robust Turn-Taking Detection in Full-Duplex Spoken Dialogue Systems | 2026 | Component — turn detection | — | arXiv |
| TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving | 2025 | Method | Code | arXiv/Github/Demo |
| Qwen3.5-Omni | 2026 | End-to-end (omni) | API / closed | Official Blog |
| Covo-Audio | 2026 | End-to-end | Code + weights | arXiv/Github/Huggingface |
| SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation | 2026 | Component — semantic VAD | Code + weights | arXiv/Github/Huggingface/Demo |
| PHOENIX-VAD: STREAMING SEMANTIC ENDPOINT DETECTION FOR FULL-DUPLEX SPEECH INTERACTION | 2025 | Component — VAD/endpointing | — | arXiv |
| Turnsense: A Lightweight End-of-Utterance Detection Model | 2025 | Component — turn detection | Code + weights | Github/Huggingface |
| EASY TURN: INTEGRATING ACOUSTIC AND LINGUISTIC MODALITIES FOR ROBUST TURN-TAKING IN FULL-DUPLEX SPOKEN DIALOGUE SYSTEMS | 2025 | Component — turn detection | Code | arXiv/Github/Demo |
| Fun-Audio-Chat | 2025 | End-to-end | Code | arXiv/Github/Demo |
| FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations | 2025 | Cascaded | Code | arXiv/Github/Demo |
| PERSONAPLEX: VOICE AND ROLE CONTROL FOR FULL DUPLEX CONVERSATIONAL SPEECH MODELS | 2026 | End-to-end | Code | arXiv/Github/Demo |
| VITA: Towards Open-Source Interactive Omni Multimodal LLM | 2024 | End-to-end (omni) | Code + weights | arXiv/Github/Demo |
| A Full-duplex Speech Dialogue Scheme Based On Large Language Models | 2024 | End-to-end | — | arXiv |
| Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM | 2024 | End-to-end | Code + weights | arXiv/Github |
| Moshi: a speech-text foundation model for real-time dialogue | 2024 | End-to-end | Code + weights | arXiv/Github |
| FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems | 2025 | Component — duplex control | — | arXiv |
| MinMo: A Multimodal Large Language Model for Seamless Voice Interaction | 2025 | End-to-end | — | arXiv/Demo |
| SoulX-DuoVoice | 2025 | End-to-end | ? | Unofficial Intro |
| SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation | 2025 | End-to-end | Code | arXiv/Github |
| CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction | 2025 | Framework | Code | arXiv/Github |
| Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities | 2024 | End-to-end (omni) | Code + weights | arXiv/Github |
| Real-Time Textless Dialogue Generation | 2025 | End-to-end | Code | arXiv/Github/Demo |
| Language Model Can Listen While Speaking | 2024 | End-to-end | — | arXiv/Demo |
| Parrot: Seamless Spoken Dialogue Interaction with Double-Channel Large Language Models | 2024 | End-to-end | — | Paper |
| Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents | 2024 | End-to-end | — | arXiv/Demo |
| LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems | 2025 | Component — dialogue mgmt | — | arXiv |
| Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems | 2022 | Cascaded | — | arXiv |
| Generative Spoken Dialogue Language Modeling | 2022 | End-to-end | Code + weights | arXiv/Github/Demo |
| Duplex Conversation in Outbound Agent System | 2021 | Cascaded | — | Paper |
| OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation | 2024 | End-to-end | — | arXiv/Demo |
| DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities | 2025 | End-to-end | Code | arXiv/Github |
| Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models | 2024 | Method | Code + data | arXiv/Github/HuggingFace |
| KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI | 2025 | Component — retrieval | Code + weights | arXiv/Github/HuggingFace |
| VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency | 2025 | Component — streaming TTS | Code | arXiv/Github/Demo |
| VoXtream2: Full-stream TTS with dynamic speaking rate control | 2026 | Component — streaming TTS | Code | arXiv/Github/Demo |
| Title | Year | Relevant Resources |
|---|---|---|
| TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue | 2026 | arXiv/Demo/Dataset/Github |
| Game-Time: Evaluating Temporal Dynamics in Spoken Language Models | 2025 | arXiv/Demo/Dataset |
| Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model | 2026 | arXiv/Github |
| Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge | 2026 | arXiv/Github/Dataset |
| Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency | 2026 | arXiv/Github/Demo |
| Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner | 2025 | arXiv/Github |
| FULL-DUPLEX-BENCH V1.5: Evaluating Overlap Handling for Full-Duplex Speech Models | 2025 | arXiv/Github |
| Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities | 2025 | arXiv/Github |
| MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models | 2025 | arXiv |
| FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems | 2025 | arXiv/Github/Dataset |
| Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics | 2025 | arXiv |
A curated list of full-duplex spoken dialogue models & benchmarks
247
49 commits
updated Sep 21, 2026
A curated list of full-duplex spoken dialogue models.
Welcome to PR if you want to add some resources.
Type — what the entry actually gives you:
| Type | Meaning |
|---|---|
| End-to-end | A standalone full-duplex model: audio in, audio out, duplex behaviour learned inside the model. (omni) means it also takes vision. |
| Cascaded | A complete system assembled from separate parts (streaming ASR + LLM + TTS + a duplex controller). |
| Component | One piece of the puzzle — turn detection, semantic VAD, streaming TTS, retrieval, dialogue management. Needs a host system. |
| Method | A training recipe, reward model, or data scheme rather than a deployable system. |
| Framework | Engineering scaffolding for building/serving duplex systems. |
Open — what has been released:
| Open | Meaning |
|---|---|
| Code + weights | Public repo and released checkpoints. |
| Code | Public repo, no checkpoints released (or none linked). |
| API / closed | Usable as a product or API; nothing released. |
| — | Paper or tech report only. |
| ? | Could not verify. |
Year is the year of first public release (arXiv v1, blog post, or repo). Entries without a dated paper are marked with the year they appeared publicly.
| Title | Year | Open | Relevant Resources |
|---|---|---|---|
| ConversationalVoice: Full-Duplex Speech Data from Real Conversations | 2026 | Code | Github/Demo |
| DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues | 2026 | Data + code | arXiv/Github/Huggingface/Demo |
| DuplexChat | 2026 | Data + code | Github/Huggingface |
| SOMMELIER: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models | 2026 | Code (pipeline) | arXiv/Demo/Github/Blog |
| SmoothConv & DuplexConv: Large-Scale Chinese Full-Duplex Speech Datasets for Conversational AI | 2026 | Data + code | Github/Demo/Huggingface |
| TURNS-2K | 2025 | Data | Github/Huggingface |
| Title | Year | Type | Open | Relevant Resources |
|---|---|---|---|---|
| Realtime-Venus: A full-duplex interaction system with asynchronous delegation | 2026 | End-to-end | Code + weights | arXiv/Github/Huggingface/Demo |
| SteerDuplex: Steerable Duplex Speech Dialogue Models | 2026 | End-to-end | To be released | arXiv/Github |
| Omni Interaction Agent Technical Report | 2026 | Cascaded | Code + weights | arXiv/Github/Huggingface/Demo |
| X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction | 2026 | Component — semantic VAD | Code + weights | arXiv/Github/Huggingface |
| JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents | 2026 | Cascaded | — | arXiv |
| Qwen Audio Agent | 2026 | Cascaded | Code | Github |
| Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models | 2026 | Method | Code | arXiv/Github |
| DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction | 2026 | End-to-end (omni) | Code | arXiv/Github |
| Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs | 2026 | End-to-end | Code + weights | arXiv/Github/Huggingface/Demo |
| BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM | 2026 | End-to-end | Code + weights | arXiv/Github/Huggingface |
| TurnSense and TurnSense 1.1: Three-Class Semantic Turn Detection for Chinese and English Speech Interaction | 2026 | Component — turn detection | Code + weights | Github/Huggingface |
| DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action | 2026 | End-to-end | Code | Github/Demo |
| X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System | 2025 | Cascaded | Code | arXiv/Github/Demo |
| Unmute | 2025 | Cascaded | Code + weights | arXiv/Github/Demo |
| Reinforcement Learning Enhanced Full-Duplex Spoken Dialogue Language Models for Conversational Interactions | 2026 | Method | — | paper |
| MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models | 2026 | Component — retrieval | Code + weights | arXiv/Github/Huggingface |
| Seeduplex: Native Full-Duplex Speech LLM | 2026 | End-to-end | API / closed | Release/Blog |
| FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection | 2026 | Component — turn detection | — | arXiv |
| JAL-Turn: Joint Acoustic-Linguistic Modeling for Real-Time and Robust Turn-Taking Detection in Full-Duplex Spoken Dialogue Systems | 2026 | Component — turn detection | — | arXiv |
| TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving | 2025 | Method | Code | arXiv/Github/Demo |
| Qwen3.5-Omni | 2026 | End-to-end (omni) | API / closed | Official Blog |
| Covo-Audio | 2026 | End-to-end | Code + weights | arXiv/Github/Huggingface |
| SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation | 2026 | Component — semantic VAD | Code + weights | arXiv/Github/Huggingface/Demo |
| PHOENIX-VAD: STREAMING SEMANTIC ENDPOINT DETECTION FOR FULL-DUPLEX SPEECH INTERACTION | 2025 | Component — VAD/endpointing | — | arXiv |
| Turnsense: A Lightweight End-of-Utterance Detection Model | 2025 | Component — turn detection | Code + weights | Github/Huggingface |
| EASY TURN: INTEGRATING ACOUSTIC AND LINGUISTIC MODALITIES FOR ROBUST TURN-TAKING IN FULL-DUPLEX SPOKEN DIALOGUE SYSTEMS | 2025 | Component — turn detection | Code | arXiv/Github/Demo |
| Fun-Audio-Chat | 2025 | End-to-end | Code | arXiv/Github/Demo |
| FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations | 2025 | Cascaded | Code | arXiv/Github/Demo |
| PERSONAPLEX: VOICE AND ROLE CONTROL FOR FULL DUPLEX CONVERSATIONAL SPEECH MODELS | 2026 | End-to-end | Code | arXiv/Github/Demo |
| VITA: Towards Open-Source Interactive Omni Multimodal LLM | 2024 | End-to-end (omni) | Code + weights | arXiv/Github/Demo |
| A Full-duplex Speech Dialogue Scheme Based On Large Language Models | 2024 | End-to-end | — | arXiv |
| Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM | 2024 | End-to-end | Code + weights | arXiv/Github |
| Moshi: a speech-text foundation model for real-time dialogue | 2024 | End-to-end | Code + weights | arXiv/Github |
| FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems | 2025 | Component — duplex control | — | arXiv |
| MinMo: A Multimodal Large Language Model for Seamless Voice Interaction | 2025 | End-to-end | — | arXiv/Demo |
| SoulX-DuoVoice | 2025 | End-to-end | ? | Unofficial Intro |
| SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation | 2025 | End-to-end | Code | arXiv/Github |
| CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction | 2025 | Framework | Code | arXiv/Github |
| Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities | 2024 | End-to-end (omni) | Code + weights | arXiv/Github |
| Real-Time Textless Dialogue Generation | 2025 | End-to-end | Code | arXiv/Github/Demo |
| Language Model Can Listen While Speaking | 2024 | End-to-end | — | arXiv/Demo |
| Parrot: Seamless Spoken Dialogue Interaction with Double-Channel Large Language Models | 2024 | End-to-end | — | Paper |
| Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents | 2024 | End-to-end | — | arXiv/Demo |
| LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems | 2025 | Component — dialogue mgmt | — | arXiv |
| Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems | 2022 | Cascaded | — | arXiv |
| Generative Spoken Dialogue Language Modeling | 2022 | End-to-end | Code + weights | arXiv/Github/Demo |
| Duplex Conversation in Outbound Agent System | 2021 | Cascaded | — | Paper |
| OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation | 2024 | End-to-end | — | arXiv/Demo |
| DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities | 2025 | End-to-end | Code | arXiv/Github |
| Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models | 2024 | Method | Code + data | arXiv/Github/HuggingFace |
| KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI | 2025 | Component — retrieval | Code + weights | arXiv/Github/HuggingFace |
| VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency | 2025 | Component — streaming TTS | Code | arXiv/Github/Demo |
| VoXtream2: Full-stream TTS with dynamic speaking rate control | 2026 | Component — streaming TTS | Code | arXiv/Github/Demo |
| Title | Year | Relevant Resources |
|---|---|---|
| TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue | 2026 | arXiv/Demo/Dataset/Github |
| Game-Time: Evaluating Temporal Dynamics in Spoken Language Models | 2025 | arXiv/Demo/Dataset |
| Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model | 2026 | arXiv/Github |
| Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge | 2026 | arXiv/Github/Dataset |
| Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency | 2026 | arXiv/Github/Demo |
| Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner | 2025 | arXiv/Github |
| FULL-DUPLEX-BENCH V1.5: Evaluating Overlap Handling for Full-Duplex Speech Models | 2025 | arXiv/Github |
| Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities | 2025 | arXiv/Github |
| MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models | 2025 | arXiv |
| FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems | 2025 | arXiv/Github/Dataset |
| Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics | 2025 | arXiv |