dreamtheater123/Awesome-SpeechLM-Survey

Github repository for ACL 2025 paper: Recent Advances in Speech Language Models: A Survey.

225

8 commits

updated Aug 21, 2026

See the code

README

Awesome-SpeechLM-Survey

arXiv

πŸŽ‰πŸŽ‰πŸŽ‰Our survey paper "Recent Advances in Speech Language Models: A Survey" has been accepted to ACL 2025 main conference!

This is the Github repository for paper: Recent Advances in Speech Language Models: A Survey. In this paper, we survey the field of Speech Language Models (SpeechLMs), which are capable of performing end-to-end speech interactions with humans and serve as autoregressive foundation models.

News

Introduction

Why SpeechLMs? SpeechLMs are used for end-to-end speech-based interactions. Traditional ASR + LLM + TTS setups suffer from information loss and cumulative errors during conversion. SpeechLMs directly model speech data, capturing both semantic and paralinguistic information for richer interactions!

In this repository, we distinguish between two closely related model families:

  • Speech Language Models (SpeechLMs) directly accept speech/audio and generate speech/audio. They may operate in turn-based, streaming, or full-duplex settings, but do not rely on an external ASR-LLM-TTS cascade as their core interaction mechanism.
  • Large Audio Language Models (LALMs) use a generative language model to understand speech, environmental sound, or music and produce text. Speech-centric LLM-ASR models are included as a dedicated LALM category, while conventional encoder-only ASR models are outside the scope of this list.

Models supporting both text and speech output are listed under SpeechLMs to avoid duplication.

Intro Figure

Taxonomy

We introduce a novel taxonomy for SpeechLMs, categorizing them based on their architecture and training recipes.

taxonomy

Speech Language Models (SpeechLMs)

The models below support speech/audio input and speech/audio output within the model. The list covers turn-based, streaming, and full-duplex interaction.

ModelPaper / ReleaseYearInteraction
Lychee-FDHierarchical Acoustic-Semantic Modeling for Full-Duplex Speech Interaction2026Full-duplex
DuplexSLADuplexSLA Β· GitHub2026Full-duplex speech-language-action
MiniCPM-o 4.5MiniCPM-o 4.52026Full-duplex omni-modal
Qwen3.5-OmniQwen3.5-Omni Technical Report2026Streaming omni-modal
PersonaPlexPersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models Β· GitHub2026Full-duplex
F-ActorF-Actor: Controllable Full-Duplex Conversational Speech Generation2026Full-duplex
MiMo-AudioMiMo-Audio Technical Report2025Spoken dialogue and audio generation
LFM2-AudioLFM2 Technical Report2025Real-time speech-to-speech
MOSS-SpeechMOSS-Speech2025Native speech-to-speech
Qwen3-OmniQwen3-Omni Technical Report2025Streaming omni-modal
TurnGuideEnhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving Β· GitHub2025Full-duplex
Step-Audio 2Step-Audio 2 Technical Report2025End-to-end speech interaction
Audio Flamingo 3Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models2025General audio and voice-to-voice
OpenS2SOpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model2025End-to-end speech-to-speech
NTPPGenerative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction2025Dual-channel dialogue
SALMONN-omniA Standalone Speech LLM without Codec Injection for Full-duplex Conversation2025Full-duplex
VITA-AudioFast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model2025Streaming speech interaction
VoilaVoice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play2025Real-time interaction
LLaMA-Omni 2LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis2025Streaming speech-to-speech
Kimi-AudioKimi-Audio Technical Report2025Audio understanding and generation
Qwen2.5-OmniQwen2.5-Omni Technical Report2025Streaming omni-modal
SlammingTraining a Speech Language Model on One GPU in a Day2025Speech-to-speech
Baichuan-AudioA Unified Framework for End-to-End Speech Interaction2025End-to-end speech interaction
Step-AudioUnified Understanding and Generation in Intelligent Speech Interaction2025Speech interaction
MinMoA Multimodal Large Language Model for Seamless Voice Interaction2025Streaming voice interaction
VITA-1.5Towards GPT-4o Level Real-Time Vision and Speech Interaction2025Real-time omni-modal
MiniCPM-oA GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming on Your Phone2025Streaming omni-modal
SLAM-OmniTimbre-Controllable Voice Interaction System with Single-Stage Training2024Voice interaction
LyraAn Efficient and Speech-Centric Framework for Omni-Cognition2024Speech interaction
Flow-OmniContinuous Speech Tokens Makes LLMs Robust Multi-Modality Learners2024Speech understanding and generation
GLM-4-VoiceTowards Intelligent and Human-Like End-to-End Spoken Chatbot2024End-to-end speech interaction
Scaling Speech-Text Pre-trainingScaling Speech-Text Pre-training with Synthetic Interleaved Data2024Speech-text generation
Freeze-OmniA Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM2024Streaming speech-to-speech
OmniFlattenAn End-to-end GPT Model for Seamless Voice Conversation2024End-to-end voice interaction
Mini-Omni2Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities2024Duplex omni-modal
IntrinsicVoiceEmpowering LLMs with Intrinsic Real-time Voice Interaction Abilities2024Real-time voice interaction
MoshiA Speech-Text Foundation Model for Real-Time Dialogue2024Full-duplex
SyncLLMBeyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents2024Full-duplex
EMOVAEmpowering Language Models to See, Hear and Speak with Vivid Emotions2024Emotional speech interaction
LLaMA-OmniSeamless Speech Interaction with Large Language Models2024Streaming speech-to-speech
Mini-OmniLanguage Models Can Hear, Talk While Thinking in Streaming2024Streaming speech-to-speech
VITATowards Open-Source Interactive Omni Multimodal LLM2024Real-time omni-modal
LSLMLanguage Model Can Listen While Speaking2024Full-duplex
Spirit-LMInterleaved Spoken and Written Language Model2024Interleaved speech-text generation
SpeechGPT-GenScaling Chain-of-Information Speech Generation2024Speech generation
GPSTGenerative Pre-trained Speech Language Model with Efficient Hierarchical Transformer2024Speech generation and continuation
ParrotAutoregressive Spoken Dialogue Language Modeling with Decoder-only Transformers2023Spoken dialogue modeling
SUTLMToward Joint Language Modeling for Speech Units and Text2023Speech-text generation
tGSLMGenerative Spoken Language Model Based on Continuous Word-Sized Audio Tokens2023Spoken language generation
LauraGPTListen, Attend, Understand, and Regenerate Audio with GPT2023Audio understanding and regeneration
SpectronSpoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM2023Spoken QA and continuation
VoxtLMUnified Decoder-Only Models for Speech Recognition, Synthesis, and Continuation2023Speech-text generation
TWISTTextually Pretrained Speech Language Models2023Spoken language modeling
AudioPaLMA Large Language Model That Can Speak and Listen2023Speech understanding and generation
VioLAUnified Codec Language Models for Speech Recognition, Synthesis, and Translation2023Speech understanding and generation
SpeechGPTEmpowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities2023Speech conversation
dGSLMGenerative Spoken Dialogue Language Modeling2023Spoken dialogue modeling
pGSLMText-Free Prosody-Aware Generative Spoken Language Modeling2021Prosody-aware spoken language modeling
GSLMOn Generative Spoken Language Modeling from Raw Audio2021Spoken language modeling

Commercial and Closed Voice Systems

These entries are tracked separately because public information may describe a product or API rather than disclose a standalone end-to-end model architecture.

SystemReleaseYearNotes
OpenAI GPT-LiveIntroducing GPT-Live2026Full-duplex voice model
OpenAI GPT-Realtime-2Advancing Voice Intelligence with New Models in the API2026Native audio input and output
OpenAI gpt-realtimeIntroducing gpt-realtime2025Native audio input and output
Gemini Native Audio / Live APILive API Capabilities2025Native real-time audio interaction
OpenAI Advanced Voice ModeAdvanced Voice Mode FAQ2024Commercial voice experience
Claude Voice ModeUsing Voice Mode on Claude Mobile Apps2025Product feature; native end-to-end architecture not publicly documented
MindGPT-4o-AudioMindGPT-4o-Audio Release2025Commercial real-time voice model

SpeechLM Training and Alignment

These papers improve the interaction behavior of full-duplex SpeechLMs rather than introducing a separate foundation-model family.

PaperYearFocus
Aligning Spoken Dialogue Models from User Interactions2025Preference alignment from user interactions
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models2026Reinforcement learning for full-duplex interactivity

Large Audio Language Models (LALMs)

LALMs accept speech, environmental sound, and/or music and generate text. Models that additionally generate speech are listed only in the SpeechLM section above.

General Audio Understanding

ModelPaper / ReleaseYearAudio Scope
MOSS-AudioMOSS-Audio2026Speech, sound, and music
Audio Flamingo NextAudio Flamingo Next2026Speech, sound, music, and long audio
VoxtralVoxtral2025Speech and long audio
Gemma 3nGemma 3n2025Speech and multimodal understanding
Audio Flamingo 2Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities2025Speech, sound, music, and long audio
Phi-4-MultimodalPhi-4-Multimodal Technical Report2025Speech, audio, vision, and text
Qwen2-AudioQwen2-Audio Technical Report2024Speech, sound, and music
GAMAA Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities2024Sound and music
WavLLMTowards Robust and Adaptive Speech Large Language Model2024Speech and paralinguistics
Audio FlamingoA Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities2024General audio
Qwen-AudioAdvancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models2023Speech, sound, music, and song
SALMONNTowards Generic Hearing Abilities for Large Language Models2023Speech, sound, and music
LTU-ASJoint Audio and Speech Understanding2023Speech, paralinguistics, and sound
PengiAn Audio Language Model for Audio Tasks2023General audio
LTUAn Audio-Enhanced Large Language Model for Long-Tail Audio Understanding2023General audio

Audio Reasoning, Music, and Prosody

ModelPaperYearFocus
TextPro-SLMMinimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM2026Speech and prosody understanding
Audio-CogitoAudio-Cogito2026Deep audio reasoning
Step-Audio-R1Step-Audio-R12025Speech, sound, and music reasoning
Music FlamingoMusic Flamingo2025Music understanding and reasoning
Audio-ReasonerImproving Reasoning Capability in Large Audio Language Models2025Chain-of-thought audio reasoning
LLarkA Multimodal Instruction-Following Language Model for Music2023Music understanding and reasoning
MusiLingoBridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response2023Music captioning and QA
MU-LLaMAMusic Understanding Large Language Model2023Music understanding

Speech-Centric LALMs and LLM-ASR

This category contains speech-conditioned generative language models whose primary output is text. Conventional encoder-only ASR systems are excluded.

ModelPaper / ReleaseYearFocus
FireRedASR2SA Fully Open-Source Industrial-Grade All-in-One Speech Recognition System2026Multilingual ASR
Qwen3-ASRQwen3-ASR Technical Report2026Multilingual and multidialect ASR
VibeVoice-ASRVibeVoice-ASR Technical Report2026Long-form ASR, diarization, and timestamps
GLM-ASR-NanoGLM-ASR2025LLM-based ASR
Fun-ASRFun-ASR: A Large Language Model Based Speech Recognition System2025Multilingual speech-to-text
FireRedASROpen-Source Industrial-Grade Speech Recognition Models from Encoder-Decoder to LLM Integration2025ASR and LLM integration
Seed-ASRUnderstanding Diverse Speech and Contexts with LLM-based Speech Recognition2024Context-aware LLM-ASR

The following works are closely related but are not included in the strict SpeechLM or LALM lists because they do not provide a unified speech-input/speech-output dialogue model or an audio-conditioned generative text LLM.

Model / FamilyPaper / ReleaseScope
CSMConversational Speech Generation ModelConversational speech generation / TTS
FunAudioLLMVoice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMsSenseVoice and CosyVoice model family
UniAudioAn Audio Foundation Model Toward Universal Audio GenerationUniversal audio generation
VoiceboxText-Guided Multilingual Universal Speech Generation at ScaleText-guided speech generation

SpeechLM Tokenizers

Semantic Tokenizers

NameTitleUrl
WhisperRobust Speech Recognition via Large-Scale Weak SupervisionLink
CosyVoiceCosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic TokensLink
Google USMGoogle USM: Scaling Automatic Speech Recognition Beyond 100 LanguagesLink
WavLMWavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech ProcessingLink
HuBERTHuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden UnitsLink
W2v-bertW2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-TrainingLink
Wav2vec 2.0wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsLink

Acoustic Tokenizers

NameTitleUrl
WavTokenizerWavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingLink
SNACSNAC: Multi-Scale Neural Audio CodecLink
EncodecHigh Fidelity Neural Audio CompressionLink
SoundStreamSoundStream: An End-to-End Neural Audio CodecLink

Mixed Tokenizers

NameTitleUrl
SpeechTokenizerSpeechTokenizer: Unified Speech Tokenizer for Speech Large Language ModelsLink
MimiMoshi: a speech-text foundation model for real-time dialogueLink
DatasetTypePhaseHoursYear
LibriSpeechASRPre-Training1k2015
Multilingual LibriSpeechASRPre-Training50.5k2020
LibriLightASRPre-Training60k2019
People datasetASRPre-Training30k2021
VoxPopuliASRPre-Training1.6k2021
GigaspeechASRPre-Training40k2021
Common VoiceASRPre-Training2.5k2019
VCTKASRPre-Training0.3k2017
WenetSpeechASRPre-Training22k2022
LibriTTSTTSPre-Training0.6k2019
CoVoST2S2TTPre-Training2.8k2020
CVSSS2STPre-Training1.9k2022
VoxCelebSpeaker IdentificationPre-Training0.4k2017
VoxCeleb2Speaker IdentificationPre-Training2.4k2018
Spotify PodcastsPodcastPre-Training47k2020
FisherTelephone conversationPre-Training2k2004
SpeechInstructInstruction-followingInstruction-Tuning-2023
InstructS2S-200KInstruction-followingInstruction-Tuning-2024
VoiceAssistant-400KInstruction-followingInstruction-Tuning-2024

Evaluation Benchmarks

NameEval Type# TasksAudio TypeI/O
ABXRepresentation1Speech$A \rightarrow -$
sWUGGYLinguistic1Speech$A \rightarrow -$
sBLIMPLinguistic1Speech$A \rightarrow -$
sStoryClozeLinguistic1Speech$A/T \rightarrow -$
STSPParalinguistic1Speech$A/T \rightarrow A/T$
MMAUDownstream27Speech, Sound, Music$A \rightarrow T$
AudiobenchDownstream8Speech, Sound$A \rightarrow T$
AIR-BenchDownstream20Speech, Sound, Music$A \rightarrow T$
SD-EvalDownstream4Speech$A \rightarrow T$
SUPERBDownstream10Speech$A \rightarrow T$
Dynamic-SUPERBDownstream180Speech, Sound, Music$A \rightarrow T$
SALMONDownstream8Speech$A \rightarrow -$
VoiceBenchDownstream8Speech$A \rightarrow A$
VoxEvalDownstream56Speech$A \rightarrow A$

Citation

@article{cui2024recent,
  title={Recent advances in speech language models: A survey},
  author={Cui, Wenqian and Yu, Dianzhi and Jiao, Xiaoqi and Meng, Ziqiao and Zhang, Guangyan and Wang, Qichao and Guo, Yiwen and King, Irwin},
  journal={arXiv preprint arXiv:2410.03751},
  year={2024}
}

Contributors

dreamtheater123/Awesome-SpeechLM-Survey

Github repository for ACL 2025 paper: Recent Advances in Speech Language Models: A Survey.

225

8 commits

updated Aug 21, 2026

See the code

README

Awesome-SpeechLM-Survey

arXiv

πŸŽ‰πŸŽ‰πŸŽ‰Our survey paper "Recent Advances in Speech Language Models: A Survey" has been accepted to ACL 2025 main conference!

This is the Github repository for paper: Recent Advances in Speech Language Models: A Survey. In this paper, we survey the field of Speech Language Models (SpeechLMs), which are capable of performing end-to-end speech interactions with humans and serve as autoregressive foundation models.

News

Introduction

Why SpeechLMs? SpeechLMs are used for end-to-end speech-based interactions. Traditional ASR + LLM + TTS setups suffer from information loss and cumulative errors during conversion. SpeechLMs directly model speech data, capturing both semantic and paralinguistic information for richer interactions!

In this repository, we distinguish between two closely related model families:

  • Speech Language Models (SpeechLMs) directly accept speech/audio and generate speech/audio. They may operate in turn-based, streaming, or full-duplex settings, but do not rely on an external ASR-LLM-TTS cascade as their core interaction mechanism.
  • Large Audio Language Models (LALMs) use a generative language model to understand speech, environmental sound, or music and produce text. Speech-centric LLM-ASR models are included as a dedicated LALM category, while conventional encoder-only ASR models are outside the scope of this list.

Models supporting both text and speech output are listed under SpeechLMs to avoid duplication.

Intro Figure

Taxonomy

We introduce a novel taxonomy for SpeechLMs, categorizing them based on their architecture and training recipes.

taxonomy

Speech Language Models (SpeechLMs)

The models below support speech/audio input and speech/audio output within the model. The list covers turn-based, streaming, and full-duplex interaction.

ModelPaper / ReleaseYearInteraction
Lychee-FDHierarchical Acoustic-Semantic Modeling for Full-Duplex Speech Interaction2026Full-duplex
DuplexSLADuplexSLA Β· GitHub2026Full-duplex speech-language-action
MiniCPM-o 4.5MiniCPM-o 4.52026Full-duplex omni-modal
Qwen3.5-OmniQwen3.5-Omni Technical Report2026Streaming omni-modal
PersonaPlexPersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models Β· GitHub2026Full-duplex
F-ActorF-Actor: Controllable Full-Duplex Conversational Speech Generation2026Full-duplex
MiMo-AudioMiMo-Audio Technical Report2025Spoken dialogue and audio generation
LFM2-AudioLFM2 Technical Report2025Real-time speech-to-speech
MOSS-SpeechMOSS-Speech2025Native speech-to-speech
Qwen3-OmniQwen3-Omni Technical Report2025Streaming omni-modal
TurnGuideEnhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving Β· GitHub2025Full-duplex
Step-Audio 2Step-Audio 2 Technical Report2025End-to-end speech interaction
Audio Flamingo 3Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models2025General audio and voice-to-voice
OpenS2SOpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model2025End-to-end speech-to-speech
NTPPGenerative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction2025Dual-channel dialogue
SALMONN-omniA Standalone Speech LLM without Codec Injection for Full-duplex Conversation2025Full-duplex
VITA-AudioFast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model2025Streaming speech interaction
VoilaVoice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play2025Real-time interaction
LLaMA-Omni 2LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis2025Streaming speech-to-speech
Kimi-AudioKimi-Audio Technical Report2025Audio understanding and generation
Qwen2.5-OmniQwen2.5-Omni Technical Report2025Streaming omni-modal
SlammingTraining a Speech Language Model on One GPU in a Day2025Speech-to-speech
Baichuan-AudioA Unified Framework for End-to-End Speech Interaction2025End-to-end speech interaction
Step-AudioUnified Understanding and Generation in Intelligent Speech Interaction2025Speech interaction
MinMoA Multimodal Large Language Model for Seamless Voice Interaction2025Streaming voice interaction
VITA-1.5Towards GPT-4o Level Real-Time Vision and Speech Interaction2025Real-time omni-modal
MiniCPM-oA GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming on Your Phone2025Streaming omni-modal
SLAM-OmniTimbre-Controllable Voice Interaction System with Single-Stage Training2024Voice interaction
LyraAn Efficient and Speech-Centric Framework for Omni-Cognition2024Speech interaction
Flow-OmniContinuous Speech Tokens Makes LLMs Robust Multi-Modality Learners2024Speech understanding and generation
GLM-4-VoiceTowards Intelligent and Human-Like End-to-End Spoken Chatbot2024End-to-end speech interaction
Scaling Speech-Text Pre-trainingScaling Speech-Text Pre-training with Synthetic Interleaved Data2024Speech-text generation
Freeze-OmniA Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM2024Streaming speech-to-speech
OmniFlattenAn End-to-end GPT Model for Seamless Voice Conversation2024End-to-end voice interaction
Mini-Omni2Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities2024Duplex omni-modal
IntrinsicVoiceEmpowering LLMs with Intrinsic Real-time Voice Interaction Abilities2024Real-time voice interaction
MoshiA Speech-Text Foundation Model for Real-Time Dialogue2024Full-duplex
SyncLLMBeyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents2024Full-duplex
EMOVAEmpowering Language Models to See, Hear and Speak with Vivid Emotions2024Emotional speech interaction
LLaMA-OmniSeamless Speech Interaction with Large Language Models2024Streaming speech-to-speech
Mini-OmniLanguage Models Can Hear, Talk While Thinking in Streaming2024Streaming speech-to-speech
VITATowards Open-Source Interactive Omni Multimodal LLM2024Real-time omni-modal
LSLMLanguage Model Can Listen While Speaking2024Full-duplex
Spirit-LMInterleaved Spoken and Written Language Model2024Interleaved speech-text generation
SpeechGPT-GenScaling Chain-of-Information Speech Generation2024Speech generation
GPSTGenerative Pre-trained Speech Language Model with Efficient Hierarchical Transformer2024Speech generation and continuation
ParrotAutoregressive Spoken Dialogue Language Modeling with Decoder-only Transformers2023Spoken dialogue modeling
SUTLMToward Joint Language Modeling for Speech Units and Text2023Speech-text generation
tGSLMGenerative Spoken Language Model Based on Continuous Word-Sized Audio Tokens2023Spoken language generation
LauraGPTListen, Attend, Understand, and Regenerate Audio with GPT2023Audio understanding and regeneration
SpectronSpoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM2023Spoken QA and continuation
VoxtLMUnified Decoder-Only Models for Speech Recognition, Synthesis, and Continuation2023Speech-text generation
TWISTTextually Pretrained Speech Language Models2023Spoken language modeling
AudioPaLMA Large Language Model That Can Speak and Listen2023Speech understanding and generation
VioLAUnified Codec Language Models for Speech Recognition, Synthesis, and Translation2023Speech understanding and generation
SpeechGPTEmpowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities2023Speech conversation
dGSLMGenerative Spoken Dialogue Language Modeling2023Spoken dialogue modeling
pGSLMText-Free Prosody-Aware Generative Spoken Language Modeling2021Prosody-aware spoken language modeling
GSLMOn Generative Spoken Language Modeling from Raw Audio2021Spoken language modeling

Commercial and Closed Voice Systems

These entries are tracked separately because public information may describe a product or API rather than disclose a standalone end-to-end model architecture.

SystemReleaseYearNotes
OpenAI GPT-LiveIntroducing GPT-Live2026Full-duplex voice model
OpenAI GPT-Realtime-2Advancing Voice Intelligence with New Models in the API2026Native audio input and output
OpenAI gpt-realtimeIntroducing gpt-realtime2025Native audio input and output
Gemini Native Audio / Live APILive API Capabilities2025Native real-time audio interaction
OpenAI Advanced Voice ModeAdvanced Voice Mode FAQ2024Commercial voice experience
Claude Voice ModeUsing Voice Mode on Claude Mobile Apps2025Product feature; native end-to-end architecture not publicly documented
MindGPT-4o-AudioMindGPT-4o-Audio Release2025Commercial real-time voice model

SpeechLM Training and Alignment

These papers improve the interaction behavior of full-duplex SpeechLMs rather than introducing a separate foundation-model family.

PaperYearFocus
Aligning Spoken Dialogue Models from User Interactions2025Preference alignment from user interactions
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models2026Reinforcement learning for full-duplex interactivity

Large Audio Language Models (LALMs)

LALMs accept speech, environmental sound, and/or music and generate text. Models that additionally generate speech are listed only in the SpeechLM section above.

General Audio Understanding

ModelPaper / ReleaseYearAudio Scope
MOSS-AudioMOSS-Audio2026Speech, sound, and music
Audio Flamingo NextAudio Flamingo Next2026Speech, sound, music, and long audio
VoxtralVoxtral2025Speech and long audio
Gemma 3nGemma 3n2025Speech and multimodal understanding
Audio Flamingo 2Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities2025Speech, sound, music, and long audio
Phi-4-MultimodalPhi-4-Multimodal Technical Report2025Speech, audio, vision, and text
Qwen2-AudioQwen2-Audio Technical Report2024Speech, sound, and music
GAMAA Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities2024Sound and music
WavLLMTowards Robust and Adaptive Speech Large Language Model2024Speech and paralinguistics
Audio FlamingoA Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities2024General audio
Qwen-AudioAdvancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models2023Speech, sound, music, and song
SALMONNTowards Generic Hearing Abilities for Large Language Models2023Speech, sound, and music
LTU-ASJoint Audio and Speech Understanding2023Speech, paralinguistics, and sound
PengiAn Audio Language Model for Audio Tasks2023General audio
LTUAn Audio-Enhanced Large Language Model for Long-Tail Audio Understanding2023General audio

Audio Reasoning, Music, and Prosody

ModelPaperYearFocus
TextPro-SLMMinimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM2026Speech and prosody understanding
Audio-CogitoAudio-Cogito2026Deep audio reasoning
Step-Audio-R1Step-Audio-R12025Speech, sound, and music reasoning
Music FlamingoMusic Flamingo2025Music understanding and reasoning
Audio-ReasonerImproving Reasoning Capability in Large Audio Language Models2025Chain-of-thought audio reasoning
LLarkA Multimodal Instruction-Following Language Model for Music2023Music understanding and reasoning
MusiLingoBridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response2023Music captioning and QA
MU-LLaMAMusic Understanding Large Language Model2023Music understanding

Speech-Centric LALMs and LLM-ASR

This category contains speech-conditioned generative language models whose primary output is text. Conventional encoder-only ASR systems are excluded.

ModelPaper / ReleaseYearFocus
FireRedASR2SA Fully Open-Source Industrial-Grade All-in-One Speech Recognition System2026Multilingual ASR
Qwen3-ASRQwen3-ASR Technical Report2026Multilingual and multidialect ASR
VibeVoice-ASRVibeVoice-ASR Technical Report2026Long-form ASR, diarization, and timestamps
GLM-ASR-NanoGLM-ASR2025LLM-based ASR
Fun-ASRFun-ASR: A Large Language Model Based Speech Recognition System2025Multilingual speech-to-text
FireRedASROpen-Source Industrial-Grade Speech Recognition Models from Encoder-Decoder to LLM Integration2025ASR and LLM integration
Seed-ASRUnderstanding Diverse Speech and Contexts with LLM-based Speech Recognition2024Context-aware LLM-ASR

The following works are closely related but are not included in the strict SpeechLM or LALM lists because they do not provide a unified speech-input/speech-output dialogue model or an audio-conditioned generative text LLM.

Model / FamilyPaper / ReleaseScope
CSMConversational Speech Generation ModelConversational speech generation / TTS
FunAudioLLMVoice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMsSenseVoice and CosyVoice model family
UniAudioAn Audio Foundation Model Toward Universal Audio GenerationUniversal audio generation
VoiceboxText-Guided Multilingual Universal Speech Generation at ScaleText-guided speech generation

SpeechLM Tokenizers

Semantic Tokenizers

NameTitleUrl
WhisperRobust Speech Recognition via Large-Scale Weak SupervisionLink
CosyVoiceCosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic TokensLink
Google USMGoogle USM: Scaling Automatic Speech Recognition Beyond 100 LanguagesLink
WavLMWavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech ProcessingLink
HuBERTHuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden UnitsLink
W2v-bertW2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-TrainingLink
Wav2vec 2.0wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsLink

Acoustic Tokenizers

NameTitleUrl
WavTokenizerWavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language ModelingLink
SNACSNAC: Multi-Scale Neural Audio CodecLink
EncodecHigh Fidelity Neural Audio CompressionLink
SoundStreamSoundStream: An End-to-End Neural Audio CodecLink

Mixed Tokenizers

NameTitleUrl
SpeechTokenizerSpeechTokenizer: Unified Speech Tokenizer for Speech Large Language ModelsLink
MimiMoshi: a speech-text foundation model for real-time dialogueLink
DatasetTypePhaseHoursYear
LibriSpeechASRPre-Training1k2015
Multilingual LibriSpeechASRPre-Training50.5k2020
LibriLightASRPre-Training60k2019
People datasetASRPre-Training30k2021
VoxPopuliASRPre-Training1.6k2021
GigaspeechASRPre-Training40k2021
Common VoiceASRPre-Training2.5k2019
VCTKASRPre-Training0.3k2017
WenetSpeechASRPre-Training22k2022
LibriTTSTTSPre-Training0.6k2019
CoVoST2S2TTPre-Training2.8k2020
CVSSS2STPre-Training1.9k2022
VoxCelebSpeaker IdentificationPre-Training0.4k2017
VoxCeleb2Speaker IdentificationPre-Training2.4k2018
Spotify PodcastsPodcastPre-Training47k2020
FisherTelephone conversationPre-Training2k2004
SpeechInstructInstruction-followingInstruction-Tuning-2023
InstructS2S-200KInstruction-followingInstruction-Tuning-2024
VoiceAssistant-400KInstruction-followingInstruction-Tuning-2024

Evaluation Benchmarks

NameEval Type# TasksAudio TypeI/O
ABXRepresentation1Speech$A \rightarrow -$
sWUGGYLinguistic1Speech$A \rightarrow -$
sBLIMPLinguistic1Speech$A \rightarrow -$
sStoryClozeLinguistic1Speech$A/T \rightarrow -$
STSPParalinguistic1Speech$A/T \rightarrow A/T$
MMAUDownstream27Speech, Sound, Music$A \rightarrow T$
AudiobenchDownstream8Speech, Sound$A \rightarrow T$
AIR-BenchDownstream20Speech, Sound, Music$A \rightarrow T$
SD-EvalDownstream4Speech$A \rightarrow T$
SUPERBDownstream10Speech$A \rightarrow T$
Dynamic-SUPERBDownstream180Speech, Sound, Music$A \rightarrow T$
SALMONDownstream8Speech$A \rightarrow -$
VoiceBenchDownstream8Speech$A \rightarrow A$
VoxEvalDownstream56Speech$A \rightarrow A$

Citation

@article{cui2024recent,
  title={Recent advances in speech language models: A survey},
  author={Cui, Wenqian and Yu, Dianzhi and Jiao, Xiaoqi and Meng, Ziqiao and Zhang, Guangyan and Wang, Qichao and Guo, Yiwen and King, Irwin},
  journal={arXiv preprint arXiv:2410.03751},
  year={2024}
}

Contributors