kadirnar/awesome-omni-architectures

Visual catalog of omni, speech and audio language models with public code and weights, short descriptions, architecture figures, license labels and VoiceBench coverage.

Python

14

9 commits

updated Sep 13, 2026

See the code

README

Awesome Omni Architectures Awesome

A visual catalog of omni, speech and audio language models with public code and weights. Every entry has a short description, a diagram, code and checkpoint links.

140 models + 4 cascaded systems · Reviewed 2026-09-13

Model list · All diagrams · VoiceBench coverage · Research · Timeline · Scope

Licenses: 68 open · 48 custom/restricted · 28 unclear. Downloadable weights do not always mean an unrestricted open-source license; each card shows the applicable terms.

Models

T: text · I: image · V: video · A: audio · S: speech · M: music · X: other. Scope and labels.

Alphabetical model list · 144 entries
ModelGroupInput → outputRelease
AnyGPTOmniT, I, S, M → T, I, S, MLicense unclear
Audio FlamingoAudioT, A → TCustom / restricted
Audio Flamingo 2AudioT, A → TCustom / restricted
Audio Flamingo 3 ChatSpeechT, A → T, SCustom / restricted
Audio Flamingo NextAudioT, A → TCustom / restricted
Audio-InteractionAudioT, A → TLicense unclear
Audio-ReasonerAudioT, A → TOpen license
Audio-Visual FlamingoOmniT, I, V, A → TCustom / restricted
AzeroSAudioT, S → TOpen license
Baichuan-AudioSpeechT, A → T, SOpen license
Baichuan-Omni 1.5OmniT, I, V, A → T, SCustom / restricted
BLSP-EmoAudioT, S → TCustom / restricted
BR-Voice-ReasonerOmniT, I, V, A → TOpen license
Canary-QwenLLM-ASRT, S → TOpen license
Covo-Audio ChatSpeechT, S → T, SCustom / restricted
DeSTA2AudioT, A → TLicense unclear
DeSTA2.5-AudioAudioT, A → TLicense unclear
DialectS2SSpeechT, S → T, SOpen license
DIFFAAudioT, A → TLicense unclear
DIFFA-2AudioT, A → TLicense unclear
DiVAAudioT, S → TCustom / restricted
DuplexCascadeSpeechT, S → T, SOpen license
DuplexMambaAudioT, S → TOpen license
EMOVA (Qwen2.5-7B)OmniT, I, S → T, SOpen license
Freeze-OmniSpeechT, S → T, SCustom / restricted
Fun-Audio-ChatSpeechT, S → T, SOpen license
Gemma 3n E4BOmniT, I, V, A → TCustom / restricted
Gemma 4 E4BOmniT, I, V, A → TOpen license
GLM-4-VoiceSpeechT, S → T, SCustom / restricted
GLM-ASR NanoLLM-ASRS → TOpen license
Granite Speech 3.3LLM-ASRT, S → TOpen license
Granite Speech 4.0LLM-ASRT, S → TOpen license
Granite Speech 4.1LLM-ASRT, S → TOpen license
HibikiSpeechT, S → T, SOpen license
HumanOmniOmniT, V, A → TLicense unclear
HumanOmniV2OmniT, V, A → TOpen license
IchigoAudioT, S → TLicense unclear
Kimi-AudioSpeechT, A → T, SOpen license
LFG-1OmniT, I, S → TOpen license
LFG-2AudioT, S → TLicense unclear
LFG-3AudioT, S → TLicense unclear
LFM2-Audio-1.5BSpeechT, S → T, SCustom / restricted
LFM2.5-Audio-1.5BSpeechT, S → T, SCustom / restricted
LFM2.5-Audio-1.5B-JPSpeechT, S → T, SCustom / restricted
LLaMA-OmniSpeechT, S → T, SCustom / restricted
LLaMA-Omni 2SpeechT, S → T, SCustom / restricted
LLaSMAudioT, S → TCustom / restricted
LongCat-Flash-OmniOmniT, I, V, A → T, SOpen license
LongCat-NextOmniT, I, A → T, I, AOpen license
Lychee-FDSpeechT, S → T, SOpen license
Lyra-BaseOmniT, I, V, A → T, SCustom / restricted
Lyra-MiniOmniT, I, V, A → T, SCustom / restricted
Mair-hub-0.5B-OmniSpeechT, S → T, SLicense unclear
Megrez-OmniOmniT, I, A → TOpen license
MellowAudioT, A → TOpen license
MERaLiON 2AudioT, A → TCustom / restricted
MERaLiON 3AudioT, S → TCustom / restricted
MERaLiON 3 ASRLLM-ASRS → TCustom / restricted
MERaLiON-AudioLLMAudioT, A → TCustom / restricted
MiDashengLMAudioT, A → TLicense unclear
MiMo-AudioSpeechT, A → T, SOpen license
MiMo-V2.5OmniT, I, V, A → TOpen license
Ming-Flash-OmniOmniT, I, V, A → T, I, AOpen license
Ming-Flash-Omni 2.0OmniT, I, V, A → T, I, AOpen license
Ming-Lite-OmniOmniT, I, V, A → T, I, SOpen license
Mini-OmniSpeechT, S → T, SOpen license
Mini-Omni2OmniT, I, S → T, SOpen license
MiniCPM-o 2.6OmniT, I, V, A → T, SOpen license
MiniCPM-o 4.5OmniT, I, V, A → T, SOpen license
MiniMind-OOmniT, I, S → T, SOpen license
MIOOmniT, I, V, A → T, I, SLicense unclear
MoshiSpeechT, S → T, SOpen license
MOSS-Audio InstructAudioT, A → TLicense unclear
MOSS-Audio ThinkingAudioT, A → TLicense unclear
MOSS-SpeechSpeechS → SLicense unclear
Music FlamingoAudioT, M → TCustom / restricted
Nemotron 3 Nano OmniOmniT, I, V, A → TCustom / restricted
Nemotron-Labs-Audex 2BSpeechT, A → T, ALicense unclear
Nemotron-Labs-Audex 30B-A3BSpeechT, A → T, ALicense unclear
NemotronLabs VoiceChat 11BSpeechT, S → T, SOpen license
NExT-GPTOmniT, I, V, A → T, I, V, ACustom / restricted
Nexus-OOmniT, I, V, A → T, SCustom / restricted
OlaOmniT, I, V, A → TOpen license
Omni-AutoThinkOmniT, I, V, A → TOpen license
Omni-R1 (audio reasoning)AudioT, A → TLicense unclear
Omni-R1 (two-system collaboration)OmniT, I, V, A → TLicense unclear
OmniVinciOmniT, I, V, A → TCustom / restricted
OneLLMOmniT, I, V, A → TCustom / restricted
OpenOmniOmniT, I, V, A → T, SLicense unclear
OpenS2SSpeechT, S → T, SOpen license
OpenS2S 1.5SpeechT, S → T, SOpen license
OSUMAudioT, S → TOpen license
OSUM-EChatSpeechT, S → T, SOpen license
ParaBridgeAudioT, A → TOpen license
Parakeet-TDT-v2 + Qwen3-8BCascadeS → TOpen license
PersonaPlexSpeechT, S → T, SCustom / restricted
Phi-4-multimodalOmniT, I, A → TOpen license
Qwen-AudioAudioT, A → TCustom / restricted
Qwen2-AudioAudioT, A → TOpen license
Qwen2.5-OmniOmniT, I, V, A → T, SOpen license
Qwen3-ASRLLM-ASRT, S → TOpen license
Qwen3-Omni InstructOmniT, I, V, A → T, SOpen license
Qwen3-Omni ThinkingOmniT, I, V, A → TOpen license
R1-OmniOmniT, V, A → TLicense unclear
SALMONNAudioT, A → TCustom / restricted
SALMONN-2AudioT, A → TOpen license
SLAM-OmniSpeechT, S → T, SOpen license
SoulX-DuplugAudioT, S → TOpen license
SpeechGPTSpeechT, S → T, SLicense unclear
SpeechGPT 2.0 PreviewSpeechT, S → T, SLicense unclear
Step-AudioSpeechT, A → T, SOpen license
Step-Audio 2 MiniSpeechT, A → T, SOpen license
Step-Audio 2 Mini ThinkAudioT, A → TOpen license
Step-Audio-AQAASpeechT, A → T, SOpen license
Step-Audio-R1AudioT, A → TOpen license
Step-Audio-R1.1SpeechT, A → T, SOpen license
Stream-OmniOmniT, I, S → T, SCustom / restricted
SympatheiaSpeechT, S → T, SLicense unclear
TurnGuideSpeechT, S → T, SLicense unclear
Ultravox 0.4.1 Llama-3.1-8BAudioT, S → TCustom / restricted
Ultravox 0.5 Llama-3.1-8BAudioT, S → TCustom / restricted
Ultravox 0.5 Llama-3.2-1BAudioT, S → TCustom / restricted
Ultravox 0.6 Gemma-3-27BAudioT, S → TCustom / restricted
Ultravox 0.6 Llama-3.1-8BAudioT, S → TCustom / restricted
Ultravox 0.6 Llama-3.3-70BAudioT, S → TCustom / restricted
Ultravox 0.6 Qwen3-32BAudioT, S → TOpen license
Ultravox 0.7 GLM-4.6AudioT, S → TOpen license
Uni-MoE 2.0 OmniOmniT, I, V, A → T, I, ALicense unclear
VibeVoice-ASRLLM-ASRT, S → TOpen license
video-SALMONN 2+ (7B)OmniT, V, A → TOpen license
VideoLLaMA 2.1 AVOmniT, I, V, A → TCustom / restricted
VITA 1.0OmniT, I, V, A → TCustom / restricted
VITA 1.5OmniT, I, V, A → T, SCustom / restricted
VITA-AudioSpeechT, S → T, SCustom / restricted
VITA-Audio PlusSpeechT, S → T, SCustom / restricted
Voila AutonomousSpeechT, S → T, SOpen license
Voila ChatSpeechT, S → T, SOpen license
VoxMindSpeechT, S → T, SLicense unclear
Voxtral MiniAudioT, A → TOpen license
Voxtral RealtimeLLM-ASRS → TOpen license
Voxtral SmallAudioT, A → TOpen license
Whisper-v3-large + Llama-3.1-8BCascadeS → TCustom / restricted
Whisper-v3-turbo + Llama-3.1-8BCascadeS → TCustom / restricted
Whisper-v3-turbo + Llama-3.2-3BCascadeS → TCustom / restricted

Model figures

Figures are credited to their sources. Editorial input/output diagrams are labeled. Credits.

AnyGPT

Understands and generates text, images, speech and music using a shared token vocabulary.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 + Llama 2 terms — Base-model community terms also apply; repository-wide code licensing is not stated.

AnyGPT — Figure 1

Figure 1 · Source

Audio Flamingo

Answers questions about sounds and music and supports audio-grounded dialogue and examples.

Paper · Code · Weights · Details

Custom / restricted: code MIT · weights OPT-IML noncommercial license — Noncommercial terms apply to model weights and relevant bundled components.

Audio Flamingo — Figure 2

Figure 2 · Source

Audio Flamingo 2

Reasons about environmental sounds and music, including long recordings.

Paper · Code · Weights · Details

Custom / restricted: code MIT · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Audio Flamingo 2 — Figure 2

Figure 2 · Source

Audio Flamingo 3 Chat

Discusses speech, sounds and music across multiple recordings and can reply with streaming speech.

Paper · Code · Weights · Details

Custom / restricted: code MIT · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Audio Flamingo 3 Chat — Figure 2

Figure 2 · Source

Audio Flamingo Next

Analyzes long audio recordings and reasons about speech, sounds and music.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Audio Flamingo Next — Figure 3

Figure 3 · Source

Audio-Interaction

Monitors an ongoing audio stream and decides when a text response or proactive intervention is useful.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

Audio-Interaction — Figure 3

Figure 3 · Source

Audio-Reasoner

Reasons about speech, music and environmental sounds to answer complex audio questions.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Audio-Reasoner — Figure 3

Figure 3 · Source

Audio-Visual Flamingo

Explains long videos by reasoning jointly about visual events and their soundtracks.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Audio-Visual Flamingo — Figure 2

Figure 2 · Source

AzeroS

Follows spoken instructions using speech adaptation learned without curated instruction-answer pairs.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

AzeroS — Figure 1

Figure 1 · Source

Baichuan-Audio

Follows spoken instructions and generates text and expressive speech responses.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Baichuan-Audio — Figure 1

Figure 1 · Source

Baichuan-Omni 1.5

Understands images, video and sound and answers through text or speech.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Baichuan model terms — The model card adds commercial-use conditions beyond its Apache metadata.

Baichuan-Omni 1.5 — Figure 2

Figure 2 · Source

BLSP-Emo

Understands vocal emotion and writes empathetic responses to spoken requests.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Apache-2.0 + Qwen terms — The Qwen-7B-Chat backbone retains its custom license.

BLSP-Emo — Figure 2

Figure 2 · Source

BR-Voice-Reasoner

Answers spoken knowledge and reasoning questions using a speech-adapted Qwen3-Omni Thinker.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

BR-Voice-Reasoner — Input/output diagram

Input/output diagram · Source

Canary-Qwen

Transcribes English speech and offers a separate text-language mode for follow-up processing.

Code · Weights · Details

Open license: code Apache-2.0 · weights CC-BY-4.0

Canary-Qwen — Input/output diagram

Input/output diagram · Source

Covo-Audio Chat

Understands spoken requests and generates context-aware, expressive voice responses.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

Covo-Audio Chat — Figure 2, PDF p. 4

Figure 2, PDF p. 4 · Source

DeSTA2

Describes and answers questions about audio using a language model aligned with speech representations.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Llama community / component terms — Base-model and dataset terms apply; no unrestricted standalone weight license is stated.

DeSTA2 — Figure 1

Figure 1 · Source

DeSTA2.5-Audio

Performs general audio understanding and follows spoken instructions without an external transcript pipeline.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Llama community / component terms — Base-model and dataset terms apply; no unrestricted standalone weight license is stated.

DeSTA2.5-Audio — Figure 2

Figure 2 · Source

DialectS2S

Understands and speaks low-resource Chinese dialects in end-to-end voice conversations.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights MIT

DialectS2S — Family figure — Figure 1

Family figure — Figure 1 · Source

DIFFA

Answers audio questions using diffusion-based text generation instead of left-to-right decoding.

Paper · Code · Weights · Details

License unclear: code Not stated · weights CC-BY-NC-SA-4.0 — The checkpoint uses a noncommercial share-alike license; repository code terms are not stated.

DIFFA — Figure 2

Figure 2 · Source

DIFFA-2

Understands speech, sound and music with a diffusion language model and richer acoustic representations.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Not stated — No release-specific code or weight license was located; do not assume DIFFA 1 terms apply.

DIFFA-2 — Figure 1

Figure 1 · Source

DiVA

Answers spoken questions by transferring a text assistant's behavior into an audio-input model.

Code · Weights · Details

Custom / restricted: code MPL-2.0 · weights MPL-2.0 + Llama 3 terms — The Llama 3 backbone retains its community license.

DiVA — Input/output diagram

Input/output diagram · Source

DuplexCascade

Coordinates streaming transcription, language generation and speech synthesis for interruptible voice conversations.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

DuplexCascade — Figure 1

Figure 1 · Source

DuplexMamba

Reads incoming speech incrementally and supports responsive text generation during speech interaction.

Paper · Code · Weights · Details

Open license: code GPL-3.0 · weights Apache-2.0

DuplexMamba — Figure 1

Figure 1 · Source

EMOVA (Qwen2.5-7B)

Discusses images and spoken requests, replying with text and expressive speech.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

EMOVA (Qwen2.5-7B) — Family figure · Figure 2

Family figure · Figure 2 · Source

Freeze-Omni

Enables streaming spoken dialogue while retaining a frozen text model's language abilities.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

Freeze-Omni — Figure 1

Figure 1 · Source

Fun-Audio-Chat

Holds low-latency spoken conversations while responding to vocal style and emotion.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Fun-Audio-Chat — Figure 2, PDF p. 4

Figure 2, PDF p. 4 · Source

Gemma 3n E4B

Understands text, images, video and audio in a model designed for mobile devices.

Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Gemma terms — The weights use custom Gemma terms.

Gemma 3n E4B — Input/output diagram

Input/output diagram · Source

Gemma 4 E4B

Answers questions about images, video and audio in a compact model designed for local use.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Gemma 4 E4B — Input/output diagram

Input/output diagram · Source

GLM-4-Voice

Converses in Chinese and English with controllable emotion, speaking rate and vocal style.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights GLM-4 model license — Custom GLM model terms apply to the speech checkpoint.

GLM-4-Voice — Figure 2

Figure 2 · Source

GLM-ASR Nano

Transcribes speech in English, Mandarin and Chinese dialects, including noisy recordings.

Code · Weights · Details

Open license: code MIT · weights MIT

GLM-ASR Nano — Input/output diagram

Input/output diagram · Source

Granite Speech 3.3

Transcribes and translates speech while retaining a text language-model interface.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Granite Speech 3.3 — Figure 1

Figure 1 · Source

Granite Speech 4.0

Provides compact speech transcription and translation with a 1B language backbone.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Granite Speech 4.0 — Input/output diagram

Input/output diagram · Source

Granite Speech 4.1

Transcribes and translates multilingual speech using an updated speech-language architecture.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Granite Speech 4.1 — Input/output diagram

Input/output diagram · Source

Hibiki

Translates French speech into English as it arrives while preserving the speaker's voice.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights CC-BY-4.0

Hibiki — Figure 2

Figure 2 · Source

HumanOmni

Interprets human behavior, expressions and interactions from video and audio.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Not stated — Public code and weights were located, but a complete release license was not stated.

HumanOmni — Figure 1

Figure 1 · Source

HumanOmniV2

Interprets people's actions, emotions and social interactions from video and audio.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

HumanOmniV2 — Figure 4

Figure 4 · Source

Ichigo

Answers spoken questions using audio tokens embedded directly into a Llama conversation.

Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 + Llama 3.1 terms — The Llama 3.1 backbone retains its community license; no standalone repository license was located.

Ichigo — Input/output diagram

Input/output diagram · Source

Kimi-Audio

Understands speech, sound and music and supports natural voice conversations.

Paper · Code · Weights · Details

Open license: code MIT / Apache-2.0 · weights MIT

Kimi-Audio — Figure 2

Figure 2 · Source

LFG-1

Answers English spoken questions locally on Apple Silicon and also accepts images.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

LFG-1 — Input/output diagram

Input/output diagram · Source

LFG-2

Reasons about English spoken instructions and returns a text answer with an optional thought channel.

Code · Weights · Details

License unclear: code Not stated · weights Not stated — Code and weights are downloadable, but this release does not state a license.

LFG-2 — Input/output diagram

Input/output diagram · Source

LFG-3

Answers English spoken questions using acoustic features instead of an intermediate transcript.

Code · Weights · Details

License unclear: code Not stated · weights Not stated — Code and weights are downloadable, but this release does not state a license.

LFG-3 — Input/output diagram

Input/output diagram · Source

LFM2-Audio-1.5B

Transcribes speech, synthesizes speech and supports interleaved voice conversations on local devices.

Paper · Code · Weights · Details

Custom / restricted: code LFM Open License v1.0 · weights LFM Open License v1.0 — Revenue-based commercial-use limit; see the LFM license.

LFM2-Audio-1.5B — Figure 7

Figure 7 · Source

LFM2.5-Audio-1.5B

Handles transcription, speech synthesis and voice chat with a faster audio decoder for local inference.

Paper · Code · Weights · Details

Custom / restricted: code LFM Open License v1.0 · weights LFM Open License v1.0 — Revenue-based commercial-use limit; see the LFM license.

LFM2.5-Audio-1.5B — Input/output diagram

Input/output diagram · Source

LFM2.5-Audio-1.5B-JP

Brings Japanese transcription, speech synthesis and spoken conversation to the LFM2.5 audio family.

Paper · Code · Weights · Details

Custom / restricted: code LFM Open License v1.0 · weights LFM Open License v1.0 — Revenue-based commercial-use limit; see the LFM license.

LFM2.5-Audio-1.5B-JP — Input/output diagram

Input/output diagram · Source

LLaMA-Omni

Answers spoken requests with synchronized text and streaming speech.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Research-only — The model card excludes commercial use of the weights.

LLaMA-Omni — Figure 2

Figure 2 · Source

LLaMA-Omni 2

Supports multi-turn spoken chat with streaming speech generation.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Research-only — The model card excludes commercial use of the weights.

LLaMA-Omni 2 — Figure 1

Figure 1 · Source

LLaSM

Answers spoken questions and supports instruction following across speech and text.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights OpenRAIL — Custom release or component terms apply; consult the linked license evidence.

LLaSM — Figure 1

Figure 1 · Source

LongCat-Flash-Omni

Understands text, images, video and audio and produces text or spoken responses.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

LongCat-Flash-Omni — Figure 2

Figure 2 · Source

LongCat-Next

Understands and generates text, images and audio within one multimodal model.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

LongCat-Next — Figure 2

Figure 2 · Source

Lychee-FD

Listens and speaks continuously while deciding when to respond, stop or handle interruptions.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Lychee-FD — Figure 3

Figure 3 · Source

Lyra-Base

Handles long speech, audiovisual questions and spoken dialogue with a 9B configuration.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights CC-BY-NC-4.0 — The repository limits model weights to noncommercial research despite the model-card badge.

Lyra-Base — Figure 2

Figure 2 · Source

Lyra-Mini

Offers audiovisual understanding and spoken interaction in Lyra's smaller 3B configuration.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights CC-BY-NC-4.0 — The repository limits model weights to noncommercial research despite the model-card badge.

Lyra-Mini — Family figure — Figure 2

Family figure — Figure 2 · Source

Mair-hub-0.5B-Omni

Provides a small speech-to-speech assistant trained with a reproducible Qwen-Omni-like recipe.

Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — The training recipe is Apache-2.0; the published checkpoint does not state a weight license.

Mair-hub-0.5B-Omni — Input/output diagram

Input/output diagram · Source

Megrez-Omni

Answers questions about pictures and audio using a compact 3B language backbone.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Megrez-Omni — Figure 2

Figure 2 · Source

Mellow

Uses reasoning over audio to answer questions that require more than recognizing a sound.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Mellow — Figure 2

Figure 2 · Source

MERaLiON 2

Understands multilingual Southeast Asian speech, code-switching, emotion and audio scenes.

Code · Weights · Details

Custom / restricted: code MERaLiON public license · weights MERaLiON public license — Custom MERaLiON terms and incorporated base-model licenses apply.

MERaLiON 2 — Input/output diagram

Input/output diagram · Source

MERaLiON 3

Understands multilingual Southeast Asian speech, including local dialects and conversational code-switching.

Code · Weights · Details

Custom / restricted: code MERaLiON public license · weights MERaLiON public license — Custom MERaLiON terms and incorporated base-model licenses apply.

MERaLiON 3 — Input/output diagram

Input/output diagram · Source

MERaLiON 3 ASR

Transcribes Southeast Asian languages, regional dialects and code-switched speech.

Code · Weights · Details

Custom / restricted: code MERaLiON public license · weights MERaLiON public license — Custom MERaLiON terms and incorporated base-model licenses apply.

MERaLiON 3 ASR — Input/output diagram

Input/output diagram · Source

MERaLiON-AudioLLM

Transcribes, translates and explains audio, including Singaporean speech and local language usage.

Paper · Code · Weights · Details

Custom / restricted: code MERaLiON public license · weights MERaLiON public license — Custom MERaLiON terms and incorporated base-model licenses apply.

MERaLiON-AudioLLM — Figure 1

Figure 1 · Source

MiDashengLM

Describes speech, environmental sounds and music and answers questions about recordings.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

MiDashengLM — Figure 2

Figure 2 · Source

MiMo-Audio

Understands recordings and conducts spoken conversations using instructions and audio examples.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights MIT

MiMo-Audio — Figure 3

Figure 3 · Source

MiMo-V2.5

Answers questions and follows instructions involving text, images, videos and audio.

Code · Weights · Details

Open license: code Apache-2.0 · weights MIT

MiMo-V2.5 — Official architecture diagram

Official architecture diagram · Source

Ming-Flash-Omni

Combines audiovisual understanding with text, image and audio generation.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Ming-Flash-Omni — Figure 2

Figure 2 · Source

Ming-Flash-Omni 2.0

Interprets mixed audiovisual inputs and generates text, images or expressive audio.

Code · Weights · Details

Open license: code MIT · weights MIT

Ming-Flash-Omni 2.0 — Official architecture diagram

Official architecture diagram · Source

Ming-Lite-Omni

Handles audiovisual questions and creates text, images and spoken responses.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Ming-Lite-Omni — Family figure · Figure 2

Family figure · Figure 2 · Source

Mini-Omni

Listens to spoken questions and streams text and speech answers together.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Mini-Omni — Figure 1

Figure 1 · Source

Mini-Omni2

Sees images, follows spoken instructions and streams spoken replies with interruption handling.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Mini-Omni2 — Figure 1

Figure 1 · Source

MiniCPM-o 2.6

Chats about images, video and audio, with speech output and voice imitation.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

MiniCPM-o 2.6 — Official architecture diagram

Official architecture diagram · Source

MiniCPM-o 4.5

Supports live audiovisual conversations while listening and speaking at the same time.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

MiniCPM-o 4.5 — Figure 4

Figure 4 · Source

MiniMind-O

Provides a small, trainable model for image questions and spoken conversations.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

MiniMind-O — Figure 1

Figure 1 · Source

MIO

Understands audiovisual content and generates text, images or speech using multimodal tokens.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 + Yi terms — The Yi-6B-Chat base retains its license; repository-wide code terms are not stated.

MIO — Figure 1

Figure 1 · Source

Moshi

Listens and speaks simultaneously, handling conversational overlap and interruptions.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights CC-BY-4.0

Moshi — Figure 1

Figure 1 · Source

MOSS-Audio Instruct

Captions recordings, answers time-specific questions and transcribes speech with timestamps.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

MOSS-Audio Instruct — Figure 2

Figure 2 · Source

MOSS-Audio Thinking

Reasons about speech, environmental sound and music before producing a text answer.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

MOSS-Audio Thinking — Family figure — Figure 2

Family figure — Figure 2 · Source

MOSS-Speech

Converses directly in Chinese and English speech without first generating a text response.

Paper · Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — The repository licenses code, but a release-specific weight license was not located.

MOSS-Speech — Figure 3

Figure 3 · Source

Music Flamingo

Explains musical structure, instruments and expressive qualities in recordings.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Music Flamingo — Figure 2

Figure 2 · Source

Nemotron 3 Nano Omni

Reasons over documents, images, video and audio and returns text responses.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA Open Model Agreement — Custom model terms apply to the weights.

Nemotron 3 Nano Omni — Figure 1

Figure 1 · Source

Nemotron-Labs-Audex 2B

Performs audio understanding, speech tasks and sound generation with a smaller dense backbone.

Paper · Code · Weights · Details

License unclear: code Not stated · weights NVIDIA OneWay Noncommercial — Noncommercial weight license; inspect the included code files for component-specific terms.

Nemotron-Labs-Audex 2B — Family figure — 30B variant

Family figure — 30B variant · Source

Nemotron-Labs-Audex 30B-A3B

Understands recordings, transcribes and translates speech, and generates speech or general sounds.

Paper · Code · Weights · Details

License unclear: code Not stated · weights NVIDIA OneWay Noncommercial — Noncommercial weight license; inspect the included code files for component-specific terms.

Nemotron-Labs-Audex 30B-A3B — Official family architecture diagram

Official family architecture diagram · Source

NemotronLabs VoiceChat 11B

Holds interruptible, full-duplex voice conversations and can call tools while maintaining the dialogue.

Code · Weights · Details

Open license: code Apache-2.0 · weights OpenMDW-1.1

NemotronLabs VoiceChat 11B — Input/output diagram

Input/output diagram · Source

NExT-GPT

Accepts and generates combinations of text, images, video and audio through connected modality models.

Paper · Code · Weights · Details

Custom / restricted: code BSD-3-Clause · weights CC-BY-NC-SA-4.0 — Noncommercial share-alike weight license; downstream modality decoders retain their own terms.

NExT-GPT — Figure 1, PDF p. 2

Figure 1, PDF p. 2 · Source

Nexus-O

Combines image, video and audio understanding with text and speech interaction.

Paper · Code · Weights · Details

Custom / restricted: code Research-only / CC-BY-NC-4.0 · weights Research-only / CC-BY-NC-4.0 — The model-card body restricts research use despite its Apache metadata.

Nexus-O — Figure 1

Figure 1 · Source

Ola

Answers questions that combine images, video, audio and text.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Ola — Figure 3

Figure 3 · Source

Omni-AutoThink

Adapts how much reasoning it uses before answering audiovisual questions.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Omni-AutoThink — Figure 1

Figure 1 · Source

Omni-R1 (audio reasoning)

Answers questions about speech, sounds and music using reinforcement-trained audio reasoning.

Paper · Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — Public checkpoint archive; no separate weight-license declaration was located.

Omni-R1 (audio reasoning) — Input/output diagram

Input/output diagram · Source

Omni-R1 (two-system collaboration)

Combines global audiovisual reasoning with detailed visual grounding to answer multimodal questions.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Academic use — The model card asks commercial users to contact the authors.

Omni-R1 (two-system collaboration) — Figure 1, PDF p. 2

Figure 1, PDF p. 2 · Source

OmniVinci

Answers questions about combined visual and audio content with temporal alignment.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

OmniVinci — Figure 2

Figure 2 · Source

OneLLM

Answers questions across images, video and audio through one aligned language model.

Paper · Code · Weights · Details

Custom / restricted: code Llama 2 community license · weights Llama 2 community license — Custom base-model community terms apply.

OneLLM — Figure 2

Figure 2 · Source

OpenOmni

Combines audiovisual understanding with multilingual and emotionally aware spoken responses.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

OpenOmni — Figure 1

Figure 1 · Source

OpenS2S

Recognizes emotional cues in speech and generates empathetic, streaming spoken replies.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

OpenS2S — Figure 1

Figure 1 · Source

OpenS2S 1.5

Provides the updated OpenS2S checkpoint for expressive, empathetic spoken dialogue.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

OpenS2S 1.5 — Family figure — Figure 1

Family figure — Figure 1 · Source

OSUM

Recognizes words, timestamps, vocal events, emotions and speaking style from speech.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

OSUM — Figure 2

Figure 2 · Source

OSUM-EChat

Uses speech understanding to produce empathetic text and voice responses.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

OSUM-EChat — Figure 2

Figure 2 · Source

ParaBridge

Adapts replies to vocal emotion and other nonverbal cues in a spoken request.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

ParaBridge — Figure 3

Figure 3 · Source

Parakeet-TDT-v2 + Qwen3-8B

Transcribes speech with Parakeet and sends the transcript to Qwen3 for a text response.

Code · Weights · ASR weights · Details

Open license: code Apache-2.0 · weights Apache-2.0 · ASR CC-BY-4.0

Parakeet-TDT-v2 + Qwen3-8B — Input/output diagram

Input/output diagram · Source

PersonaPlex

Holds full-duplex conversations with a role set by text and a voice set by an audio prompt.

Paper · Code · Weights · Details

Custom / restricted: code MIT · weights NVIDIA Open Model License — Custom NVIDIA weight terms; access requires accepting the model agreement.

PersonaPlex — Figure 1

Figure 1 · Source

Phi-4-multimodal

Reads images, transcribes or translates speech, and answers multimodal questions in text.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights MIT

Phi-4-multimodal — Figure 1

Figure 1 · Source

Qwen-Audio

Transcribes, translates and describes speech, sounds and music through natural-language responses.

Paper · Code · Weights · Details

Custom / restricted: code Tongyi Qianwen license · weights Tongyi Qianwen license — Custom Qwen terms include a large-user-base commercial permission requirement.

Qwen-Audio — Figure 3

Figure 3 · Source

Qwen2-Audio

Answers spoken instructions and analyzes speech, environmental sounds and music in text.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen2-Audio — Figure 2

Figure 2 · Source

Qwen2.5-Omni

Understands images, video and sound and streams coordinated text and speech replies.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen2.5-Omni — Figure 2

Figure 2 · Source

Qwen3-ASR

Identifies languages and transcribes speech across languages and dialects.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen3-ASR — Figure 2

Figure 2 · Source

Qwen3-Omni Instruct

Converses about text, images, video and audio and generates streaming speech.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen3-Omni Instruct — Figure 2

Figure 2 · Source

Qwen3-Omni Thinking

Uses explicit reasoning to answer questions about text, images, video and audio.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen3-Omni Thinking — Family figure — Figure 2

Family figure — Figure 2 · Source

R1-Omni

Reasons about emotion using facial, vocal and contextual cues in videos.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

R1-Omni — Input/output diagram

Input/output diagram · Source

SALMONN

Transcribes speech, captions sound and music, and answers audio-grounded questions.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Apache-2.0 + Llama 2 terms — The Vicuna/Llama 2 language backbone retains its community license.

SALMONN — Figure 1

Figure 1 · Source

SALMONN-2

Understands speech, sounds and music and supports context-aware transcription and audio questions.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

SALMONN-2 — Figure 2

Figure 2 · Source

SLAM-Omni

Conducts multi-turn voice conversations while allowing the output speaker's timbre to be changed.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

SLAM-Omni — Figure 2

Figure 2 · Source

SoulX-Duplug

Predicts when a voice assistant should keep listening, respond or handle an interruption.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

SoulX-Duplug — Figure 4

Figure 4 · Source

SpeechGPT

Follows spoken instructions and generates speech using discrete speech units in a language model.

Paper · Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — The repository licenses code, but a release-specific weight license was not located.

SpeechGPT — Figure 2

Figure 2 · Source

SpeechGPT 2.0 Preview

Understands spoken requests and produces natural spoken responses through a speech-native language model.

Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — The repository licenses code, but a release-specific weight license was not located.

SpeechGPT 2.0 Preview — Official architecture diagram

Official architecture diagram · Source

Step-Audio

Supports voice conversations with expressive speech, role-play and spoken tool requests.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio — Figure 2

Figure 2 · Source

Step-Audio 2 Mini

Understands audio and produces expressive spoken replies, with support for tool-assisted conversations.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio 2 Mini — Figure 3

Figure 3 · Source

Step-Audio 2 Mini Think

Reasons before answering spoken questions and complex audio-understanding tasks.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio 2 Mini Think — Family figure — Figure 3

Family figure — Figure 3 · Source

Step-Audio-AQAA

Produces expressive spoken answers directly from audio questions.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio-AQAA — Figure 1

Figure 1 · Source

Step-Audio-R1

Uses reinforcement-learned reasoning to answer questions about speech, sound and music.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio-R1 — Figure 2

Figure 2 · Source

Step-Audio-R1.1

Extends audio reasoning into responsive spoken dialogue and voice interaction.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio-R1.1 — Related method — Mind-Paced Speaking, Figure 3

Related method — Mind-Paced Speaking, Figure 3 · Source

Stream-Omni

Supports simultaneous text, image and spoken interaction while aligning text and speech outputs.

Paper · Code · Weights · Details

Custom / restricted: code GPL-3.0 · weights GPL-3.0 + Llama 3.1 terms — The included Llama 3.1 backbone retains its community license.

Stream-Omni — Figure 2

Figure 2 · Source

Sympatheia

Adjusts spoken replies to a speaker's changing emotional state.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 + GLM-4 terms — GLM-4-Voice base-model conditions also apply; repository code licensing is not stated.

Sympatheia — Figure 3

Figure 3 · Source

TurnGuide

Maintains coherent spoken conversations while managing overlaps and turn changes.

Paper · Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — No release-specific weight license is stated; the GLM-4-Voice base has custom terms.

TurnGuide — Figure 2

Figure 2 · Source

Ultravox 0.4.1 Llama-3.1-8B

Provides an earlier Ultravox speech-understanding checkpoint for text-based voice responses.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.4.1 Llama-3.1-8B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.5 Llama-3.1-8B

Handles spoken questions and instruction following with text output.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.5 Llama-3.1-8B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.5 Llama-3.2-1B

Provides speech-to-text-answer interaction with a small 1B language backbone.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.5 Llama-3.2-1B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.6 Gemma-3-27B

Answers voice requests using Gemma 3's language capabilities.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Gemma terms — The included Gemma backbone retains its custom usage terms.

Ultravox 0.6 Gemma-3-27B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.6 Llama-3.1-8B

Responds to spoken instructions in text using the 8B Ultravox 0.6 release.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.6 Llama-3.1-8B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.6 Llama-3.3-70B

Turns spoken questions into text answers using a large Llama 3.3 backbone.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.6 Llama-3.3-70B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.6 Qwen3-32B

Adds spoken-instruction understanding and text responses to Qwen3-32B.

Code · Weights · Details

Open license: code MIT · weights MIT

Ultravox 0.6 Qwen3-32B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.7 GLM-4.6

Answers spoken questions and follows voice instructions using a GLM-4.6 language backbone.

Code · Weights · Details

Open license: code MIT · weights MIT

Ultravox 0.7 GLM-4.6 — Family figure — Llama-based example

Family figure — Llama-based example · Source

Uni-MoE 2.0 Omni

Understands audiovisual requests and supports text, image and speech generation.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Weights are Apache-2.0; no license for this implementation directory was located.

Uni-MoE 2.0 Omni — Figure 2

Figure 2 · Source

VibeVoice-ASR

Transcribes long recordings with speaker labels and timestamps in one pass.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

VibeVoice-ASR — Figure 2

Figure 2 · Source

video-SALMONN 2+ (7B)

Captions videos and answers questions using both what is seen and heard.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

video-SALMONN 2+ (7B) — Family figure · Figure 1

Family figure · Figure 1 · Source

VideoLLaMA 2.1 AV

Answers questions about videos using synchronized visual and audio information.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Research-only — The repository restricts use to noncommercial research despite the checkpoint badge.

VideoLLaMA 2.1 AV — Figure 1

Figure 1 · Source

VITA 1.0

Understands images, video and speech in a multimodal assistant with interruptible interaction.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

VITA 1.0 — Figure 2

Figure 2 · Source

VITA 1.5

Provides visual and spoken interaction with streaming speech responses.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

VITA 1.5 — Figure 2

Figure 2 · Source

VITA-Audio

Streams spoken responses quickly by generating multiple audio tokens per language-model step.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

VITA-Audio — Figure 2

Figure 2 · Source

VITA-Audio Plus

Provides the later VITA-Audio release with accelerated interleaved speech and text generation.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

VITA-Audio Plus — Family figure — Figure 2

Family figure — Figure 2 · Source

Voila Autonomous

Supports autonomous voice interaction with simultaneous listening and speaking.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Voila Autonomous — Family figure — Figure 2

Family figure — Figure 2 · Source

Voila Chat

Holds voice conversations with speaker identity and speaking style controlled by prompts.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Voila Chat — Figure 2

Figure 2 · Source

VoxMind

Handles spoken requests that require reasoning, planning and external tool use.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Not stated — Public code and weights were located, but a complete release license was not stated.

VoxMind — Figure 2

Figure 2 · Source

Voxtral Mini

Provides transcription and audio question answering in a smaller 3B configuration.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Voxtral Mini — Family figure — Figure 1

Family figure — Figure 1 · Source

Voxtral Realtime

Transcribes live audio incrementally for low-latency applications.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Voxtral Realtime — Input/output diagram

Input/output diagram · Source

Voxtral Small

Transcribes and translates speech and answers questions about long audio recordings.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Voxtral Small — Figure 1

Figure 1 · Source

Whisper-v3-large + Llama-3.1-8B

Uses Whisper to transcribe a spoken request and Llama 3.1 to answer it.

Code · Weights · ASR weights · Details

Custom / restricted: code Apache-2.0 · weights Llama community license · ASR Apache-2.0 — The text backbone uses the Llama community license.

Whisper-v3-large + Llama-3.1-8B — Input/output diagram

Input/output diagram · Source

Whisper-v3-turbo + Llama-3.1-8B

Pairs faster Whisper transcription with Llama 3.1 text responses.

Code · Weights · ASR weights · Details

Custom / restricted: code Apache-2.0 · weights Llama community license · ASR MIT — The text backbone uses the Llama community license.

Whisper-v3-turbo + Llama-3.1-8B — Input/output diagram

Input/output diagram · Source

Whisper-v3-turbo + Llama-3.2-3B

Combines Whisper Turbo with a smaller Llama 3.2 model for spoken question answering.

Code · Weights · ASR weights · Details

Custom / restricted: code Apache-2.0 · weights Llama community license · ASR MIT — The text backbone uses the Llama community license.

Whisper-v3-turbo + Llama-3.2-3B — Input/output diagram

Input/output diagram · Source


JSON catalog · Apache 2.0 · Third-party figure notice

Contributors

kadirnar

9 commits

kadirnar/awesome-omni-architectures

Visual catalog of omni, speech and audio language models with public code and weights, short descriptions, architecture figures, license labels and VoiceBench coverage.

Python

14

9 commits

updated Sep 13, 2026

See the code

README

Awesome Omni Architectures Awesome

A visual catalog of omni, speech and audio language models with public code and weights. Every entry has a short description, a diagram, code and checkpoint links.

140 models + 4 cascaded systems · Reviewed 2026-09-13

Model list · All diagrams · VoiceBench coverage · Research · Timeline · Scope

Licenses: 68 open · 48 custom/restricted · 28 unclear. Downloadable weights do not always mean an unrestricted open-source license; each card shows the applicable terms.

Models

T: text · I: image · V: video · A: audio · S: speech · M: music · X: other. Scope and labels.

Alphabetical model list · 144 entries
ModelGroupInput → outputRelease
AnyGPTOmniT, I, S, M → T, I, S, MLicense unclear
Audio FlamingoAudioT, A → TCustom / restricted
Audio Flamingo 2AudioT, A → TCustom / restricted
Audio Flamingo 3 ChatSpeechT, A → T, SCustom / restricted
Audio Flamingo NextAudioT, A → TCustom / restricted
Audio-InteractionAudioT, A → TLicense unclear
Audio-ReasonerAudioT, A → TOpen license
Audio-Visual FlamingoOmniT, I, V, A → TCustom / restricted
AzeroSAudioT, S → TOpen license
Baichuan-AudioSpeechT, A → T, SOpen license
Baichuan-Omni 1.5OmniT, I, V, A → T, SCustom / restricted
BLSP-EmoAudioT, S → TCustom / restricted
BR-Voice-ReasonerOmniT, I, V, A → TOpen license
Canary-QwenLLM-ASRT, S → TOpen license
Covo-Audio ChatSpeechT, S → T, SCustom / restricted
DeSTA2AudioT, A → TLicense unclear
DeSTA2.5-AudioAudioT, A → TLicense unclear
DialectS2SSpeechT, S → T, SOpen license
DIFFAAudioT, A → TLicense unclear
DIFFA-2AudioT, A → TLicense unclear
DiVAAudioT, S → TCustom / restricted
DuplexCascadeSpeechT, S → T, SOpen license
DuplexMambaAudioT, S → TOpen license
EMOVA (Qwen2.5-7B)OmniT, I, S → T, SOpen license
Freeze-OmniSpeechT, S → T, SCustom / restricted
Fun-Audio-ChatSpeechT, S → T, SOpen license
Gemma 3n E4BOmniT, I, V, A → TCustom / restricted
Gemma 4 E4BOmniT, I, V, A → TOpen license
GLM-4-VoiceSpeechT, S → T, SCustom / restricted
GLM-ASR NanoLLM-ASRS → TOpen license
Granite Speech 3.3LLM-ASRT, S → TOpen license
Granite Speech 4.0LLM-ASRT, S → TOpen license
Granite Speech 4.1LLM-ASRT, S → TOpen license
HibikiSpeechT, S → T, SOpen license
HumanOmniOmniT, V, A → TLicense unclear
HumanOmniV2OmniT, V, A → TOpen license
IchigoAudioT, S → TLicense unclear
Kimi-AudioSpeechT, A → T, SOpen license
LFG-1OmniT, I, S → TOpen license
LFG-2AudioT, S → TLicense unclear
LFG-3AudioT, S → TLicense unclear
LFM2-Audio-1.5BSpeechT, S → T, SCustom / restricted
LFM2.5-Audio-1.5BSpeechT, S → T, SCustom / restricted
LFM2.5-Audio-1.5B-JPSpeechT, S → T, SCustom / restricted
LLaMA-OmniSpeechT, S → T, SCustom / restricted
LLaMA-Omni 2SpeechT, S → T, SCustom / restricted
LLaSMAudioT, S → TCustom / restricted
LongCat-Flash-OmniOmniT, I, V, A → T, SOpen license
LongCat-NextOmniT, I, A → T, I, AOpen license
Lychee-FDSpeechT, S → T, SOpen license
Lyra-BaseOmniT, I, V, A → T, SCustom / restricted
Lyra-MiniOmniT, I, V, A → T, SCustom / restricted
Mair-hub-0.5B-OmniSpeechT, S → T, SLicense unclear
Megrez-OmniOmniT, I, A → TOpen license
MellowAudioT, A → TOpen license
MERaLiON 2AudioT, A → TCustom / restricted
MERaLiON 3AudioT, S → TCustom / restricted
MERaLiON 3 ASRLLM-ASRS → TCustom / restricted
MERaLiON-AudioLLMAudioT, A → TCustom / restricted
MiDashengLMAudioT, A → TLicense unclear
MiMo-AudioSpeechT, A → T, SOpen license
MiMo-V2.5OmniT, I, V, A → TOpen license
Ming-Flash-OmniOmniT, I, V, A → T, I, AOpen license
Ming-Flash-Omni 2.0OmniT, I, V, A → T, I, AOpen license
Ming-Lite-OmniOmniT, I, V, A → T, I, SOpen license
Mini-OmniSpeechT, S → T, SOpen license
Mini-Omni2OmniT, I, S → T, SOpen license
MiniCPM-o 2.6OmniT, I, V, A → T, SOpen license
MiniCPM-o 4.5OmniT, I, V, A → T, SOpen license
MiniMind-OOmniT, I, S → T, SOpen license
MIOOmniT, I, V, A → T, I, SLicense unclear
MoshiSpeechT, S → T, SOpen license
MOSS-Audio InstructAudioT, A → TLicense unclear
MOSS-Audio ThinkingAudioT, A → TLicense unclear
MOSS-SpeechSpeechS → SLicense unclear
Music FlamingoAudioT, M → TCustom / restricted
Nemotron 3 Nano OmniOmniT, I, V, A → TCustom / restricted
Nemotron-Labs-Audex 2BSpeechT, A → T, ALicense unclear
Nemotron-Labs-Audex 30B-A3BSpeechT, A → T, ALicense unclear
NemotronLabs VoiceChat 11BSpeechT, S → T, SOpen license
NExT-GPTOmniT, I, V, A → T, I, V, ACustom / restricted
Nexus-OOmniT, I, V, A → T, SCustom / restricted
OlaOmniT, I, V, A → TOpen license
Omni-AutoThinkOmniT, I, V, A → TOpen license
Omni-R1 (audio reasoning)AudioT, A → TLicense unclear
Omni-R1 (two-system collaboration)OmniT, I, V, A → TLicense unclear
OmniVinciOmniT, I, V, A → TCustom / restricted
OneLLMOmniT, I, V, A → TCustom / restricted
OpenOmniOmniT, I, V, A → T, SLicense unclear
OpenS2SSpeechT, S → T, SOpen license
OpenS2S 1.5SpeechT, S → T, SOpen license
OSUMAudioT, S → TOpen license
OSUM-EChatSpeechT, S → T, SOpen license
ParaBridgeAudioT, A → TOpen license
Parakeet-TDT-v2 + Qwen3-8BCascadeS → TOpen license
PersonaPlexSpeechT, S → T, SCustom / restricted
Phi-4-multimodalOmniT, I, A → TOpen license
Qwen-AudioAudioT, A → TCustom / restricted
Qwen2-AudioAudioT, A → TOpen license
Qwen2.5-OmniOmniT, I, V, A → T, SOpen license
Qwen3-ASRLLM-ASRT, S → TOpen license
Qwen3-Omni InstructOmniT, I, V, A → T, SOpen license
Qwen3-Omni ThinkingOmniT, I, V, A → TOpen license
R1-OmniOmniT, V, A → TLicense unclear
SALMONNAudioT, A → TCustom / restricted
SALMONN-2AudioT, A → TOpen license
SLAM-OmniSpeechT, S → T, SOpen license
SoulX-DuplugAudioT, S → TOpen license
SpeechGPTSpeechT, S → T, SLicense unclear
SpeechGPT 2.0 PreviewSpeechT, S → T, SLicense unclear
Step-AudioSpeechT, A → T, SOpen license
Step-Audio 2 MiniSpeechT, A → T, SOpen license
Step-Audio 2 Mini ThinkAudioT, A → TOpen license
Step-Audio-AQAASpeechT, A → T, SOpen license
Step-Audio-R1AudioT, A → TOpen license
Step-Audio-R1.1SpeechT, A → T, SOpen license
Stream-OmniOmniT, I, S → T, SCustom / restricted
SympatheiaSpeechT, S → T, SLicense unclear
TurnGuideSpeechT, S → T, SLicense unclear
Ultravox 0.4.1 Llama-3.1-8BAudioT, S → TCustom / restricted
Ultravox 0.5 Llama-3.1-8BAudioT, S → TCustom / restricted
Ultravox 0.5 Llama-3.2-1BAudioT, S → TCustom / restricted
Ultravox 0.6 Gemma-3-27BAudioT, S → TCustom / restricted
Ultravox 0.6 Llama-3.1-8BAudioT, S → TCustom / restricted
Ultravox 0.6 Llama-3.3-70BAudioT, S → TCustom / restricted
Ultravox 0.6 Qwen3-32BAudioT, S → TOpen license
Ultravox 0.7 GLM-4.6AudioT, S → TOpen license
Uni-MoE 2.0 OmniOmniT, I, V, A → T, I, ALicense unclear
VibeVoice-ASRLLM-ASRT, S → TOpen license
video-SALMONN 2+ (7B)OmniT, V, A → TOpen license
VideoLLaMA 2.1 AVOmniT, I, V, A → TCustom / restricted
VITA 1.0OmniT, I, V, A → TCustom / restricted
VITA 1.5OmniT, I, V, A → T, SCustom / restricted
VITA-AudioSpeechT, S → T, SCustom / restricted
VITA-Audio PlusSpeechT, S → T, SCustom / restricted
Voila AutonomousSpeechT, S → T, SOpen license
Voila ChatSpeechT, S → T, SOpen license
VoxMindSpeechT, S → T, SLicense unclear
Voxtral MiniAudioT, A → TOpen license
Voxtral RealtimeLLM-ASRS → TOpen license
Voxtral SmallAudioT, A → TOpen license
Whisper-v3-large + Llama-3.1-8BCascadeS → TCustom / restricted
Whisper-v3-turbo + Llama-3.1-8BCascadeS → TCustom / restricted
Whisper-v3-turbo + Llama-3.2-3BCascadeS → TCustom / restricted

Model figures

Figures are credited to their sources. Editorial input/output diagrams are labeled. Credits.

AnyGPT

Understands and generates text, images, speech and music using a shared token vocabulary.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 + Llama 2 terms — Base-model community terms also apply; repository-wide code licensing is not stated.

AnyGPT — Figure 1

Figure 1 · Source

Audio Flamingo

Answers questions about sounds and music and supports audio-grounded dialogue and examples.

Paper · Code · Weights · Details

Custom / restricted: code MIT · weights OPT-IML noncommercial license — Noncommercial terms apply to model weights and relevant bundled components.

Audio Flamingo — Figure 2

Figure 2 · Source

Audio Flamingo 2

Reasons about environmental sounds and music, including long recordings.

Paper · Code · Weights · Details

Custom / restricted: code MIT · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Audio Flamingo 2 — Figure 2

Figure 2 · Source

Audio Flamingo 3 Chat

Discusses speech, sounds and music across multiple recordings and can reply with streaming speech.

Paper · Code · Weights · Details

Custom / restricted: code MIT · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Audio Flamingo 3 Chat — Figure 2

Figure 2 · Source

Audio Flamingo Next

Analyzes long audio recordings and reasons about speech, sounds and music.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Audio Flamingo Next — Figure 3

Figure 3 · Source

Audio-Interaction

Monitors an ongoing audio stream and decides when a text response or proactive intervention is useful.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

Audio-Interaction — Figure 3

Figure 3 · Source

Audio-Reasoner

Reasons about speech, music and environmental sounds to answer complex audio questions.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Audio-Reasoner — Figure 3

Figure 3 · Source

Audio-Visual Flamingo

Explains long videos by reasoning jointly about visual events and their soundtracks.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Audio-Visual Flamingo — Figure 2

Figure 2 · Source

AzeroS

Follows spoken instructions using speech adaptation learned without curated instruction-answer pairs.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

AzeroS — Figure 1

Figure 1 · Source

Baichuan-Audio

Follows spoken instructions and generates text and expressive speech responses.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Baichuan-Audio — Figure 1

Figure 1 · Source

Baichuan-Omni 1.5

Understands images, video and sound and answers through text or speech.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Baichuan model terms — The model card adds commercial-use conditions beyond its Apache metadata.

Baichuan-Omni 1.5 — Figure 2

Figure 2 · Source

BLSP-Emo

Understands vocal emotion and writes empathetic responses to spoken requests.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Apache-2.0 + Qwen terms — The Qwen-7B-Chat backbone retains its custom license.

BLSP-Emo — Figure 2

Figure 2 · Source

BR-Voice-Reasoner

Answers spoken knowledge and reasoning questions using a speech-adapted Qwen3-Omni Thinker.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

BR-Voice-Reasoner — Input/output diagram

Input/output diagram · Source

Canary-Qwen

Transcribes English speech and offers a separate text-language mode for follow-up processing.

Code · Weights · Details

Open license: code Apache-2.0 · weights CC-BY-4.0

Canary-Qwen — Input/output diagram

Input/output diagram · Source

Covo-Audio Chat

Understands spoken requests and generates context-aware, expressive voice responses.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

Covo-Audio Chat — Figure 2, PDF p. 4

Figure 2, PDF p. 4 · Source

DeSTA2

Describes and answers questions about audio using a language model aligned with speech representations.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Llama community / component terms — Base-model and dataset terms apply; no unrestricted standalone weight license is stated.

DeSTA2 — Figure 1

Figure 1 · Source

DeSTA2.5-Audio

Performs general audio understanding and follows spoken instructions without an external transcript pipeline.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Llama community / component terms — Base-model and dataset terms apply; no unrestricted standalone weight license is stated.

DeSTA2.5-Audio — Figure 2

Figure 2 · Source

DialectS2S

Understands and speaks low-resource Chinese dialects in end-to-end voice conversations.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights MIT

DialectS2S — Family figure — Figure 1

Family figure — Figure 1 · Source

DIFFA

Answers audio questions using diffusion-based text generation instead of left-to-right decoding.

Paper · Code · Weights · Details

License unclear: code Not stated · weights CC-BY-NC-SA-4.0 — The checkpoint uses a noncommercial share-alike license; repository code terms are not stated.

DIFFA — Figure 2

Figure 2 · Source

DIFFA-2

Understands speech, sound and music with a diffusion language model and richer acoustic representations.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Not stated — No release-specific code or weight license was located; do not assume DIFFA 1 terms apply.

DIFFA-2 — Figure 1

Figure 1 · Source

DiVA

Answers spoken questions by transferring a text assistant's behavior into an audio-input model.

Code · Weights · Details

Custom / restricted: code MPL-2.0 · weights MPL-2.0 + Llama 3 terms — The Llama 3 backbone retains its community license.

DiVA — Input/output diagram

Input/output diagram · Source

DuplexCascade

Coordinates streaming transcription, language generation and speech synthesis for interruptible voice conversations.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

DuplexCascade — Figure 1

Figure 1 · Source

DuplexMamba

Reads incoming speech incrementally and supports responsive text generation during speech interaction.

Paper · Code · Weights · Details

Open license: code GPL-3.0 · weights Apache-2.0

DuplexMamba — Figure 1

Figure 1 · Source

EMOVA (Qwen2.5-7B)

Discusses images and spoken requests, replying with text and expressive speech.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

EMOVA (Qwen2.5-7B) — Family figure · Figure 2

Family figure · Figure 2 · Source

Freeze-Omni

Enables streaming spoken dialogue while retaining a frozen text model's language abilities.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

Freeze-Omni — Figure 1

Figure 1 · Source

Fun-Audio-Chat

Holds low-latency spoken conversations while responding to vocal style and emotion.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Fun-Audio-Chat — Figure 2, PDF p. 4

Figure 2, PDF p. 4 · Source

Gemma 3n E4B

Understands text, images, video and audio in a model designed for mobile devices.

Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Gemma terms — The weights use custom Gemma terms.

Gemma 3n E4B — Input/output diagram

Input/output diagram · Source

Gemma 4 E4B

Answers questions about images, video and audio in a compact model designed for local use.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Gemma 4 E4B — Input/output diagram

Input/output diagram · Source

GLM-4-Voice

Converses in Chinese and English with controllable emotion, speaking rate and vocal style.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights GLM-4 model license — Custom GLM model terms apply to the speech checkpoint.

GLM-4-Voice — Figure 2

Figure 2 · Source

GLM-ASR Nano

Transcribes speech in English, Mandarin and Chinese dialects, including noisy recordings.

Code · Weights · Details

Open license: code MIT · weights MIT

GLM-ASR Nano — Input/output diagram

Input/output diagram · Source

Granite Speech 3.3

Transcribes and translates speech while retaining a text language-model interface.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Granite Speech 3.3 — Figure 1

Figure 1 · Source

Granite Speech 4.0

Provides compact speech transcription and translation with a 1B language backbone.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Granite Speech 4.0 — Input/output diagram

Input/output diagram · Source

Granite Speech 4.1

Transcribes and translates multilingual speech using an updated speech-language architecture.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Granite Speech 4.1 — Input/output diagram

Input/output diagram · Source

Hibiki

Translates French speech into English as it arrives while preserving the speaker's voice.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights CC-BY-4.0

Hibiki — Figure 2

Figure 2 · Source

HumanOmni

Interprets human behavior, expressions and interactions from video and audio.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Not stated — Public code and weights were located, but a complete release license was not stated.

HumanOmni — Figure 1

Figure 1 · Source

HumanOmniV2

Interprets people's actions, emotions and social interactions from video and audio.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

HumanOmniV2 — Figure 4

Figure 4 · Source

Ichigo

Answers spoken questions using audio tokens embedded directly into a Llama conversation.

Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 + Llama 3.1 terms — The Llama 3.1 backbone retains its community license; no standalone repository license was located.

Ichigo — Input/output diagram

Input/output diagram · Source

Kimi-Audio

Understands speech, sound and music and supports natural voice conversations.

Paper · Code · Weights · Details

Open license: code MIT / Apache-2.0 · weights MIT

Kimi-Audio — Figure 2

Figure 2 · Source

LFG-1

Answers English spoken questions locally on Apple Silicon and also accepts images.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

LFG-1 — Input/output diagram

Input/output diagram · Source

LFG-2

Reasons about English spoken instructions and returns a text answer with an optional thought channel.

Code · Weights · Details

License unclear: code Not stated · weights Not stated — Code and weights are downloadable, but this release does not state a license.

LFG-2 — Input/output diagram

Input/output diagram · Source

LFG-3

Answers English spoken questions using acoustic features instead of an intermediate transcript.

Code · Weights · Details

License unclear: code Not stated · weights Not stated — Code and weights are downloadable, but this release does not state a license.

LFG-3 — Input/output diagram

Input/output diagram · Source

LFM2-Audio-1.5B

Transcribes speech, synthesizes speech and supports interleaved voice conversations on local devices.

Paper · Code · Weights · Details

Custom / restricted: code LFM Open License v1.0 · weights LFM Open License v1.0 — Revenue-based commercial-use limit; see the LFM license.

LFM2-Audio-1.5B — Figure 7

Figure 7 · Source

LFM2.5-Audio-1.5B

Handles transcription, speech synthesis and voice chat with a faster audio decoder for local inference.

Paper · Code · Weights · Details

Custom / restricted: code LFM Open License v1.0 · weights LFM Open License v1.0 — Revenue-based commercial-use limit; see the LFM license.

LFM2.5-Audio-1.5B — Input/output diagram

Input/output diagram · Source

LFM2.5-Audio-1.5B-JP

Brings Japanese transcription, speech synthesis and spoken conversation to the LFM2.5 audio family.

Paper · Code · Weights · Details

Custom / restricted: code LFM Open License v1.0 · weights LFM Open License v1.0 — Revenue-based commercial-use limit; see the LFM license.

LFM2.5-Audio-1.5B-JP — Input/output diagram

Input/output diagram · Source

LLaMA-Omni

Answers spoken requests with synchronized text and streaming speech.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Research-only — The model card excludes commercial use of the weights.

LLaMA-Omni — Figure 2

Figure 2 · Source

LLaMA-Omni 2

Supports multi-turn spoken chat with streaming speech generation.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Research-only — The model card excludes commercial use of the weights.

LLaMA-Omni 2 — Figure 1

Figure 1 · Source

LLaSM

Answers spoken questions and supports instruction following across speech and text.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights OpenRAIL — Custom release or component terms apply; consult the linked license evidence.

LLaSM — Figure 1

Figure 1 · Source

LongCat-Flash-Omni

Understands text, images, video and audio and produces text or spoken responses.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

LongCat-Flash-Omni — Figure 2

Figure 2 · Source

LongCat-Next

Understands and generates text, images and audio within one multimodal model.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

LongCat-Next — Figure 2

Figure 2 · Source

Lychee-FD

Listens and speaks continuously while deciding when to respond, stop or handle interruptions.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Lychee-FD — Figure 3

Figure 3 · Source

Lyra-Base

Handles long speech, audiovisual questions and spoken dialogue with a 9B configuration.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights CC-BY-NC-4.0 — The repository limits model weights to noncommercial research despite the model-card badge.

Lyra-Base — Figure 2

Figure 2 · Source

Lyra-Mini

Offers audiovisual understanding and spoken interaction in Lyra's smaller 3B configuration.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights CC-BY-NC-4.0 — The repository limits model weights to noncommercial research despite the model-card badge.

Lyra-Mini — Family figure — Figure 2

Family figure — Figure 2 · Source

Mair-hub-0.5B-Omni

Provides a small speech-to-speech assistant trained with a reproducible Qwen-Omni-like recipe.

Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — The training recipe is Apache-2.0; the published checkpoint does not state a weight license.

Mair-hub-0.5B-Omni — Input/output diagram

Input/output diagram · Source

Megrez-Omni

Answers questions about pictures and audio using a compact 3B language backbone.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Megrez-Omni — Figure 2

Figure 2 · Source

Mellow

Uses reasoning over audio to answer questions that require more than recognizing a sound.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Mellow — Figure 2

Figure 2 · Source

MERaLiON 2

Understands multilingual Southeast Asian speech, code-switching, emotion and audio scenes.

Code · Weights · Details

Custom / restricted: code MERaLiON public license · weights MERaLiON public license — Custom MERaLiON terms and incorporated base-model licenses apply.

MERaLiON 2 — Input/output diagram

Input/output diagram · Source

MERaLiON 3

Understands multilingual Southeast Asian speech, including local dialects and conversational code-switching.

Code · Weights · Details

Custom / restricted: code MERaLiON public license · weights MERaLiON public license — Custom MERaLiON terms and incorporated base-model licenses apply.

MERaLiON 3 — Input/output diagram

Input/output diagram · Source

MERaLiON 3 ASR

Transcribes Southeast Asian languages, regional dialects and code-switched speech.

Code · Weights · Details

Custom / restricted: code MERaLiON public license · weights MERaLiON public license — Custom MERaLiON terms and incorporated base-model licenses apply.

MERaLiON 3 ASR — Input/output diagram

Input/output diagram · Source

MERaLiON-AudioLLM

Transcribes, translates and explains audio, including Singaporean speech and local language usage.

Paper · Code · Weights · Details

Custom / restricted: code MERaLiON public license · weights MERaLiON public license — Custom MERaLiON terms and incorporated base-model licenses apply.

MERaLiON-AudioLLM — Figure 1

Figure 1 · Source

MiDashengLM

Describes speech, environmental sounds and music and answers questions about recordings.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

MiDashengLM — Figure 2

Figure 2 · Source

MiMo-Audio

Understands recordings and conducts spoken conversations using instructions and audio examples.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights MIT

MiMo-Audio — Figure 3

Figure 3 · Source

MiMo-V2.5

Answers questions and follows instructions involving text, images, videos and audio.

Code · Weights · Details

Open license: code Apache-2.0 · weights MIT

MiMo-V2.5 — Official architecture diagram

Official architecture diagram · Source

Ming-Flash-Omni

Combines audiovisual understanding with text, image and audio generation.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Ming-Flash-Omni — Figure 2

Figure 2 · Source

Ming-Flash-Omni 2.0

Interprets mixed audiovisual inputs and generates text, images or expressive audio.

Code · Weights · Details

Open license: code MIT · weights MIT

Ming-Flash-Omni 2.0 — Official architecture diagram

Official architecture diagram · Source

Ming-Lite-Omni

Handles audiovisual questions and creates text, images and spoken responses.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Ming-Lite-Omni — Family figure · Figure 2

Family figure · Figure 2 · Source

Mini-Omni

Listens to spoken questions and streams text and speech answers together.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Mini-Omni — Figure 1

Figure 1 · Source

Mini-Omni2

Sees images, follows spoken instructions and streams spoken replies with interruption handling.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Mini-Omni2 — Figure 1

Figure 1 · Source

MiniCPM-o 2.6

Chats about images, video and audio, with speech output and voice imitation.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

MiniCPM-o 2.6 — Official architecture diagram

Official architecture diagram · Source

MiniCPM-o 4.5

Supports live audiovisual conversations while listening and speaking at the same time.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

MiniCPM-o 4.5 — Figure 4

Figure 4 · Source

MiniMind-O

Provides a small, trainable model for image questions and spoken conversations.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

MiniMind-O — Figure 1

Figure 1 · Source

MIO

Understands audiovisual content and generates text, images or speech using multimodal tokens.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 + Yi terms — The Yi-6B-Chat base retains its license; repository-wide code terms are not stated.

MIO — Figure 1

Figure 1 · Source

Moshi

Listens and speaks simultaneously, handling conversational overlap and interruptions.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights CC-BY-4.0

Moshi — Figure 1

Figure 1 · Source

MOSS-Audio Instruct

Captions recordings, answers time-specific questions and transcribes speech with timestamps.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

MOSS-Audio Instruct — Figure 2

Figure 2 · Source

MOSS-Audio Thinking

Reasons about speech, environmental sound and music before producing a text answer.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

MOSS-Audio Thinking — Family figure — Figure 2

Family figure — Figure 2 · Source

MOSS-Speech

Converses directly in Chinese and English speech without first generating a text response.

Paper · Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — The repository licenses code, but a release-specific weight license was not located.

MOSS-Speech — Figure 3

Figure 3 · Source

Music Flamingo

Explains musical structure, instruments and expressive qualities in recordings.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

Music Flamingo — Figure 2

Figure 2 · Source

Nemotron 3 Nano Omni

Reasons over documents, images, video and audio and returns text responses.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA Open Model Agreement — Custom model terms apply to the weights.

Nemotron 3 Nano Omni — Figure 1

Figure 1 · Source

Nemotron-Labs-Audex 2B

Performs audio understanding, speech tasks and sound generation with a smaller dense backbone.

Paper · Code · Weights · Details

License unclear: code Not stated · weights NVIDIA OneWay Noncommercial — Noncommercial weight license; inspect the included code files for component-specific terms.

Nemotron-Labs-Audex 2B — Family figure — 30B variant

Family figure — 30B variant · Source

Nemotron-Labs-Audex 30B-A3B

Understands recordings, transcribes and translates speech, and generates speech or general sounds.

Paper · Code · Weights · Details

License unclear: code Not stated · weights NVIDIA OneWay Noncommercial — Noncommercial weight license; inspect the included code files for component-specific terms.

Nemotron-Labs-Audex 30B-A3B — Official family architecture diagram

Official family architecture diagram · Source

NemotronLabs VoiceChat 11B

Holds interruptible, full-duplex voice conversations and can call tools while maintaining the dialogue.

Code · Weights · Details

Open license: code Apache-2.0 · weights OpenMDW-1.1

NemotronLabs VoiceChat 11B — Input/output diagram

Input/output diagram · Source

NExT-GPT

Accepts and generates combinations of text, images, video and audio through connected modality models.

Paper · Code · Weights · Details

Custom / restricted: code BSD-3-Clause · weights CC-BY-NC-SA-4.0 — Noncommercial share-alike weight license; downstream modality decoders retain their own terms.

NExT-GPT — Figure 1, PDF p. 2

Figure 1, PDF p. 2 · Source

Nexus-O

Combines image, video and audio understanding with text and speech interaction.

Paper · Code · Weights · Details

Custom / restricted: code Research-only / CC-BY-NC-4.0 · weights Research-only / CC-BY-NC-4.0 — The model-card body restricts research use despite its Apache metadata.

Nexus-O — Figure 1

Figure 1 · Source

Ola

Answers questions that combine images, video, audio and text.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Ola — Figure 3

Figure 3 · Source

Omni-AutoThink

Adapts how much reasoning it uses before answering audiovisual questions.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Omni-AutoThink — Figure 1

Figure 1 · Source

Omni-R1 (audio reasoning)

Answers questions about speech, sounds and music using reinforcement-trained audio reasoning.

Paper · Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — Public checkpoint archive; no separate weight-license declaration was located.

Omni-R1 (audio reasoning) — Input/output diagram

Input/output diagram · Source

Omni-R1 (two-system collaboration)

Combines global audiovisual reasoning with detailed visual grounding to answer multimodal questions.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Academic use — The model card asks commercial users to contact the authors.

Omni-R1 (two-system collaboration) — Figure 1, PDF p. 2

Figure 1, PDF p. 2 · Source

OmniVinci

Answers questions about combined visual and audio content with temporal alignment.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights NVIDIA OneWay Noncommercial — Noncommercial terms apply to model weights and relevant bundled components.

OmniVinci — Figure 2

Figure 2 · Source

OneLLM

Answers questions across images, video and audio through one aligned language model.

Paper · Code · Weights · Details

Custom / restricted: code Llama 2 community license · weights Llama 2 community license — Custom base-model community terms apply.

OneLLM — Figure 2

Figure 2 · Source

OpenOmni

Combines audiovisual understanding with multilingual and emotionally aware spoken responses.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

OpenOmni — Figure 1

Figure 1 · Source

OpenS2S

Recognizes emotional cues in speech and generates empathetic, streaming spoken replies.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

OpenS2S — Figure 1

Figure 1 · Source

OpenS2S 1.5

Provides the updated OpenS2S checkpoint for expressive, empathetic spoken dialogue.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

OpenS2S 1.5 — Family figure — Figure 1

Family figure — Figure 1 · Source

OSUM

Recognizes words, timestamps, vocal events, emotions and speaking style from speech.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

OSUM — Figure 2

Figure 2 · Source

OSUM-EChat

Uses speech understanding to produce empathetic text and voice responses.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

OSUM-EChat — Figure 2

Figure 2 · Source

ParaBridge

Adapts replies to vocal emotion and other nonverbal cues in a spoken request.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

ParaBridge — Figure 3

Figure 3 · Source

Parakeet-TDT-v2 + Qwen3-8B

Transcribes speech with Parakeet and sends the transcript to Qwen3 for a text response.

Code · Weights · ASR weights · Details

Open license: code Apache-2.0 · weights Apache-2.0 · ASR CC-BY-4.0

Parakeet-TDT-v2 + Qwen3-8B — Input/output diagram

Input/output diagram · Source

PersonaPlex

Holds full-duplex conversations with a role set by text and a voice set by an audio prompt.

Paper · Code · Weights · Details

Custom / restricted: code MIT · weights NVIDIA Open Model License — Custom NVIDIA weight terms; access requires accepting the model agreement.

PersonaPlex — Figure 1

Figure 1 · Source

Phi-4-multimodal

Reads images, transcribes or translates speech, and answers multimodal questions in text.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights MIT

Phi-4-multimodal — Figure 1

Figure 1 · Source

Qwen-Audio

Transcribes, translates and describes speech, sounds and music through natural-language responses.

Paper · Code · Weights · Details

Custom / restricted: code Tongyi Qianwen license · weights Tongyi Qianwen license — Custom Qwen terms include a large-user-base commercial permission requirement.

Qwen-Audio — Figure 3

Figure 3 · Source

Qwen2-Audio

Answers spoken instructions and analyzes speech, environmental sounds and music in text.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen2-Audio — Figure 2

Figure 2 · Source

Qwen2.5-Omni

Understands images, video and sound and streams coordinated text and speech replies.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen2.5-Omni — Figure 2

Figure 2 · Source

Qwen3-ASR

Identifies languages and transcribes speech across languages and dialects.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen3-ASR — Figure 2

Figure 2 · Source

Qwen3-Omni Instruct

Converses about text, images, video and audio and generates streaming speech.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen3-Omni Instruct — Figure 2

Figure 2 · Source

Qwen3-Omni Thinking

Uses explicit reasoning to answer questions about text, images, video and audio.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Qwen3-Omni Thinking — Family figure — Figure 2

Family figure — Figure 2 · Source

R1-Omni

Reasons about emotion using facial, vocal and contextual cues in videos.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Public code and weights were located, but a complete release license was not stated.

R1-Omni — Input/output diagram

Input/output diagram · Source

SALMONN

Transcribes speech, captions sound and music, and answers audio-grounded questions.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Apache-2.0 + Llama 2 terms — The Vicuna/Llama 2 language backbone retains its community license.

SALMONN — Figure 1

Figure 1 · Source

SALMONN-2

Understands speech, sounds and music and supports context-aware transcription and audio questions.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

SALMONN-2 — Figure 2

Figure 2 · Source

SLAM-Omni

Conducts multi-turn voice conversations while allowing the output speaker's timbre to be changed.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

SLAM-Omni — Figure 2

Figure 2 · Source

SoulX-Duplug

Predicts when a voice assistant should keep listening, respond or handle an interruption.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

SoulX-Duplug — Figure 4

Figure 4 · Source

SpeechGPT

Follows spoken instructions and generates speech using discrete speech units in a language model.

Paper · Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — The repository licenses code, but a release-specific weight license was not located.

SpeechGPT — Figure 2

Figure 2 · Source

SpeechGPT 2.0 Preview

Understands spoken requests and produces natural spoken responses through a speech-native language model.

Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — The repository licenses code, but a release-specific weight license was not located.

SpeechGPT 2.0 Preview — Official architecture diagram

Official architecture diagram · Source

Step-Audio

Supports voice conversations with expressive speech, role-play and spoken tool requests.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio — Figure 2

Figure 2 · Source

Step-Audio 2 Mini

Understands audio and produces expressive spoken replies, with support for tool-assisted conversations.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio 2 Mini — Figure 3

Figure 3 · Source

Step-Audio 2 Mini Think

Reasons before answering spoken questions and complex audio-understanding tasks.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio 2 Mini Think — Family figure — Figure 3

Family figure — Figure 3 · Source

Step-Audio-AQAA

Produces expressive spoken answers directly from audio questions.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio-AQAA — Figure 1

Figure 1 · Source

Step-Audio-R1

Uses reinforcement-learned reasoning to answer questions about speech, sound and music.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio-R1 — Figure 2

Figure 2 · Source

Step-Audio-R1.1

Extends audio reasoning into responsive spoken dialogue and voice interaction.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Step-Audio-R1.1 — Related method — Mind-Paced Speaking, Figure 3

Related method — Mind-Paced Speaking, Figure 3 · Source

Stream-Omni

Supports simultaneous text, image and spoken interaction while aligning text and speech outputs.

Paper · Code · Weights · Details

Custom / restricted: code GPL-3.0 · weights GPL-3.0 + Llama 3.1 terms — The included Llama 3.1 backbone retains its community license.

Stream-Omni — Figure 2

Figure 2 · Source

Sympatheia

Adjusts spoken replies to a speaker's changing emotional state.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 + GLM-4 terms — GLM-4-Voice base-model conditions also apply; repository code licensing is not stated.

Sympatheia — Figure 3

Figure 3 · Source

TurnGuide

Maintains coherent spoken conversations while managing overlaps and turn changes.

Paper · Code · Weights · Details

License unclear: code Apache-2.0 · weights Not stated — No release-specific weight license is stated; the GLM-4-Voice base has custom terms.

TurnGuide — Figure 2

Figure 2 · Source

Ultravox 0.4.1 Llama-3.1-8B

Provides an earlier Ultravox speech-understanding checkpoint for text-based voice responses.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.4.1 Llama-3.1-8B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.5 Llama-3.1-8B

Handles spoken questions and instruction following with text output.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.5 Llama-3.1-8B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.5 Llama-3.2-1B

Provides speech-to-text-answer interaction with a small 1B language backbone.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.5 Llama-3.2-1B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.6 Gemma-3-27B

Answers voice requests using Gemma 3's language capabilities.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Gemma terms — The included Gemma backbone retains its custom usage terms.

Ultravox 0.6 Gemma-3-27B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.6 Llama-3.1-8B

Responds to spoken instructions in text using the 8B Ultravox 0.6 release.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.6 Llama-3.1-8B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.6 Llama-3.3-70B

Turns spoken questions into text answers using a large Llama 3.3 backbone.

Code · Weights · Details

Custom / restricted: code MIT · weights MIT + Llama community terms — Ultravox code is MIT; the included Llama backbone retains its community license.

Ultravox 0.6 Llama-3.3-70B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.6 Qwen3-32B

Adds spoken-instruction understanding and text responses to Qwen3-32B.

Code · Weights · Details

Open license: code MIT · weights MIT

Ultravox 0.6 Qwen3-32B — Family figure — Llama-based example

Family figure — Llama-based example · Source

Ultravox 0.7 GLM-4.6

Answers spoken questions and follows voice instructions using a GLM-4.6 language backbone.

Code · Weights · Details

Open license: code MIT · weights MIT

Ultravox 0.7 GLM-4.6 — Family figure — Llama-based example

Family figure — Llama-based example · Source

Uni-MoE 2.0 Omni

Understands audiovisual requests and supports text, image and speech generation.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Apache-2.0 — Weights are Apache-2.0; no license for this implementation directory was located.

Uni-MoE 2.0 Omni — Figure 2

Figure 2 · Source

VibeVoice-ASR

Transcribes long recordings with speaker labels and timestamps in one pass.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

VibeVoice-ASR — Figure 2

Figure 2 · Source

video-SALMONN 2+ (7B)

Captions videos and answers questions using both what is seen and heard.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

video-SALMONN 2+ (7B) — Family figure · Figure 1

Family figure · Figure 1 · Source

VideoLLaMA 2.1 AV

Answers questions about videos using synchronized visual and audio information.

Paper · Code · Weights · Details

Custom / restricted: code Apache-2.0 · weights Research-only — The repository restricts use to noncommercial research despite the checkpoint badge.

VideoLLaMA 2.1 AV — Figure 1

Figure 1 · Source

VITA 1.0

Understands images, video and speech in a multimodal assistant with interruptible interaction.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

VITA 1.0 — Figure 2

Figure 2 · Source

VITA 1.5

Provides visual and spoken interaction with streaming speech responses.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

VITA 1.5 — Figure 2

Figure 2 · Source

VITA-Audio

Streams spoken responses quickly by generating multiple audio tokens per language-model step.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

VITA-Audio — Figure 2

Figure 2 · Source

VITA-Audio Plus

Provides the later VITA-Audio release with accelerated interleaved speech and text generation.

Paper · Code · Weights · Details

Custom / restricted: code Tencent research license · weights Tencent research license — Academic/research use only; commercial and production use are excluded.

VITA-Audio Plus — Family figure — Figure 2

Family figure — Figure 2 · Source

Voila Autonomous

Supports autonomous voice interaction with simultaneous listening and speaking.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Voila Autonomous — Family figure — Figure 2

Family figure — Figure 2 · Source

Voila Chat

Holds voice conversations with speaker identity and speaking style controlled by prompts.

Paper · Code · Weights · Details

Open license: code MIT · weights MIT

Voila Chat — Figure 2

Figure 2 · Source

VoxMind

Handles spoken requests that require reasoning, planning and external tool use.

Paper · Code · Weights · Details

License unclear: code Not stated · weights Not stated — Public code and weights were located, but a complete release license was not stated.

VoxMind — Figure 2

Figure 2 · Source

Voxtral Mini

Provides transcription and audio question answering in a smaller 3B configuration.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Voxtral Mini — Family figure — Figure 1

Family figure — Figure 1 · Source

Voxtral Realtime

Transcribes live audio incrementally for low-latency applications.

Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Voxtral Realtime — Input/output diagram

Input/output diagram · Source

Voxtral Small

Transcribes and translates speech and answers questions about long audio recordings.

Paper · Code · Weights · Details

Open license: code Apache-2.0 · weights Apache-2.0

Voxtral Small — Figure 1

Figure 1 · Source

Whisper-v3-large + Llama-3.1-8B

Uses Whisper to transcribe a spoken request and Llama 3.1 to answer it.

Code · Weights · ASR weights · Details

Custom / restricted: code Apache-2.0 · weights Llama community license · ASR Apache-2.0 — The text backbone uses the Llama community license.

Whisper-v3-large + Llama-3.1-8B — Input/output diagram

Input/output diagram · Source

Whisper-v3-turbo + Llama-3.1-8B

Pairs faster Whisper transcription with Llama 3.1 text responses.

Code · Weights · ASR weights · Details

Custom / restricted: code Apache-2.0 · weights Llama community license · ASR MIT — The text backbone uses the Llama community license.

Whisper-v3-turbo + Llama-3.1-8B — Input/output diagram

Input/output diagram · Source

Whisper-v3-turbo + Llama-3.2-3B

Combines Whisper Turbo with a smaller Llama 3.2 model for spoken question answering.

Code · Weights · ASR weights · Details

Custom / restricted: code Apache-2.0 · weights Llama community license · ASR MIT — The text backbone uses the Llama community license.

Whisper-v3-turbo + Llama-3.2-3B — Input/output diagram

Input/output diagram · Source


JSON catalog · Apache 2.0 · Third-party figure notice

Contributors

kadirnar

9 commits

Languages

Python

100.0%