18
stars
2
commits
1
repos using this model
1
linked in READMEs
Aug 24, 2026
updated
TLive-Omni is an omni-modal understanding model for e-commerce live-stream, mapping image, video, audio, and text into a unified text-output interface. Built on a Qwen3.5 backbone with a grafted AuT audio encoder, it supports up to 256K tokens of context, trained via a three-stage SFT recipe followed by Faithful-RFT reinforcement fine-tuning.
TLive-Omni is built on a Qwen3.5 backbone and extends it with a audio encoder through a lightweight MLP aligner, forming a unified text-output omni-modal understanding model. For video inputs with audio, each temporal grid is organized into a timestamped grid that interleaves video and audio token blocks, keeping audio segments adjacent to their corresponding visual content. The model supports up to 256K tokens of context at inference.
We evaluate TLive-Omni-4B and TLive-Omni-9B on both live-commerce tasks and general benchmarks. Dash (-) denotes an unreported result or undisclosed parameter count. The Best results among the compared open-source models are marked in bold, while the second-best results are in underlined.
| Task | Metric | TLive-Omni 4B | TLive-Omni 9B | Gemini 2.5 Flash | Gemini 2.5 Pro | Gemini 3 Flash | Gemini 3 Pro | Gemini 3.5 Flash | Qwen3.5-Omni Flash | OmniVinci 9B | Nemotron 3 Nano Omni 30B-A3B | Ming-Lite-Omni v1.5 20B-A3B | MiniCPM-o 2.6 8B | MiniCPM-o 4.5 9B | Qwen2.5-Omni 7B | Qwen3-Omni 30B-A3B |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Audio | ||||||||||||||||
| Live-Commerce ASR | CER β | 6.66 | 6.46 | 16.30 | 11.48 | 15.18 | 12.09 | 13.09 | 6.81 | β | 12.10 | 10.06 | 13.88 | 10.70 | 7.86 | 6.75 |
| Speaker-Attributed ASR | cpWER β | 12.88 | 12.27 | 17.14 | 12.17 | 19.04 | 11.67 | 11.99 | 13.23 | β | 17.65 | β | β | 18.89 | β | 27.84 |
| Audio Description | Acc. β | 76.12 | 75.96 | 65.21 | 81.10 | 68.27 | 85.07 | 79.97 | 62.82 | 39.90 | 33.01 | 45.99 | 49.84 | 47.59 | 47.92 | 61.06 |
| Audio Description | Hal. β | 20.97 | 21.00 | 26.19 | 14.16 | 26.17 | 10.92 | 14.36 | 27.81 | 47.36 | 39.77 | 44.05 | 41.76 | 52.41 | 36.92 | 30.22 |
| Audio QA | Acc. β | 72.60 | 76.28 | 76.28 | 82.85 | 74.68 | 88.62 | 87.99 | 78.04 | 66.51 | 64.90 | 40.54 | 39.74 | 42.47 | 61.38 | 76.76 |
| Image | ||||||||||||||||
| Visual Grounding | Live AP β | 82.85 | 82.33 | 61.08 | 51.98 | 80.38 | 73.80 | 84.15 | 79.96 | 34.86 | 73.08 | 52.46 | 3.82 | 23.90 | 75.61 | 79.22 |
| Visual Grounding | Prod AP β | 91.45 | 89.96 | 28.81 | 32.63 | 65.67 | 58.83 | 74.89 | 60.44 | 8.93 | 48.62 | 40.73 | 1.77 | 53.63 | 22.85 | 68.88 |
| Text Understanding | Loc. F1 β | 86.99 | 87.59 | 20.52 | 31.60 | 61.11 | 68.60 | 64.44 | 74.07 | 50.25 | 52.91 | 13.27 | 5.74 | 5.43 | 42.64 | 30.46 |
| Text Understanding | Rec. NED β | 4.72 | 4.24 | 43.28 | 27.82 | 16.25 | 9.72 | 16.64 | 12.48 | 32.77 | 29.42 | 59.16 | 77.58 | 71.65 | 37.79 | 14.83 |
| Text Understanding | Cls. Acc. β | 79.06 | 79.85 | 51.21 | 61.86 | 69.11 | 76.86 | 69.76 | 53.25 | 57.29 | 37.86 | 32.94 | 15.92 | 11.62 | 51.25 | 69.46 |
| Video | ||||||||||||||||
| Temporal Grounding | mIoU β | 77.63 | 81.49 | 76.50 | 76.22 | 77.43 | 77.90 | 77.90 | 62.10 | 13.10 | 23.39 | 14.34 | 14.56 | 43.20 | 30.83 | 39.22 |
| Dense Caption | Acc. β | 69.23 | 74.63 | 54.60 | 41.95 | 32.21 | 37.80 | 33.80 | 32.94 | 18.59 | 17.96 | 13.81 | 10.53 | 21.06 | 16.51 | 21.44 |
| Dense Caption | Hal. β | 9.57 | 8.76 | 10.97 | 16.88 | 20.76 | 20.99 | 17.30 | 20.91 | 27.13 | 16.62 | 39.33 | 26.93 | 28.61 | 36.44 | 25.82 |
| Video QA | Acc. β | 92.31 | 93.23 | 88.21 | 92.62 | 89.64 | 84.36 | 86.90 | 87.28 | 72.51 | 82.56 | 64.51 | 60.30 | 84.62 | 75.48 | 81.62 |
| Shot Understanding | Layout β | 78.40 | 77.00 | 80.00 | 85.20 | 76.80 | 80.40 | 83.40 | 84.40 | 73.60 | 79.20 | 74.60 | 66.60 | 78.20 | 74.40 | 82.20 |
| Shot Understanding | Shot Size β | 51.20 | 51.00 | 46.80 | 50.80 | 45.70 | 43.40 | 44.20 | 48.90 | 52.70 | 34.00 | 41.70 | 38.30 | 42.80 | 38.10 | 37.40 |
| Shot Understanding | Camera β | 80.90 | 82.00 | 84.20 | 76.00 | 78.50 | 75.70 | 75.50 | 85.50 | 68.10 | 80.20 | 72.80 | 68.30 | 79.20 | 81.20 | 76.10 |
| Shot Understanding | Content β | 69.80 | 71.00 | 68.60 | 70.40 | 71.60 | 74.80 | 70.20 | 66.40 | 49.20 | 58.60 | 51.60 | 49.00 | 66.60 | 68.00 | 63.60 |
| Model | Params | MMMU | MathVista | DynaMath | VLMsAreBlind | MMBench(EN-DEV-v1.1) | RealWorldQA | MMStar | SimpleVQA | Hallusion | AI2D | OCRBench | CC-OCR | CharXiv(RQ) | RefCOCO | ERQA | EmbSpatial |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Open-source VLM models | |||||||||||||||||
| MiMo-VL-SFT | 7B | 64.6 | 81.8 | 46.9 | 78.0 | 84.5 | β | β | β | β | 83.2 | 87.6 | β | 54.4 | 85.7 | β | β |
| SAIL-VL2 | 8B | 55.4 | 76.4 | 17.8 | β | β | 76.3 | 70.7 | β | 55.1 | 87.7 | 91.3 | β | β | 74.0 | β | β |
| Valley2.5 | 8B | 62.1 | 74.4 | 32.7 | β | 85.5 | 70.5 | 67.3 | β | 56.3 | 84.4 | 87.0 | β | β | β | β | β |
| LLaVA-OneVision-2 | 8B | β | β | β | β | 85.7 | 69.7 | 64.8 | β | β | 84.3 | 78.2 | β | β | β | 43.3 | 78.1 |
| InternVL3.5 | 4B | 66.6 | 77.1 | 35.7 | β | 80.3 | 66.3 | 65.0 | β | 44.8 | 82.6 | 82.2 | β | 39.6 | 89.4 | 38.5 | β |
| InternVL3.5 | 8B | 73.4 | 78.4 | 37.7 | β | 79.5 | 67.5 | 69.3 | β | 54.5 | 84.0 | 84.0 | β | 44.4 | 89.7 | 41.0 | 73.2 |
| Qwen3-VL | 4B | 67.4 | 73.7 | 65.3 | 71.9 | 83.9 | 70.9 | 69.8 | 48.0 | 57.6 | 84.1 | 88.1 | 76.2 | 39.7 | 89.0 | 41.3 | 79.6 |
| Qwen3-VL | 8B | 69.6 | 77.2 | 67.7 | 74.0 | 84.5 | 71.5 | 70.9 | 50.2 | 61.1 | 85.7 | 89.6 | 79.9 | 46.4 | 89.1 | 45.8 | 78.5 |
| Qwen3.5 | 4B | 72.1 | 81.0 | 69.6 | 62.3 | 86.3 | 72.5 | 74.8 | 44.6 | 76.9 | 87.1 | 85.9 | 71.1 | 62.9 | 87.6 | 46.8 | 76.6 |
| Qwen3.5 | 9B | 74.2 | 82.2 | 74.6 | 71.8 | 87.7 | 72.9 | 76.3 | 48.9 | 76.0 | 88.0 | 88.5 | 73.4 | 67.5 | 90.0 | 47.3 | 78.7 |
| Open-source Omni models | |||||||||||||||||
| InteractiveOmni | 4B | 61.1 | 61.7 | β | β | 78.9 | β | 62.6 | β | 52.2 | 83.8 | 80.0 | β | β | β | β | β |
| InteractiveOmni | 8B | 66.9 | 68.0 | β | β | 81.4 | β | 66.8 | β | 61.3 | 84.3 | 83.7 | β | β | β | β | β |
| VITA-1.5 | 7B | 52.1 | 66.2 | β | β | 76.7 | β | 59.9 | β | 44.9 | 79.3 | 73.2 | β | β | β | β | β |
| Valley3 | 8B | 69.3 | β | β | β | β | β | β | β | 55.9 | β | β | β | β | β | β | β |
| OmniVinci | 9B | 49.7 | 63.5 | β | β | β | 67.5 | β | β | β | β | β | β | β | β | β | β |
| Nemotron 3 Nano Omni | 30B-A3B | 55.2 | 71.9 | β | β | β | β | β | β | β | 88.5 | 88.3 | β | 49.1 | 80.6 | β | β |
| Ming-Lite-Omni v1.5 | 20B-A3B | 54.3 | 72.0 | β | β | β | β | 65.1 | β | 54.6 | 84.9 | 88.9 | β | β | 87.8 | β | β |
| MiniCPM-o 2.6 | 8B | 50.4 | 71.9 | β | β | 80.5 | β | 64.0 | β | 51.9 | 85.8 | 89.7 | β | β | β | β | β |
| MiniCPM-o 4.5 | 9B | 67.6 | β | β | β | 87.6 | β | 73.1 | β | 63.2 | 87.6 | 87.6 | β | β | β | β | β |
| Qwen2.5-Omni | 7B | 59.2 | 67.9 | β | β | 81.8 | 70.3 | 64.0 | β | β | 83.2 | β | β | β | 87.7 | β | β |
| Qwen3-Omni | 30B-A3B | 69.1 | 75.9 | β | β | β | β | 68.5 | β | 59.7 | 85.2 | 86.0 | β | 61.1 | β | β | β |
| Ours | |||||||||||||||||
| TLive-Omni | 4B | 70.9 | 79.9 | 72.5 | 71.8 | 87.0 | 77.7 | 73.9 | 47.6 | 77.7 | 86.6 | 86.6 | 80.5 | 61.3 | 87.4 | 42.3 | 79.3 |
| TLive-Omni | 9B | 73.4 | 81.9 | 73.3 | 75.5 | 88.9 | 76.6 | 75.1 | 50.0 | 76.0 | 88.6 | 90.3 | 81.3 | 63.1 | 90.0 | 48.0 | 80.4 |
| Model | Params | MVBench | MLVU | Video-MME | LongVideoBench | LVBench | MMVU | VideoMMMU | Charades-TL | ActivityNet-TL | QVHighlights-TL |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Open-source VLM models | |||||||||||
| MiMo-VL-SFT | 7B | β | β | 66.9 | β | β | β | 53.1 | 39.6 | 35.5 | 41.5 |
| SAIL-VL2 | 8B | β | β | 62.7 | 58.3 | β | β | β | β | β | β |
| LLaVA-OneVision-2 | 8B | 66.2 | 76.6 | 71.9 | 66.9 | 55.5 | 56.2 | β | 53.5 | 53.8 | 66.4 |
| LLaVA-Video | 7B | 58.6 | 70.8 | 63.3 | 58.2 | 44.2 | 47.1 | 36.1 | 15.2 | 14.6 | 10.4 |
| InternVL3.5 | 4B | 71.2 | 70.4 | 65.4 | 60.8 | 43.2 | 47.6 | 57.6 | 16.0 | 14.9 | 17.7 |
| InternVL3.5 | 8B | 72.1 | 70.2 | 66.0 | 62.1 | 46.7 | 60.2 | β | 27.8 | 31.3 | 31.3 |
| MiniCPM-V 4.5 | 8B | β | 75.1 | 67.9 | 63.9 | 50.4 | 58.9 | 57.1 | 31.9 | 32.3 | 46.1 |
| LongVU | 7B | 66.9 | 65.4 | 60.6 | β | β | β | β | β | β | β |
| LongVILA | 7B | 67.1 | β | 60.1 | 57.1 | β | β | β | β | β | β |
| Mage-VL | 4B | 65.1 | 68.7 | 64.0 | 61.3 | 41.8 | β | β | 50.7 | 45.4 | 57.4 |
| Molmo2 | 4B | 75.1 | 63.0 | 69.6 | 68.0 | 53.9 | 51.2 | 50.7 | 33.3 | 39.8 | 58.7 |
| Molmo2 | 8B | 75.9 | 60.2 | 69.9 | 67.5 | 52.8 | β | β | β | β | β |
| NVILA | 8B | 68.1 | 70.1 | 64.2 | 57.7 | β | β | β | β | β | β |
| Kangaroo | 8B | 61.1 | 61.0 | 56.0 | 54.8 | 39.4 | β | β | β | β | β |
| Video-XL2 | 8B | β | 74.8 | 66.6 | 61.0 | 48.4 | 50.0 | 39.9 | 38.9 | 30.0 | 46.2 |
| VideoChat3 | 4B | β | β | 70.1 | β | 56.7 | 56.4 | 57.4 | 56.1 | 54.6 | 67.0 |
| VideoLLaMA 3 | 7B | 69.7 | 73.0 | 66.2 | 59.8 | 45.3 | 44.1 | 34.6 | 39.8 | 29.8 | 36.9 |
| Qwen3-VL | 4B | 68.9 | 75.3 | 69.3 | β | 56.2 | 50.5 | 56.2 | 46.4 | 48.2 | 58.7 |
| Qwen3-VL | 8B | 68.7 | 78.1 | 71.4 | β | 58.0 | 58.7 | 65.3 | 48.3 | 46.8 | 59.4 |
| Qwen3.5 | 4B | 66.6 | 75.1 | 71.6 | 65.1 | 55.3 | 57.8 | 69.8 | 48.7 | 51.6 | 55.0 |
| Qwen3.5 | 9B | 75.7 | 79.7 | 66.9 | 67.9 | 60.9 | 63.7 | 70.3 | 52.0 | 54.0 | 57.2 |
| Open-source Omni models | |||||||||||
| InteractiveOmni | 4B | β | 68.0 | 63.3 | 57.0 | β | β | β | β | β | β |
| InteractiveOmni | 8B | β | 71.6 | 66.0 | 59.1 | β | β | β | β | β | β |
| VITA-1.5 | 7B | 55.4 | β | 56.1 | β | β | β | β | β | β | β |
| Valley3 | 8B | β | 55.6 | β | β | β | β | 61.2 | β | β | β |
| OmniVinci | 9B | 70.6 | β | 68.2 | 61.3 | β | β | β | β | β | β |
| Nemotron 3 Nano Omni | 30B-A3B | β | β | 70.8 | β | β | β | β | β | β | β |
| Ming-Lite-Omni v1.5 | 20B-A3B | 69.4 | β | 67.1 | 59.5 | β | β | β | β | β | β |
| MiniCPM-o 2.6 | 8B | β | β | 63.9 | β | β | β | β | β | β | β |
| MiniCPM-o 4.5 | 9B | β | 76.5 | 70.4 | 66.0 | β | β | β | β | β | β |
| Qwen2.5-Omni | 7B | 70.3 | β | 64.3 | β | β | β | β | β | β | β |
| Qwen3-Omni | 30B-A3B | β | 75.2 | 70.5 | β | β | β | β | β | β | β |
| Ours | |||||||||||
| TLive-Omni | 4B | 69.0 | 76.1 | 71.3 | 66.1 | 57.1 | 59.9 | 73.9 | 57.0 | 58.2 | 69.2 |
| TLive-Omni | 9B | 72.5 | 80.9 | 75.6 | 69.9 | 60.8 | 67.1 | 72.8 | 56.3 | 55.4 | 64.1 |
| Model | Params | AVUT | WorldSense | VideoHolmes | DailyOmni | OmniVideoBench | FutureOmni |
|---|---|---|---|---|---|---|---|
| Open-source Omni models | |||||||
| video-SALMONN 2+ | 3B | 66.2 | 48.3 | 42.2 | 67.7 | β | β |
| video-SALMONN 2+ | 7B | 69.5 | 50.9 | 46.9 | 71.8 | β | β |
| OmniVinci | 9B | β | 48.2 | β | 66.5 | 36.7 | 52.8 |
| Nemotron 3 Nano Omni | 30B-A3B | β | 55.2 | β | 74.5 | β | β |
| MiniCPM-o 4.5 | 9B | 78.6 | 55.7 | 64.3 | 80.2 | 41.1 | 56.1 |
| Qwen2.5-Omni | 7B | β | 45.4 | β | 62.4 | 36.5 | 48.9 |
| Qwen3-Omni | 30B-A3B | 74.2 | 54.0 | 50.4 | 71.9 | 43.8 | 53.4 |
| Ours | |||||||
| TLive-Omni | 4B | 78.6 | 54.0 | 57.5 | 78.6 | 41.6 | 57.2 |
| TLive-Omni | 9B | 80.0 | 56.0 | 59.3 | 80.5 | 43.2 | 58.5 |
This release targets Python 3.10 on Linux x86_64 with CUDA 12.8 and PyTorch 2.10.0.
conda create -n tlive python=3.10 -y
conda activate tlive
pip install -r https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/v1.0.0-rc1/environments/requirements.txt
The remote environments/requirements.txt includes custom wheels for the supported environment and model. If any wheel does not match your hardware, CUDA version, or Python version, replace it with a compatible build for your setup.
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
model_id = "TaoLiveAIGC/TLive-Omni-4B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto",
attn_implementation="flash_attention_2",
).eval()
def generate(messages, *, use_audio_in_video=False, videos_kwargs=None, generation_kwargs=None):
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
enable_thinking=False,
use_audio_in_video=use_audio_in_video,
videos_kwargs=videos_kwargs or {},
)
prompt_length = inputs["input_ids"].shape[-1]
inputs = inputs.to(model.device)
generation_kwargs = generation_kwargs or {}
with torch.inference_mode():
generated_ids = model.generate(
**inputs,
do_sample=False,
max_new_tokens=1024,
**generation_kwargs,
)
answer_ids = generated_ids[:, prompt_length:]
answer = processor.batch_decode(
answer_ids,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0]
return answer.strip()
Replace messages with one of the examples below for text, image, audio, or video inputs.
Text
messages = [{
"role": "user",
"content": [{"type": "text", "text": "Briefly explain why multimodal context can improve an answer."}],
}]
print(generate(messages))
Image
messages = [{
"role": "user",
"content": [
{"type": "image", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/image.jpg"},
{"type": "text", "text": "Describe this image."},
],
}]
print(generate(messages))
Audio
messages = [{
"role": "user",
"content": [
{"type": "audio", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/audio.mp3"},
{"type": "text", "text": "Transcribe and summarize this audio."},
],
}]
print(generate(messages))
Video with audio
messages = [{
"role": "user",
"content": [
{"type": "video", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/vocal_video.mp4"},
{"type": "text", "text": "Describe the video, including relevant speech and sounds."},
],
}]
print(generate(messages, use_audio_in_video=True, videos_kwargs={"fps": 1.0}))
Video without audio
messages = [{
"role": "user",
"content": [
{"type": "video", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/silence_video.mp4"},
{"type": "text", "text": "Describe the visual events in this video."},
],
}]
print(generate(messages, use_audio_in_video=False, videos_kwargs={"fps": 1.0}))
For temporal localization outputs, we recommend the MM:SS - MM:SS interval format, for example 01:23 - 01:35. For videos, set use_audio_in_video=True when the audio track should be used, and False for visual-only inference.
First install the pre-built wheel (Python 3.10 + CUDA 12.8 + Linux x86_64), built and tested on NVIDIA H20 GPUs (Hopper, sm_90):
pip install https://github.com/TaoLiveAIGC/TLive-Omni/releases/download/v1.0.0-rc1/vllm-0.19.0+cu128-cp310-cp310-linux_x86_64.whl
If your GPU, driver, or CUDA setup is not compatible with this wheel, build vLLM from source using the customized code in the vllm/ directory of the GitHub release.
from transformers import AutoProcessor
from vllm import LLM, SamplingParams
from vllm.model_executor.models.tlive_omni_processing import process_audio_info
model_id = "TaoLiveAIGC/TLive-Omni-4B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
def build_prompt(messages):
return processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=False,
enable_thinking=False,
)
def generate(inputs, *, limit_mm_per_prompt=None, vllm_kwargs=None, sampling_kwargs=None):
vllm_kwargs = vllm_kwargs or {}
sampling_kwargs = sampling_kwargs or {}
llm = LLM(
model=model_id,
trust_remote_code=True,
dtype="bfloat16",
max_model_len=32768,
tensor_parallel_size=1,
gpu_memory_utilization=0.9,
max_num_seqs=4,
max_num_batched_tokens=32768,
seed=42,
limit_mm_per_prompt=limit_mm_per_prompt,
**vllm_kwargs,
)
outputs = llm.generate(
inputs,
sampling_params=SamplingParams(
temperature=0.0,
max_tokens=1024,
**sampling_kwargs,
),
)
return outputs[0].outputs[0].text.strip()
Replace messages with one of the examples below for text, image, audio, or video inputs.
Text
messages = [{
"role": "user",
"content": [{"type": "text", "text": "Briefly explain why multimodal context can improve an answer."}],
}]
inputs = {"prompt": build_prompt(messages)}
print(generate(inputs))
Image
image_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/image.jpg"
messages = [{
"role": "user",
"content": [
{"type": "image", "path": image_path},
{"type": "text", "text": "Describe this image."},
],
}]
inputs = {
"prompt": build_prompt(messages),
"multi_modal_data": {"image": [image_path]},
}
print(generate(inputs, limit_mm_per_prompt={"image": 1}))
Audio
audio_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/audio.mp3"
messages = [{
"role": "user",
"content": [
{"type": "audio", "audio": audio_path},
{"type": "text", "text": "Transcribe and summarize this audio."},
],
}]
inputs = {
"prompt": build_prompt(messages),
"multi_modal_data": {"audio": process_audio_info(messages, use_audio_in_video=False)},
}
print(generate(inputs, limit_mm_per_prompt={"audio": 1}))
Video with audio
video_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/vocal_video.mp4"
messages = [{
"role": "user",
"content": [
{"type": "video", "video": video_path},
{"type": "text", "text": "Describe the video, including relevant speech and sounds."},
],
}]
inputs = {
"prompt": build_prompt(messages),
"multi_modal_data": {
"video": [video_path],
"audio": process_audio_info(messages, use_audio_in_video=True),
},
"mm_processor_kwargs": {"videos_kwargs": {"fps": 1.0, "use_audio_in_video": True, "return_metadata": True}},
}
print(generate(inputs, limit_mm_per_prompt={"video": 1, "audio": 1}))
Video without audio
video_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/silence_video.mp4"
messages = [{
"role": "user",
"content": [
{"type": "video", "video": video_path},
{"type": "text", "text": "Describe the visual events in this video."},
],
}]
inputs = {
"prompt": build_prompt(messages),
"multi_modal_data": {"video": [video_path]},
"mm_processor_kwargs": {"videos_kwargs": {"fps": 1.0, "use_audio_in_video": False, "return_metadata": True}},
}
print(generate(inputs, limit_mm_per_prompt={"video": 1}))
If you find our work helpful, please consider citing our paper:
@article{hu2026tliveomni,
title={TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming},
author={Hu, Yibo and Qian, Yu and Gu, Mao and Tao, Yingfan and Chen, Yuhao and Luo, Yongdong and Liu, Zhuoqun and Jin, Meiguang and Ma, Junfeng},
journal={arXiv preprint arXiv:2608.20958},
year={2026}
}
This project is released under the Apache License 2.0.
18
stars
2
commits
1
repos using this model
1
linked in READMEs
Aug 24, 2026
updated
TLive-Omni is an omni-modal understanding model for e-commerce live-stream, mapping image, video, audio, and text into a unified text-output interface. Built on a Qwen3.5 backbone with a grafted AuT audio encoder, it supports up to 256K tokens of context, trained via a three-stage SFT recipe followed by Faithful-RFT reinforcement fine-tuning.
TLive-Omni is built on a Qwen3.5 backbone and extends it with a audio encoder through a lightweight MLP aligner, forming a unified text-output omni-modal understanding model. For video inputs with audio, each temporal grid is organized into a timestamped grid that interleaves video and audio token blocks, keeping audio segments adjacent to their corresponding visual content. The model supports up to 256K tokens of context at inference.
We evaluate TLive-Omni-4B and TLive-Omni-9B on both live-commerce tasks and general benchmarks. Dash (-) denotes an unreported result or undisclosed parameter count. The Best results among the compared open-source models are marked in bold, while the second-best results are in underlined.
| Task | Metric | TLive-Omni 4B | TLive-Omni 9B | Gemini 2.5 Flash | Gemini 2.5 Pro | Gemini 3 Flash | Gemini 3 Pro | Gemini 3.5 Flash | Qwen3.5-Omni Flash | OmniVinci 9B | Nemotron 3 Nano Omni 30B-A3B | Ming-Lite-Omni v1.5 20B-A3B | MiniCPM-o 2.6 8B | MiniCPM-o 4.5 9B | Qwen2.5-Omni 7B | Qwen3-Omni 30B-A3B |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Audio | ||||||||||||||||
| Live-Commerce ASR | CER β | 6.66 | 6.46 | 16.30 | 11.48 | 15.18 | 12.09 | 13.09 | 6.81 | β | 12.10 | 10.06 | 13.88 | 10.70 | 7.86 | 6.75 |
| Speaker-Attributed ASR | cpWER β | 12.88 | 12.27 | 17.14 | 12.17 | 19.04 | 11.67 | 11.99 | 13.23 | β | 17.65 | β | β | 18.89 | β | 27.84 |
| Audio Description | Acc. β | 76.12 | 75.96 | 65.21 | 81.10 | 68.27 | 85.07 | 79.97 | 62.82 | 39.90 | 33.01 | 45.99 | 49.84 | 47.59 | 47.92 | 61.06 |
| Audio Description | Hal. β | 20.97 | 21.00 | 26.19 | 14.16 | 26.17 | 10.92 | 14.36 | 27.81 | 47.36 | 39.77 | 44.05 | 41.76 | 52.41 | 36.92 | 30.22 |
| Audio QA | Acc. β | 72.60 | 76.28 | 76.28 | 82.85 | 74.68 | 88.62 | 87.99 | 78.04 | 66.51 | 64.90 | 40.54 | 39.74 | 42.47 | 61.38 | 76.76 |
| Image | ||||||||||||||||
| Visual Grounding | Live AP β | 82.85 | 82.33 | 61.08 | 51.98 | 80.38 | 73.80 | 84.15 | 79.96 | 34.86 | 73.08 | 52.46 | 3.82 | 23.90 | 75.61 | 79.22 |
| Visual Grounding | Prod AP β | 91.45 | 89.96 | 28.81 | 32.63 | 65.67 | 58.83 | 74.89 | 60.44 | 8.93 | 48.62 | 40.73 | 1.77 | 53.63 | 22.85 | 68.88 |
| Text Understanding | Loc. F1 β | 86.99 | 87.59 | 20.52 | 31.60 | 61.11 | 68.60 | 64.44 | 74.07 | 50.25 | 52.91 | 13.27 | 5.74 | 5.43 | 42.64 | 30.46 |
| Text Understanding | Rec. NED β | 4.72 | 4.24 | 43.28 | 27.82 | 16.25 | 9.72 | 16.64 | 12.48 | 32.77 | 29.42 | 59.16 | 77.58 | 71.65 | 37.79 | 14.83 |
| Text Understanding | Cls. Acc. β | 79.06 | 79.85 | 51.21 | 61.86 | 69.11 | 76.86 | 69.76 | 53.25 | 57.29 | 37.86 | 32.94 | 15.92 | 11.62 | 51.25 | 69.46 |
| Video | ||||||||||||||||
| Temporal Grounding | mIoU β | 77.63 | 81.49 | 76.50 | 76.22 | 77.43 | 77.90 | 77.90 | 62.10 | 13.10 | 23.39 | 14.34 | 14.56 | 43.20 | 30.83 | 39.22 |
| Dense Caption | Acc. β | 69.23 | 74.63 | 54.60 | 41.95 | 32.21 | 37.80 | 33.80 | 32.94 | 18.59 | 17.96 | 13.81 | 10.53 | 21.06 | 16.51 | 21.44 |
| Dense Caption | Hal. β | 9.57 | 8.76 | 10.97 | 16.88 | 20.76 | 20.99 | 17.30 | 20.91 | 27.13 | 16.62 | 39.33 | 26.93 | 28.61 | 36.44 | 25.82 |
| Video QA | Acc. β | 92.31 | 93.23 | 88.21 | 92.62 | 89.64 | 84.36 | 86.90 | 87.28 | 72.51 | 82.56 | 64.51 | 60.30 | 84.62 | 75.48 | 81.62 |
| Shot Understanding | Layout β | 78.40 | 77.00 | 80.00 | 85.20 | 76.80 | 80.40 | 83.40 | 84.40 | 73.60 | 79.20 | 74.60 | 66.60 | 78.20 | 74.40 | 82.20 |
| Shot Understanding | Shot Size β | 51.20 | 51.00 | 46.80 | 50.80 | 45.70 | 43.40 | 44.20 | 48.90 | 52.70 | 34.00 | 41.70 | 38.30 | 42.80 | 38.10 | 37.40 |
| Shot Understanding | Camera β | 80.90 | 82.00 | 84.20 | 76.00 | 78.50 | 75.70 | 75.50 | 85.50 | 68.10 | 80.20 | 72.80 | 68.30 | 79.20 | 81.20 | 76.10 |
| Shot Understanding | Content β | 69.80 | 71.00 | 68.60 | 70.40 | 71.60 | 74.80 | 70.20 | 66.40 | 49.20 | 58.60 | 51.60 | 49.00 | 66.60 | 68.00 | 63.60 |
| Model | Params | MMMU | MathVista | DynaMath | VLMsAreBlind | MMBench(EN-DEV-v1.1) | RealWorldQA | MMStar | SimpleVQA | Hallusion | AI2D | OCRBench | CC-OCR | CharXiv(RQ) | RefCOCO | ERQA | EmbSpatial |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Open-source VLM models | |||||||||||||||||
| MiMo-VL-SFT | 7B | 64.6 | 81.8 | 46.9 | 78.0 | 84.5 | β | β | β | β | 83.2 | 87.6 | β | 54.4 | 85.7 | β | β |
| SAIL-VL2 | 8B | 55.4 | 76.4 | 17.8 | β | β | 76.3 | 70.7 | β | 55.1 | 87.7 | 91.3 | β | β | 74.0 | β | β |
| Valley2.5 | 8B | 62.1 | 74.4 | 32.7 | β | 85.5 | 70.5 | 67.3 | β | 56.3 | 84.4 | 87.0 | β | β | β | β | β |
| LLaVA-OneVision-2 | 8B | β | β | β | β | 85.7 | 69.7 | 64.8 | β | β | 84.3 | 78.2 | β | β | β | 43.3 | 78.1 |
| InternVL3.5 | 4B | 66.6 | 77.1 | 35.7 | β | 80.3 | 66.3 | 65.0 | β | 44.8 | 82.6 | 82.2 | β | 39.6 | 89.4 | 38.5 | β |
| InternVL3.5 | 8B | 73.4 | 78.4 | 37.7 | β | 79.5 | 67.5 | 69.3 | β | 54.5 | 84.0 | 84.0 | β | 44.4 | 89.7 | 41.0 | 73.2 |
| Qwen3-VL | 4B | 67.4 | 73.7 | 65.3 | 71.9 | 83.9 | 70.9 | 69.8 | 48.0 | 57.6 | 84.1 | 88.1 | 76.2 | 39.7 | 89.0 | 41.3 | 79.6 |
| Qwen3-VL | 8B | 69.6 | 77.2 | 67.7 | 74.0 | 84.5 | 71.5 | 70.9 | 50.2 | 61.1 | 85.7 | 89.6 | 79.9 | 46.4 | 89.1 | 45.8 | 78.5 |
| Qwen3.5 | 4B | 72.1 | 81.0 | 69.6 | 62.3 | 86.3 | 72.5 | 74.8 | 44.6 | 76.9 | 87.1 | 85.9 | 71.1 | 62.9 | 87.6 | 46.8 | 76.6 |
| Qwen3.5 | 9B | 74.2 | 82.2 | 74.6 | 71.8 | 87.7 | 72.9 | 76.3 | 48.9 | 76.0 | 88.0 | 88.5 | 73.4 | 67.5 | 90.0 | 47.3 | 78.7 |
| Open-source Omni models | |||||||||||||||||
| InteractiveOmni | 4B | 61.1 | 61.7 | β | β | 78.9 | β | 62.6 | β | 52.2 | 83.8 | 80.0 | β | β | β | β | β |
| InteractiveOmni | 8B | 66.9 | 68.0 | β | β | 81.4 | β | 66.8 | β | 61.3 | 84.3 | 83.7 | β | β | β | β | β |
| VITA-1.5 | 7B | 52.1 | 66.2 | β | β | 76.7 | β | 59.9 | β | 44.9 | 79.3 | 73.2 | β | β | β | β | β |
| Valley3 | 8B | 69.3 | β | β | β | β | β | β | β | 55.9 | β | β | β | β | β | β | β |
| OmniVinci | 9B | 49.7 | 63.5 | β | β | β | 67.5 | β | β | β | β | β | β | β | β | β | β |
| Nemotron 3 Nano Omni | 30B-A3B | 55.2 | 71.9 | β | β | β | β | β | β | β | 88.5 | 88.3 | β | 49.1 | 80.6 | β | β |
| Ming-Lite-Omni v1.5 | 20B-A3B | 54.3 | 72.0 | β | β | β | β | 65.1 | β | 54.6 | 84.9 | 88.9 | β | β | 87.8 | β | β |
| MiniCPM-o 2.6 | 8B | 50.4 | 71.9 | β | β | 80.5 | β | 64.0 | β | 51.9 | 85.8 | 89.7 | β | β | β | β | β |
| MiniCPM-o 4.5 | 9B | 67.6 | β | β | β | 87.6 | β | 73.1 | β | 63.2 | 87.6 | 87.6 | β | β | β | β | β |
| Qwen2.5-Omni | 7B | 59.2 | 67.9 | β | β | 81.8 | 70.3 | 64.0 | β | β | 83.2 | β | β | β | 87.7 | β | β |
| Qwen3-Omni | 30B-A3B | 69.1 | 75.9 | β | β | β | β | 68.5 | β | 59.7 | 85.2 | 86.0 | β | 61.1 | β | β | β |
| Ours | |||||||||||||||||
| TLive-Omni | 4B | 70.9 | 79.9 | 72.5 | 71.8 | 87.0 | 77.7 | 73.9 | 47.6 | 77.7 | 86.6 | 86.6 | 80.5 | 61.3 | 87.4 | 42.3 | 79.3 |
| TLive-Omni | 9B | 73.4 | 81.9 | 73.3 | 75.5 | 88.9 | 76.6 | 75.1 | 50.0 | 76.0 | 88.6 | 90.3 | 81.3 | 63.1 | 90.0 | 48.0 | 80.4 |
| Model | Params | MVBench | MLVU | Video-MME | LongVideoBench | LVBench | MMVU | VideoMMMU | Charades-TL | ActivityNet-TL | QVHighlights-TL |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Open-source VLM models | |||||||||||
| MiMo-VL-SFT | 7B | β | β | 66.9 | β | β | β | 53.1 | 39.6 | 35.5 | 41.5 |
| SAIL-VL2 | 8B | β | β | 62.7 | 58.3 | β | β | β | β | β | β |
| LLaVA-OneVision-2 | 8B | 66.2 | 76.6 | 71.9 | 66.9 | 55.5 | 56.2 | β | 53.5 | 53.8 | 66.4 |
| LLaVA-Video | 7B | 58.6 | 70.8 | 63.3 | 58.2 | 44.2 | 47.1 | 36.1 | 15.2 | 14.6 | 10.4 |
| InternVL3.5 | 4B | 71.2 | 70.4 | 65.4 | 60.8 | 43.2 | 47.6 | 57.6 | 16.0 | 14.9 | 17.7 |
| InternVL3.5 | 8B | 72.1 | 70.2 | 66.0 | 62.1 | 46.7 | 60.2 | β | 27.8 | 31.3 | 31.3 |
| MiniCPM-V 4.5 | 8B | β | 75.1 | 67.9 | 63.9 | 50.4 | 58.9 | 57.1 | 31.9 | 32.3 | 46.1 |
| LongVU | 7B | 66.9 | 65.4 | 60.6 | β | β | β | β | β | β | β |
| LongVILA | 7B | 67.1 | β | 60.1 | 57.1 | β | β | β | β | β | β |
| Mage-VL | 4B | 65.1 | 68.7 | 64.0 | 61.3 | 41.8 | β | β | 50.7 | 45.4 | 57.4 |
| Molmo2 | 4B | 75.1 | 63.0 | 69.6 | 68.0 | 53.9 | 51.2 | 50.7 | 33.3 | 39.8 | 58.7 |
| Molmo2 | 8B | 75.9 | 60.2 | 69.9 | 67.5 | 52.8 | β | β | β | β | β |
| NVILA | 8B | 68.1 | 70.1 | 64.2 | 57.7 | β | β | β | β | β | β |
| Kangaroo | 8B | 61.1 | 61.0 | 56.0 | 54.8 | 39.4 | β | β | β | β | β |
| Video-XL2 | 8B | β | 74.8 | 66.6 | 61.0 | 48.4 | 50.0 | 39.9 | 38.9 | 30.0 | 46.2 |
| VideoChat3 | 4B | β | β | 70.1 | β | 56.7 | 56.4 | 57.4 | 56.1 | 54.6 | 67.0 |
| VideoLLaMA 3 | 7B | 69.7 | 73.0 | 66.2 | 59.8 | 45.3 | 44.1 | 34.6 | 39.8 | 29.8 | 36.9 |
| Qwen3-VL | 4B | 68.9 | 75.3 | 69.3 | β | 56.2 | 50.5 | 56.2 | 46.4 | 48.2 | 58.7 |
| Qwen3-VL | 8B | 68.7 | 78.1 | 71.4 | β | 58.0 | 58.7 | 65.3 | 48.3 | 46.8 | 59.4 |
| Qwen3.5 | 4B | 66.6 | 75.1 | 71.6 | 65.1 | 55.3 | 57.8 | 69.8 | 48.7 | 51.6 | 55.0 |
| Qwen3.5 | 9B | 75.7 | 79.7 | 66.9 | 67.9 | 60.9 | 63.7 | 70.3 | 52.0 | 54.0 | 57.2 |
| Open-source Omni models | |||||||||||
| InteractiveOmni | 4B | β | 68.0 | 63.3 | 57.0 | β | β | β | β | β | β |
| InteractiveOmni | 8B | β | 71.6 | 66.0 | 59.1 | β | β | β | β | β | β |
| VITA-1.5 | 7B | 55.4 | β | 56.1 | β | β | β | β | β | β | β |
| Valley3 | 8B | β | 55.6 | β | β | β | β | 61.2 | β | β | β |
| OmniVinci | 9B | 70.6 | β | 68.2 | 61.3 | β | β | β | β | β | β |
| Nemotron 3 Nano Omni | 30B-A3B | β | β | 70.8 | β | β | β | β | β | β | β |
| Ming-Lite-Omni v1.5 | 20B-A3B | 69.4 | β | 67.1 | 59.5 | β | β | β | β | β | β |
| MiniCPM-o 2.6 | 8B | β | β | 63.9 | β | β | β | β | β | β | β |
| MiniCPM-o 4.5 | 9B | β | 76.5 | 70.4 | 66.0 | β | β | β | β | β | β |
| Qwen2.5-Omni | 7B | 70.3 | β | 64.3 | β | β | β | β | β | β | β |
| Qwen3-Omni | 30B-A3B | β | 75.2 | 70.5 | β | β | β | β | β | β | β |
| Ours | |||||||||||
| TLive-Omni | 4B | 69.0 | 76.1 | 71.3 | 66.1 | 57.1 | 59.9 | 73.9 | 57.0 | 58.2 | 69.2 |
| TLive-Omni | 9B | 72.5 | 80.9 | 75.6 | 69.9 | 60.8 | 67.1 | 72.8 | 56.3 | 55.4 | 64.1 |
| Model | Params | AVUT | WorldSense | VideoHolmes | DailyOmni | OmniVideoBench | FutureOmni |
|---|---|---|---|---|---|---|---|
| Open-source Omni models | |||||||
| video-SALMONN 2+ | 3B | 66.2 | 48.3 | 42.2 | 67.7 | β | β |
| video-SALMONN 2+ | 7B | 69.5 | 50.9 | 46.9 | 71.8 | β | β |
| OmniVinci | 9B | β | 48.2 | β | 66.5 | 36.7 | 52.8 |
| Nemotron 3 Nano Omni | 30B-A3B | β | 55.2 | β | 74.5 | β | β |
| MiniCPM-o 4.5 | 9B | 78.6 | 55.7 | 64.3 | 80.2 | 41.1 | 56.1 |
| Qwen2.5-Omni | 7B | β | 45.4 | β | 62.4 | 36.5 | 48.9 |
| Qwen3-Omni | 30B-A3B | 74.2 | 54.0 | 50.4 | 71.9 | 43.8 | 53.4 |
| Ours | |||||||
| TLive-Omni | 4B | 78.6 | 54.0 | 57.5 | 78.6 | 41.6 | 57.2 |
| TLive-Omni | 9B | 80.0 | 56.0 | 59.3 | 80.5 | 43.2 | 58.5 |
This release targets Python 3.10 on Linux x86_64 with CUDA 12.8 and PyTorch 2.10.0.
conda create -n tlive python=3.10 -y
conda activate tlive
pip install -r https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/v1.0.0-rc1/environments/requirements.txt
The remote environments/requirements.txt includes custom wheels for the supported environment and model. If any wheel does not match your hardware, CUDA version, or Python version, replace it with a compatible build for your setup.
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
model_id = "TaoLiveAIGC/TLive-Omni-4B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto",
attn_implementation="flash_attention_2",
).eval()
def generate(messages, *, use_audio_in_video=False, videos_kwargs=None, generation_kwargs=None):
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
enable_thinking=False,
use_audio_in_video=use_audio_in_video,
videos_kwargs=videos_kwargs or {},
)
prompt_length = inputs["input_ids"].shape[-1]
inputs = inputs.to(model.device)
generation_kwargs = generation_kwargs or {}
with torch.inference_mode():
generated_ids = model.generate(
**inputs,
do_sample=False,
max_new_tokens=1024,
**generation_kwargs,
)
answer_ids = generated_ids[:, prompt_length:]
answer = processor.batch_decode(
answer_ids,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0]
return answer.strip()
Replace messages with one of the examples below for text, image, audio, or video inputs.
Text
messages = [{
"role": "user",
"content": [{"type": "text", "text": "Briefly explain why multimodal context can improve an answer."}],
}]
print(generate(messages))
Image
messages = [{
"role": "user",
"content": [
{"type": "image", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/image.jpg"},
{"type": "text", "text": "Describe this image."},
],
}]
print(generate(messages))
Audio
messages = [{
"role": "user",
"content": [
{"type": "audio", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/audio.mp3"},
{"type": "text", "text": "Transcribe and summarize this audio."},
],
}]
print(generate(messages))
Video with audio
messages = [{
"role": "user",
"content": [
{"type": "video", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/vocal_video.mp4"},
{"type": "text", "text": "Describe the video, including relevant speech and sounds."},
],
}]
print(generate(messages, use_audio_in_video=True, videos_kwargs={"fps": 1.0}))
Video without audio
messages = [{
"role": "user",
"content": [
{"type": "video", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/silence_video.mp4"},
{"type": "text", "text": "Describe the visual events in this video."},
],
}]
print(generate(messages, use_audio_in_video=False, videos_kwargs={"fps": 1.0}))
For temporal localization outputs, we recommend the MM:SS - MM:SS interval format, for example 01:23 - 01:35. For videos, set use_audio_in_video=True when the audio track should be used, and False for visual-only inference.
First install the pre-built wheel (Python 3.10 + CUDA 12.8 + Linux x86_64), built and tested on NVIDIA H20 GPUs (Hopper, sm_90):
pip install https://github.com/TaoLiveAIGC/TLive-Omni/releases/download/v1.0.0-rc1/vllm-0.19.0+cu128-cp310-cp310-linux_x86_64.whl
If your GPU, driver, or CUDA setup is not compatible with this wheel, build vLLM from source using the customized code in the vllm/ directory of the GitHub release.
from transformers import AutoProcessor
from vllm import LLM, SamplingParams
from vllm.model_executor.models.tlive_omni_processing import process_audio_info
model_id = "TaoLiveAIGC/TLive-Omni-4B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
def build_prompt(messages):
return processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=False,
enable_thinking=False,
)
def generate(inputs, *, limit_mm_per_prompt=None, vllm_kwargs=None, sampling_kwargs=None):
vllm_kwargs = vllm_kwargs or {}
sampling_kwargs = sampling_kwargs or {}
llm = LLM(
model=model_id,
trust_remote_code=True,
dtype="bfloat16",
max_model_len=32768,
tensor_parallel_size=1,
gpu_memory_utilization=0.9,
max_num_seqs=4,
max_num_batched_tokens=32768,
seed=42,
limit_mm_per_prompt=limit_mm_per_prompt,
**vllm_kwargs,
)
outputs = llm.generate(
inputs,
sampling_params=SamplingParams(
temperature=0.0,
max_tokens=1024,
**sampling_kwargs,
),
)
return outputs[0].outputs[0].text.strip()
Replace messages with one of the examples below for text, image, audio, or video inputs.
Text
messages = [{
"role": "user",
"content": [{"type": "text", "text": "Briefly explain why multimodal context can improve an answer."}],
}]
inputs = {"prompt": build_prompt(messages)}
print(generate(inputs))
Image
image_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/image.jpg"
messages = [{
"role": "user",
"content": [
{"type": "image", "path": image_path},
{"type": "text", "text": "Describe this image."},
],
}]
inputs = {
"prompt": build_prompt(messages),
"multi_modal_data": {"image": [image_path]},
}
print(generate(inputs, limit_mm_per_prompt={"image": 1}))
Audio
audio_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/audio.mp3"
messages = [{
"role": "user",
"content": [
{"type": "audio", "audio": audio_path},
{"type": "text", "text": "Transcribe and summarize this audio."},
],
}]
inputs = {
"prompt": build_prompt(messages),
"multi_modal_data": {"audio": process_audio_info(messages, use_audio_in_video=False)},
}
print(generate(inputs, limit_mm_per_prompt={"audio": 1}))
Video with audio
video_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/vocal_video.mp4"
messages = [{
"role": "user",
"content": [
{"type": "video", "video": video_path},
{"type": "text", "text": "Describe the video, including relevant speech and sounds."},
],
}]
inputs = {
"prompt": build_prompt(messages),
"multi_modal_data": {
"video": [video_path],
"audio": process_audio_info(messages, use_audio_in_video=True),
},
"mm_processor_kwargs": {"videos_kwargs": {"fps": 1.0, "use_audio_in_video": True, "return_metadata": True}},
}
print(generate(inputs, limit_mm_per_prompt={"video": 1, "audio": 1}))
Video without audio
video_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/silence_video.mp4"
messages = [{
"role": "user",
"content": [
{"type": "video", "video": video_path},
{"type": "text", "text": "Describe the visual events in this video."},
],
}]
inputs = {
"prompt": build_prompt(messages),
"multi_modal_data": {"video": [video_path]},
"mm_processor_kwargs": {"videos_kwargs": {"fps": 1.0, "use_audio_in_video": False, "return_metadata": True}},
}
print(generate(inputs, limit_mm_per_prompt={"video": 1}))
If you find our work helpful, please consider citing our paper:
@article{hu2026tliveomni,
title={TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming},
author={Hu, Yibo and Qian, Yu and Gu, Mao and Tao, Yingfan and Chen, Yuhao and Luo, Yongdong and Liu, Zhuoqun and Jin, Meiguang and Ma, Junfeng},
journal={arXiv preprint arXiv:2608.20958},
year={2026}
}
This project is released under the Apache License 2.0.