TaoLiveAIGC/TLive-Omni-9B

Model

15

stars

2

commits

1

linked in READMEs

Aug 24, 2026

updated

audio
conversational
custom-code
custom_code
multimodal
safetensors
text-generation
tlive_omni
transformers
video
vision
Browse cluster: Multimedia processing and video benchmarking β†’

README

TLive-Omni

logo

Technical Report Hugging Face 4B Model Hugging Face 9B Model License

πŸ“‹ Overview

TLive-Omni is an omni-modal understanding model for e-commerce live-stream, mapping image, video, audio, and text into a unified text-output interface. Built on a Qwen3.5 backbone with a grafted AuT audio encoder, it supports up to 256K tokens of context, trained via a three-stage SFT recipe followed by Faithful-RFT reinforcement fine-tuning.

✨ Highlights

  • Timestamped Per-vGrid layout β€” Audio and video tokens are organized into timestamped grid with explicit boundaries, keeping audio segments adjacent to their corresponding visual content for fine-grained temporal alignment over long streams.
  • Three-stage SFT recipe β€” Progressive training from audio-language alignment to full multimodal SFT, developing live-commerce understanding from omni-modal perception to instruction-following responses.
  • Faithful-RFT β€” A reinforcement fine-tuning stage for faithful and real-time live-stream demands, suppressing explicit reasoning traces and directly optimizing answer quality for live-commerce tasks.
  • Rich atomic capabilities β€” A scenario-oriented taxonomy covering speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc, supported by a compact data production engine.
  • Strong live-commerce performance with competitive generalization β€” 4B and 9B variants demonstrate strong results across live-commerce audio, image, and video tasks, together with excellent generalization on general benchmarks.

πŸ—οΈ Architecture

TLive-Omni is built on a Qwen3.5 backbone and extends it with a audio encoder through a lightweight MLP aligner, forming a unified text-output omni-modal understanding model. For video inputs with audio, each temporal grid is organized into a timestamped grid that interleaves video and audio token blocks, keeping audio segments adjacent to their corresponding visual content. The model supports up to 256K tokens of context at inference.

πŸ“Š Benchmark Results

We evaluate TLive-Omni-4B and TLive-Omni-9B on both live-commerce tasks and general benchmarks. Dash (-) denotes an unreported result or undisclosed parameter count. The Best results among the compared open-source models are marked in bold, while the second-best results are in underlined.

Live-Commerce Evaluation

Click to expand
TaskMetricTLive-Omni 4BTLive-Omni 9BGemini 2.5 FlashGemini 2.5 ProGemini 3 FlashGemini 3 ProGemini 3.5 FlashQwen3.5-Omni FlashOmniVinci 9BNemotron 3 Nano Omni 30B-A3BMing-Lite-Omni v1.5 20B-A3BMiniCPM-o 2.6 8BMiniCPM-o 4.5 9BQwen2.5-Omni 7BQwen3-Omni 30B-A3B
Audio
Live-Commerce ASRCER ↓6.666.4616.3011.4815.1812.0913.096.81β€”12.1010.0613.8810.707.866.75
Speaker-Attributed ASRcpWER ↓12.8812.2717.1412.1719.0411.6711.9913.23β€”17.65β€”β€”18.89β€”27.84
Audio DescriptionAcc. ↑76.1275.9665.2181.1068.2785.0779.9762.8239.9033.0145.9949.8447.5947.9261.06
Audio DescriptionHal. ↓20.9721.0026.1914.1626.1710.9214.3627.8147.3639.7744.0541.7652.4136.9230.22
Audio QAAcc. ↑72.6076.2876.2882.8574.6888.6287.9978.0466.5164.9040.5439.7442.4761.3876.76
Image
Visual GroundingLive AP ↑82.8582.3361.0851.9880.3873.8084.1579.9634.8673.0852.463.8223.9075.6179.22
Visual GroundingProd AP ↑91.4589.9628.8132.6365.6758.8374.8960.448.9348.6240.731.7753.6322.8568.88
Text UnderstandingLoc. F1 ↑86.9987.5920.5231.6061.1168.6064.4474.0750.2552.9113.275.745.4342.6430.46
Text UnderstandingRec. NED ↓4.724.2443.2827.8216.259.7216.6412.4832.7729.4259.1677.5871.6537.7914.83
Text UnderstandingCls. Acc. ↑79.0679.8551.2161.8669.1176.8669.7653.2557.2937.8632.9415.9211.6251.2569.46
Video
Temporal GroundingmIoU ↑77.6381.4976.5076.2277.4377.9077.9062.1013.1023.3914.3414.5643.2030.8339.22
Dense CaptionAcc. ↑69.2374.6354.6041.9532.2137.8033.8032.9418.5917.9613.8110.5321.0616.5121.44
Dense CaptionHal. ↓9.578.7610.9716.8820.7620.9917.3020.9127.1316.6239.3326.9328.6136.4425.82
Video QAAcc. ↑92.3193.2388.2192.6289.6484.3686.9087.2872.5182.5664.5160.3084.6275.4881.62
Shot UnderstandingLayout ↑78.4077.0080.0085.2076.8080.4083.4084.4073.6079.2074.6066.6078.2074.4082.20
Shot UnderstandingShot Size ↑51.2051.0046.8050.8045.7043.4044.2048.9052.7034.0041.7038.3042.8038.1037.40
Shot UnderstandingCamera ↑80.9082.0084.2076.0078.5075.7075.5085.5068.1080.2072.8068.3079.2081.2076.10
Shot UnderstandingContent ↑69.8071.0068.6070.4071.6074.8070.2066.4049.2058.6051.6049.0066.6068.0063.60

General Benchmark: Image Understanding

Click to expand
ModelParamsMMMUMathVistaDynaMathVLMsAreBlindMMBench(EN-DEV-v1.1)RealWorldQAMMStarSimpleVQAHallusionAI2DOCRBenchCC-OCRCharXiv(RQ)RefCOCOERQAEmbSpatial
Open-source VLM models
MiMo-VL-SFT7B64.681.846.978.084.5β€”β€”β€”β€”83.287.6β€”54.485.7β€”β€”
SAIL-VL28B55.476.417.8β€”β€”76.370.7β€”55.187.791.3β€”β€”74.0β€”β€”
Valley2.58B62.174.432.7β€”85.570.567.3β€”56.384.487.0β€”β€”β€”β€”β€”
LLaVA-OneVision-28Bβ€”β€”β€”β€”85.769.764.8β€”β€”84.378.2β€”β€”β€”43.378.1
InternVL3.54B66.677.135.7β€”80.366.365.0β€”44.882.682.2β€”39.689.438.5β€”
InternVL3.58B73.478.437.7β€”79.567.569.3β€”54.584.084.0β€”44.489.741.073.2
Qwen3-VL4B67.473.765.371.983.970.969.848.057.684.188.176.239.789.041.379.6
Qwen3-VL8B69.677.267.774.084.571.570.950.261.185.789.679.946.489.145.878.5
Qwen3.54B72.181.069.662.386.372.574.844.676.987.185.971.162.987.646.876.6
Qwen3.59B74.282.274.671.887.772.976.348.976.088.088.573.467.590.047.378.7
Open-source Omni models
InteractiveOmni4B61.161.7β€”β€”78.9β€”62.6β€”52.283.880.0β€”β€”β€”β€”β€”
InteractiveOmni8B66.968.0β€”β€”81.4β€”66.8β€”61.384.383.7β€”β€”β€”β€”β€”
VITA-1.57B52.166.2β€”β€”76.7β€”59.9β€”44.979.373.2β€”β€”β€”β€”β€”
Valley38B69.3β€”β€”β€”β€”β€”β€”β€”55.9β€”β€”β€”β€”β€”β€”β€”
OmniVinci9B49.763.5β€”β€”β€”67.5β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”
Nemotron 3 Nano Omni30B-A3B55.271.9β€”β€”β€”β€”β€”β€”β€”88.588.3β€”49.180.6β€”β€”
Ming-Lite-Omni v1.520B-A3B54.372.0β€”β€”β€”β€”65.1β€”54.684.988.9β€”β€”87.8β€”β€”
MiniCPM-o 2.68B50.471.9β€”β€”80.5β€”64.0β€”51.985.889.7β€”β€”β€”β€”β€”
MiniCPM-o 4.59B67.6β€”β€”β€”87.6β€”73.1β€”63.287.687.6β€”β€”β€”β€”β€”
Qwen2.5-Omni7B59.267.9β€”β€”81.870.364.0β€”β€”83.2β€”β€”β€”87.7β€”β€”
Qwen3-Omni30B-A3B69.175.9β€”β€”β€”β€”68.5β€”59.785.286.0β€”61.1β€”β€”β€”
Ours
TLive-Omni4B70.979.972.571.887.077.773.947.677.786.686.680.561.387.442.379.3
TLive-Omni9B73.481.973.375.588.976.675.150.076.088.690.381.363.190.048.080.4

General Benchmark: Video Understanding

Click to expand
ModelParamsMVBenchMLVUVideo-MMELongVideoBenchLVBenchMMVUVideoMMMUCharades-TLActivityNet-TLQVHighlights-TL
Open-source VLM models
MiMo-VL-SFT7Bβ€”β€”66.9β€”β€”β€”53.139.635.541.5
SAIL-VL28Bβ€”β€”62.758.3β€”β€”β€”β€”β€”β€”
LLaVA-OneVision-28B66.276.671.966.955.556.2β€”53.553.866.4
LLaVA-Video7B58.670.863.358.244.247.136.115.214.610.4
InternVL3.54B71.270.465.460.843.247.657.616.014.917.7
InternVL3.58B72.170.266.062.146.760.2β€”27.831.331.3
MiniCPM-V 4.58Bβ€”75.167.963.950.458.957.131.932.346.1
LongVU7B66.965.460.6β€”β€”β€”β€”β€”β€”β€”
LongVILA7B67.1β€”60.157.1β€”β€”β€”β€”β€”β€”
Mage-VL4B65.168.764.061.341.8β€”β€”50.745.457.4
Molmo24B75.163.069.668.053.951.250.733.339.858.7
Molmo28B75.960.269.967.552.8β€”β€”β€”β€”β€”
NVILA8B68.170.164.257.7β€”β€”β€”β€”β€”β€”
Kangaroo8B61.161.056.054.839.4β€”β€”β€”β€”β€”
Video-XL28Bβ€”74.866.661.048.450.039.938.930.046.2
VideoChat34Bβ€”β€”70.1β€”56.756.457.456.154.667.0
VideoLLaMA 37B69.773.066.259.845.344.134.639.829.836.9
Qwen3-VL4B68.975.369.3β€”56.250.556.246.448.258.7
Qwen3-VL8B68.778.171.4β€”58.058.765.348.346.859.4
Qwen3.54B66.675.171.665.155.357.869.848.751.655.0
Qwen3.59B75.779.766.967.960.963.770.352.054.057.2
Open-source Omni models
InteractiveOmni4Bβ€”68.063.357.0β€”β€”β€”β€”β€”β€”
InteractiveOmni8Bβ€”71.666.059.1β€”β€”β€”β€”β€”β€”
VITA-1.57B55.4β€”56.1β€”β€”β€”β€”β€”β€”β€”
Valley38Bβ€”55.6β€”β€”β€”β€”61.2β€”β€”β€”
OmniVinci9B70.6β€”68.261.3β€”β€”β€”β€”β€”β€”
Nemotron 3 Nano Omni30B-A3Bβ€”β€”70.8β€”β€”β€”β€”β€”β€”β€”
Ming-Lite-Omni v1.520B-A3B69.4β€”67.159.5β€”β€”β€”β€”β€”β€”
MiniCPM-o 2.68Bβ€”β€”63.9β€”β€”β€”β€”β€”β€”β€”
MiniCPM-o 4.59Bβ€”76.570.466.0β€”β€”β€”β€”β€”β€”
Qwen2.5-Omni7B70.3β€”64.3β€”β€”β€”β€”β€”β€”β€”
Qwen3-Omni30B-A3Bβ€”75.270.5β€”β€”β€”β€”β€”β€”β€”
Ours
TLive-Omni4B69.076.171.366.157.159.973.957.058.269.2
TLive-Omni9B72.580.975.669.960.867.172.856.355.464.1

General Benchmark: Omni Understanding

Click to expand
ModelParamsAVUTWorldSenseVideoHolmesDailyOmniOmniVideoBenchFutureOmni
Open-source Omni models
video-SALMONN 2+3B66.248.342.267.7β€”β€”
video-SALMONN 2+7B69.550.946.971.8β€”β€”
OmniVinci9Bβ€”48.2β€”66.536.752.8
Nemotron 3 Nano Omni30B-A3Bβ€”55.2β€”74.5β€”β€”
MiniCPM-o 4.59B78.655.764.380.241.156.1
Qwen2.5-Omni7Bβ€”45.4β€”62.436.548.9
Qwen3-Omni30B-A3B74.254.050.471.943.853.4
Ours
TLive-Omni4B78.654.057.578.641.657.2
TLive-Omni9B80.056.059.380.543.258.5

βš™οΈ Installation

This release targets Python 3.10 on Linux x86_64 with CUDA 12.8 and PyTorch 2.10.0.

conda create -n tlive python=3.10 -y
conda activate tlive
pip install -r https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/v1.0.0-rc1/environments/requirements.txt

The remote environments/requirements.txt includes custom wheels for the supported environment and model. If any wheel does not match your hardware, CUDA version, or Python version, replace it with a compatible build for your setup.

πŸš€ Quick Start

Transformers inference

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

model_id = "TaoLiveAIGC/TLive-Omni-9B"

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="flash_attention_2",
).eval()


def generate(messages, *, use_audio_in_video=False, videos_kwargs=None, generation_kwargs=None):
    inputs = processor.apply_chat_template(
        messages,
        add_generation_prompt=True,
        tokenize=True,
        return_dict=True,
        return_tensors="pt",
        enable_thinking=False,
        use_audio_in_video=use_audio_in_video,
        videos_kwargs=videos_kwargs or {},
    )
    prompt_length = inputs["input_ids"].shape[-1]
    inputs = inputs.to(model.device)

    generation_kwargs = generation_kwargs or {}
    with torch.inference_mode():
        generated_ids = model.generate(
            **inputs,
            do_sample=False,
            max_new_tokens=1024,
            **generation_kwargs,
        )

    answer_ids = generated_ids[:, prompt_length:]
    answer = processor.batch_decode(
        answer_ids,
        skip_special_tokens=True,
        clean_up_tokenization_spaces=False,
    )[0]
    return answer.strip()

Replace messages with one of the examples below for text, image, audio, or video inputs.

Text

messages = [{
    "role": "user",
    "content": [{"type": "text", "text": "Briefly explain why multimodal context can improve an answer."}],
}]
print(generate(messages))

Image

messages = [{
    "role": "user",
    "content": [
        {"type": "image", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/image.jpg"},
        {"type": "text", "text": "Describe this image."},
    ],
}]
print(generate(messages))

Audio

messages = [{
    "role": "user",
    "content": [
        {"type": "audio", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/audio.mp3"},
        {"type": "text", "text": "Transcribe and summarize this audio."},
    ],
}]
print(generate(messages))

Video with audio

messages = [{
    "role": "user",
    "content": [
        {"type": "video", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/vocal_video.mp4"},
        {"type": "text", "text": "Describe the video, including relevant speech and sounds."},
    ],
}]
print(generate(messages, use_audio_in_video=True, videos_kwargs={"fps": 1.0}))

Video without audio

messages = [{
    "role": "user",
    "content": [
        {"type": "video", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/silence_video.mp4"},
        {"type": "text", "text": "Describe the visual events in this video."},
    ],
}]
print(generate(messages, use_audio_in_video=False, videos_kwargs={"fps": 1.0}))

For temporal localization outputs, we recommend the MM:SS - MM:SS interval format, for example 01:23 - 01:35. For videos, set use_audio_in_video=True when the audio track should be used, and False for visual-only inference.

⚑ vLLM

Installation

First install the pre-built wheel (Python 3.10 + CUDA 12.8 + Linux x86_64), built and tested on NVIDIA H20 GPUs (Hopper, sm_90):

pip install https://github.com/TaoLiveAIGC/TLive-Omni/releases/download/v1.0.0-rc1/vllm-0.19.0+cu128-cp310-cp310-linux_x86_64.whl

If your GPU, driver, or CUDA setup is not compatible with this wheel, build vLLM from source using the customized code in the vllm/ directory of the GitHub release.

Inference

from transformers import AutoProcessor
from vllm import LLM, SamplingParams
from vllm.model_executor.models.tlive_omni_processing import process_audio_info

model_id = "TaoLiveAIGC/TLive-Omni-9B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)


def build_prompt(messages):
    return processor.apply_chat_template(
        messages,
        add_generation_prompt=True,
        tokenize=False,
        enable_thinking=False,
    )


def generate(inputs, *, limit_mm_per_prompt=None, vllm_kwargs=None, sampling_kwargs=None):
    vllm_kwargs = vllm_kwargs or {}
    sampling_kwargs = sampling_kwargs or {}
    llm = LLM(
        model=model_id,
        trust_remote_code=True,
        dtype="bfloat16",
        max_model_len=32768,
        tensor_parallel_size=1,
        gpu_memory_utilization=0.9,
        max_num_seqs=4,
        max_num_batched_tokens=32768,
        seed=42,
        limit_mm_per_prompt=limit_mm_per_prompt,
        **vllm_kwargs,
    )
    outputs = llm.generate(
        inputs,
        sampling_params=SamplingParams(
            temperature=0.0,
            max_tokens=1024,
            **sampling_kwargs,
        ),
    )
    return outputs[0].outputs[0].text.strip()

Replace messages with one of the examples below for text, image, audio, or video inputs.

Text

messages = [{
    "role": "user",
    "content": [{"type": "text", "text": "Briefly explain why multimodal context can improve an answer."}],
}]
inputs = {"prompt": build_prompt(messages)}
print(generate(inputs))

Image

image_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/image.jpg"
messages = [{
    "role": "user",
    "content": [
        {"type": "image", "path": image_path},
        {"type": "text", "text": "Describe this image."},
    ],
}]
inputs = {
    "prompt": build_prompt(messages),
    "multi_modal_data": {"image": [image_path]},
}
print(generate(inputs, limit_mm_per_prompt={"image": 1}))

Audio

audio_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/audio.mp3"
messages = [{
    "role": "user",
    "content": [
        {"type": "audio", "audio": audio_path},
        {"type": "text", "text": "Transcribe and summarize this audio."},
    ],
}]
inputs = {
    "prompt": build_prompt(messages),
    "multi_modal_data": {"audio": process_audio_info(messages, use_audio_in_video=False)},
}
print(generate(inputs, limit_mm_per_prompt={"audio": 1}))

Video with audio

video_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/vocal_video.mp4"
messages = [{
    "role": "user",
    "content": [
        {"type": "video", "video": video_path},
        {"type": "text", "text": "Describe the video, including relevant speech and sounds."},
    ],
}]
inputs = {
    "prompt": build_prompt(messages),
    "multi_modal_data": {
        "video": [video_path],
        "audio": process_audio_info(messages, use_audio_in_video=True),
    },
    "mm_processor_kwargs": {"videos_kwargs": {"fps": 1.0, "use_audio_in_video": True, "return_metadata": True}},
}
print(generate(inputs, limit_mm_per_prompt={"video": 1, "audio": 1}))

Video without audio

video_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/silence_video.mp4"
messages = [{
    "role": "user",
    "content": [
        {"type": "video", "video": video_path},
        {"type": "text", "text": "Describe the visual events in this video."},
    ],
}]
inputs = {
    "prompt": build_prompt(messages),
    "multi_modal_data": {"video": [video_path]},
    "mm_processor_kwargs": {"videos_kwargs": {"fps": 1.0, "use_audio_in_video": False, "return_metadata": True}},
}
print(generate(inputs, limit_mm_per_prompt={"video": 1}))

πŸ“– Citation

If you find our work helpful, please consider citing our paper:

@article{hu2026tliveomni,
  title={TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming},
  author={Hu, Yibo and Qian, Yu and Gu, Mao and Tao, Yingfan and Chen, Yuhao and Luo, Yongdong and Liu, Zhuoqun and Jin, Meiguang and Ma, Junfeng},
  journal={arXiv preprint arXiv:2608.20958},
  year={2026}
}

πŸ“„ License

This project is released under the Apache License 2.0.

Contributors

GM
GMadeus

1 commits

Leon1207

1 commits

TaoLiveAIGC/TLive-Omni-9B

Model

15

stars

2

commits

1

linked in READMEs

Aug 24, 2026

updated

audio
conversational
custom-code
custom_code
multimodal
safetensors
text-generation
tlive_omni
transformers
video
vision
Browse cluster: Multimedia processing and video benchmarking β†’

README

TLive-Omni

logo

Technical Report Hugging Face 4B Model Hugging Face 9B Model License

πŸ“‹ Overview

TLive-Omni is an omni-modal understanding model for e-commerce live-stream, mapping image, video, audio, and text into a unified text-output interface. Built on a Qwen3.5 backbone with a grafted AuT audio encoder, it supports up to 256K tokens of context, trained via a three-stage SFT recipe followed by Faithful-RFT reinforcement fine-tuning.

✨ Highlights

  • Timestamped Per-vGrid layout β€” Audio and video tokens are organized into timestamped grid with explicit boundaries, keeping audio segments adjacent to their corresponding visual content for fine-grained temporal alignment over long streams.
  • Three-stage SFT recipe β€” Progressive training from audio-language alignment to full multimodal SFT, developing live-commerce understanding from omni-modal perception to instruction-following responses.
  • Faithful-RFT β€” A reinforcement fine-tuning stage for faithful and real-time live-stream demands, suppressing explicit reasoning traces and directly optimizing answer quality for live-commerce tasks.
  • Rich atomic capabilities β€” A scenario-oriented taxonomy covering speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc, supported by a compact data production engine.
  • Strong live-commerce performance with competitive generalization β€” 4B and 9B variants demonstrate strong results across live-commerce audio, image, and video tasks, together with excellent generalization on general benchmarks.

πŸ—οΈ Architecture

TLive-Omni is built on a Qwen3.5 backbone and extends it with a audio encoder through a lightweight MLP aligner, forming a unified text-output omni-modal understanding model. For video inputs with audio, each temporal grid is organized into a timestamped grid that interleaves video and audio token blocks, keeping audio segments adjacent to their corresponding visual content. The model supports up to 256K tokens of context at inference.

πŸ“Š Benchmark Results

We evaluate TLive-Omni-4B and TLive-Omni-9B on both live-commerce tasks and general benchmarks. Dash (-) denotes an unreported result or undisclosed parameter count. The Best results among the compared open-source models are marked in bold, while the second-best results are in underlined.

Live-Commerce Evaluation

Click to expand
TaskMetricTLive-Omni 4BTLive-Omni 9BGemini 2.5 FlashGemini 2.5 ProGemini 3 FlashGemini 3 ProGemini 3.5 FlashQwen3.5-Omni FlashOmniVinci 9BNemotron 3 Nano Omni 30B-A3BMing-Lite-Omni v1.5 20B-A3BMiniCPM-o 2.6 8BMiniCPM-o 4.5 9BQwen2.5-Omni 7BQwen3-Omni 30B-A3B
Audio
Live-Commerce ASRCER ↓6.666.4616.3011.4815.1812.0913.096.81β€”12.1010.0613.8810.707.866.75
Speaker-Attributed ASRcpWER ↓12.8812.2717.1412.1719.0411.6711.9913.23β€”17.65β€”β€”18.89β€”27.84
Audio DescriptionAcc. ↑76.1275.9665.2181.1068.2785.0779.9762.8239.9033.0145.9949.8447.5947.9261.06
Audio DescriptionHal. ↓20.9721.0026.1914.1626.1710.9214.3627.8147.3639.7744.0541.7652.4136.9230.22
Audio QAAcc. ↑72.6076.2876.2882.8574.6888.6287.9978.0466.5164.9040.5439.7442.4761.3876.76
Image
Visual GroundingLive AP ↑82.8582.3361.0851.9880.3873.8084.1579.9634.8673.0852.463.8223.9075.6179.22
Visual GroundingProd AP ↑91.4589.9628.8132.6365.6758.8374.8960.448.9348.6240.731.7753.6322.8568.88
Text UnderstandingLoc. F1 ↑86.9987.5920.5231.6061.1168.6064.4474.0750.2552.9113.275.745.4342.6430.46
Text UnderstandingRec. NED ↓4.724.2443.2827.8216.259.7216.6412.4832.7729.4259.1677.5871.6537.7914.83
Text UnderstandingCls. Acc. ↑79.0679.8551.2161.8669.1176.8669.7653.2557.2937.8632.9415.9211.6251.2569.46
Video
Temporal GroundingmIoU ↑77.6381.4976.5076.2277.4377.9077.9062.1013.1023.3914.3414.5643.2030.8339.22
Dense CaptionAcc. ↑69.2374.6354.6041.9532.2137.8033.8032.9418.5917.9613.8110.5321.0616.5121.44
Dense CaptionHal. ↓9.578.7610.9716.8820.7620.9917.3020.9127.1316.6239.3326.9328.6136.4425.82
Video QAAcc. ↑92.3193.2388.2192.6289.6484.3686.9087.2872.5182.5664.5160.3084.6275.4881.62
Shot UnderstandingLayout ↑78.4077.0080.0085.2076.8080.4083.4084.4073.6079.2074.6066.6078.2074.4082.20
Shot UnderstandingShot Size ↑51.2051.0046.8050.8045.7043.4044.2048.9052.7034.0041.7038.3042.8038.1037.40
Shot UnderstandingCamera ↑80.9082.0084.2076.0078.5075.7075.5085.5068.1080.2072.8068.3079.2081.2076.10
Shot UnderstandingContent ↑69.8071.0068.6070.4071.6074.8070.2066.4049.2058.6051.6049.0066.6068.0063.60

General Benchmark: Image Understanding

Click to expand
ModelParamsMMMUMathVistaDynaMathVLMsAreBlindMMBench(EN-DEV-v1.1)RealWorldQAMMStarSimpleVQAHallusionAI2DOCRBenchCC-OCRCharXiv(RQ)RefCOCOERQAEmbSpatial
Open-source VLM models
MiMo-VL-SFT7B64.681.846.978.084.5β€”β€”β€”β€”83.287.6β€”54.485.7β€”β€”
SAIL-VL28B55.476.417.8β€”β€”76.370.7β€”55.187.791.3β€”β€”74.0β€”β€”
Valley2.58B62.174.432.7β€”85.570.567.3β€”56.384.487.0β€”β€”β€”β€”β€”
LLaVA-OneVision-28Bβ€”β€”β€”β€”85.769.764.8β€”β€”84.378.2β€”β€”β€”43.378.1
InternVL3.54B66.677.135.7β€”80.366.365.0β€”44.882.682.2β€”39.689.438.5β€”
InternVL3.58B73.478.437.7β€”79.567.569.3β€”54.584.084.0β€”44.489.741.073.2
Qwen3-VL4B67.473.765.371.983.970.969.848.057.684.188.176.239.789.041.379.6
Qwen3-VL8B69.677.267.774.084.571.570.950.261.185.789.679.946.489.145.878.5
Qwen3.54B72.181.069.662.386.372.574.844.676.987.185.971.162.987.646.876.6
Qwen3.59B74.282.274.671.887.772.976.348.976.088.088.573.467.590.047.378.7
Open-source Omni models
InteractiveOmni4B61.161.7β€”β€”78.9β€”62.6β€”52.283.880.0β€”β€”β€”β€”β€”
InteractiveOmni8B66.968.0β€”β€”81.4β€”66.8β€”61.384.383.7β€”β€”β€”β€”β€”
VITA-1.57B52.166.2β€”β€”76.7β€”59.9β€”44.979.373.2β€”β€”β€”β€”β€”
Valley38B69.3β€”β€”β€”β€”β€”β€”β€”55.9β€”β€”β€”β€”β€”β€”β€”
OmniVinci9B49.763.5β€”β€”β€”67.5β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”
Nemotron 3 Nano Omni30B-A3B55.271.9β€”β€”β€”β€”β€”β€”β€”88.588.3β€”49.180.6β€”β€”
Ming-Lite-Omni v1.520B-A3B54.372.0β€”β€”β€”β€”65.1β€”54.684.988.9β€”β€”87.8β€”β€”
MiniCPM-o 2.68B50.471.9β€”β€”80.5β€”64.0β€”51.985.889.7β€”β€”β€”β€”β€”
MiniCPM-o 4.59B67.6β€”β€”β€”87.6β€”73.1β€”63.287.687.6β€”β€”β€”β€”β€”
Qwen2.5-Omni7B59.267.9β€”β€”81.870.364.0β€”β€”83.2β€”β€”β€”87.7β€”β€”
Qwen3-Omni30B-A3B69.175.9β€”β€”β€”β€”68.5β€”59.785.286.0β€”61.1β€”β€”β€”
Ours
TLive-Omni4B70.979.972.571.887.077.773.947.677.786.686.680.561.387.442.379.3
TLive-Omni9B73.481.973.375.588.976.675.150.076.088.690.381.363.190.048.080.4

General Benchmark: Video Understanding

Click to expand
ModelParamsMVBenchMLVUVideo-MMELongVideoBenchLVBenchMMVUVideoMMMUCharades-TLActivityNet-TLQVHighlights-TL
Open-source VLM models
MiMo-VL-SFT7Bβ€”β€”66.9β€”β€”β€”53.139.635.541.5
SAIL-VL28Bβ€”β€”62.758.3β€”β€”β€”β€”β€”β€”
LLaVA-OneVision-28B66.276.671.966.955.556.2β€”53.553.866.4
LLaVA-Video7B58.670.863.358.244.247.136.115.214.610.4
InternVL3.54B71.270.465.460.843.247.657.616.014.917.7
InternVL3.58B72.170.266.062.146.760.2β€”27.831.331.3
MiniCPM-V 4.58Bβ€”75.167.963.950.458.957.131.932.346.1
LongVU7B66.965.460.6β€”β€”β€”β€”β€”β€”β€”
LongVILA7B67.1β€”60.157.1β€”β€”β€”β€”β€”β€”
Mage-VL4B65.168.764.061.341.8β€”β€”50.745.457.4
Molmo24B75.163.069.668.053.951.250.733.339.858.7
Molmo28B75.960.269.967.552.8β€”β€”β€”β€”β€”
NVILA8B68.170.164.257.7β€”β€”β€”β€”β€”β€”
Kangaroo8B61.161.056.054.839.4β€”β€”β€”β€”β€”
Video-XL28Bβ€”74.866.661.048.450.039.938.930.046.2
VideoChat34Bβ€”β€”70.1β€”56.756.457.456.154.667.0
VideoLLaMA 37B69.773.066.259.845.344.134.639.829.836.9
Qwen3-VL4B68.975.369.3β€”56.250.556.246.448.258.7
Qwen3-VL8B68.778.171.4β€”58.058.765.348.346.859.4
Qwen3.54B66.675.171.665.155.357.869.848.751.655.0
Qwen3.59B75.779.766.967.960.963.770.352.054.057.2
Open-source Omni models
InteractiveOmni4Bβ€”68.063.357.0β€”β€”β€”β€”β€”β€”
InteractiveOmni8Bβ€”71.666.059.1β€”β€”β€”β€”β€”β€”
VITA-1.57B55.4β€”56.1β€”β€”β€”β€”β€”β€”β€”
Valley38Bβ€”55.6β€”β€”β€”β€”61.2β€”β€”β€”
OmniVinci9B70.6β€”68.261.3β€”β€”β€”β€”β€”β€”
Nemotron 3 Nano Omni30B-A3Bβ€”β€”70.8β€”β€”β€”β€”β€”β€”β€”
Ming-Lite-Omni v1.520B-A3B69.4β€”67.159.5β€”β€”β€”β€”β€”β€”
MiniCPM-o 2.68Bβ€”β€”63.9β€”β€”β€”β€”β€”β€”β€”
MiniCPM-o 4.59Bβ€”76.570.466.0β€”β€”β€”β€”β€”β€”
Qwen2.5-Omni7B70.3β€”64.3β€”β€”β€”β€”β€”β€”β€”
Qwen3-Omni30B-A3Bβ€”75.270.5β€”β€”β€”β€”β€”β€”β€”
Ours
TLive-Omni4B69.076.171.366.157.159.973.957.058.269.2
TLive-Omni9B72.580.975.669.960.867.172.856.355.464.1

General Benchmark: Omni Understanding

Click to expand
ModelParamsAVUTWorldSenseVideoHolmesDailyOmniOmniVideoBenchFutureOmni
Open-source Omni models
video-SALMONN 2+3B66.248.342.267.7β€”β€”
video-SALMONN 2+7B69.550.946.971.8β€”β€”
OmniVinci9Bβ€”48.2β€”66.536.752.8
Nemotron 3 Nano Omni30B-A3Bβ€”55.2β€”74.5β€”β€”
MiniCPM-o 4.59B78.655.764.380.241.156.1
Qwen2.5-Omni7Bβ€”45.4β€”62.436.548.9
Qwen3-Omni30B-A3B74.254.050.471.943.853.4
Ours
TLive-Omni4B78.654.057.578.641.657.2
TLive-Omni9B80.056.059.380.543.258.5

βš™οΈ Installation

This release targets Python 3.10 on Linux x86_64 with CUDA 12.8 and PyTorch 2.10.0.

conda create -n tlive python=3.10 -y
conda activate tlive
pip install -r https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/v1.0.0-rc1/environments/requirements.txt

The remote environments/requirements.txt includes custom wheels for the supported environment and model. If any wheel does not match your hardware, CUDA version, or Python version, replace it with a compatible build for your setup.

πŸš€ Quick Start

Transformers inference

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

model_id = "TaoLiveAIGC/TLive-Omni-9B"

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="flash_attention_2",
).eval()


def generate(messages, *, use_audio_in_video=False, videos_kwargs=None, generation_kwargs=None):
    inputs = processor.apply_chat_template(
        messages,
        add_generation_prompt=True,
        tokenize=True,
        return_dict=True,
        return_tensors="pt",
        enable_thinking=False,
        use_audio_in_video=use_audio_in_video,
        videos_kwargs=videos_kwargs or {},
    )
    prompt_length = inputs["input_ids"].shape[-1]
    inputs = inputs.to(model.device)

    generation_kwargs = generation_kwargs or {}
    with torch.inference_mode():
        generated_ids = model.generate(
            **inputs,
            do_sample=False,
            max_new_tokens=1024,
            **generation_kwargs,
        )

    answer_ids = generated_ids[:, prompt_length:]
    answer = processor.batch_decode(
        answer_ids,
        skip_special_tokens=True,
        clean_up_tokenization_spaces=False,
    )[0]
    return answer.strip()

Replace messages with one of the examples below for text, image, audio, or video inputs.

Text

messages = [{
    "role": "user",
    "content": [{"type": "text", "text": "Briefly explain why multimodal context can improve an answer."}],
}]
print(generate(messages))

Image

messages = [{
    "role": "user",
    "content": [
        {"type": "image", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/image.jpg"},
        {"type": "text", "text": "Describe this image."},
    ],
}]
print(generate(messages))

Audio

messages = [{
    "role": "user",
    "content": [
        {"type": "audio", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/audio.mp3"},
        {"type": "text", "text": "Transcribe and summarize this audio."},
    ],
}]
print(generate(messages))

Video with audio

messages = [{
    "role": "user",
    "content": [
        {"type": "video", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/vocal_video.mp4"},
        {"type": "text", "text": "Describe the video, including relevant speech and sounds."},
    ],
}]
print(generate(messages, use_audio_in_video=True, videos_kwargs={"fps": 1.0}))

Video without audio

messages = [{
    "role": "user",
    "content": [
        {"type": "video", "path": "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/silence_video.mp4"},
        {"type": "text", "text": "Describe the visual events in this video."},
    ],
}]
print(generate(messages, use_audio_in_video=False, videos_kwargs={"fps": 1.0}))

For temporal localization outputs, we recommend the MM:SS - MM:SS interval format, for example 01:23 - 01:35. For videos, set use_audio_in_video=True when the audio track should be used, and False for visual-only inference.

⚑ vLLM

Installation

First install the pre-built wheel (Python 3.10 + CUDA 12.8 + Linux x86_64), built and tested on NVIDIA H20 GPUs (Hopper, sm_90):

pip install https://github.com/TaoLiveAIGC/TLive-Omni/releases/download/v1.0.0-rc1/vllm-0.19.0+cu128-cp310-cp310-linux_x86_64.whl

If your GPU, driver, or CUDA setup is not compatible with this wheel, build vLLM from source using the customized code in the vllm/ directory of the GitHub release.

Inference

from transformers import AutoProcessor
from vllm import LLM, SamplingParams
from vllm.model_executor.models.tlive_omni_processing import process_audio_info

model_id = "TaoLiveAIGC/TLive-Omni-9B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)


def build_prompt(messages):
    return processor.apply_chat_template(
        messages,
        add_generation_prompt=True,
        tokenize=False,
        enable_thinking=False,
    )


def generate(inputs, *, limit_mm_per_prompt=None, vllm_kwargs=None, sampling_kwargs=None):
    vllm_kwargs = vllm_kwargs or {}
    sampling_kwargs = sampling_kwargs or {}
    llm = LLM(
        model=model_id,
        trust_remote_code=True,
        dtype="bfloat16",
        max_model_len=32768,
        tensor_parallel_size=1,
        gpu_memory_utilization=0.9,
        max_num_seqs=4,
        max_num_batched_tokens=32768,
        seed=42,
        limit_mm_per_prompt=limit_mm_per_prompt,
        **vllm_kwargs,
    )
    outputs = llm.generate(
        inputs,
        sampling_params=SamplingParams(
            temperature=0.0,
            max_tokens=1024,
            **sampling_kwargs,
        ),
    )
    return outputs[0].outputs[0].text.strip()

Replace messages with one of the examples below for text, image, audio, or video inputs.

Text

messages = [{
    "role": "user",
    "content": [{"type": "text", "text": "Briefly explain why multimodal context can improve an answer."}],
}]
inputs = {"prompt": build_prompt(messages)}
print(generate(inputs))

Image

image_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/image.jpg"
messages = [{
    "role": "user",
    "content": [
        {"type": "image", "path": image_path},
        {"type": "text", "text": "Describe this image."},
    ],
}]
inputs = {
    "prompt": build_prompt(messages),
    "multi_modal_data": {"image": [image_path]},
}
print(generate(inputs, limit_mm_per_prompt={"image": 1}))

Audio

audio_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/audio.mp3"
messages = [{
    "role": "user",
    "content": [
        {"type": "audio", "audio": audio_path},
        {"type": "text", "text": "Transcribe and summarize this audio."},
    ],
}]
inputs = {
    "prompt": build_prompt(messages),
    "multi_modal_data": {"audio": process_audio_info(messages, use_audio_in_video=False)},
}
print(generate(inputs, limit_mm_per_prompt={"audio": 1}))

Video with audio

video_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/vocal_video.mp4"
messages = [{
    "role": "user",
    "content": [
        {"type": "video", "video": video_path},
        {"type": "text", "text": "Describe the video, including relevant speech and sounds."},
    ],
}]
inputs = {
    "prompt": build_prompt(messages),
    "multi_modal_data": {
        "video": [video_path],
        "audio": process_audio_info(messages, use_audio_in_video=True),
    },
    "mm_processor_kwargs": {"videos_kwargs": {"fps": 1.0, "use_audio_in_video": True, "return_metadata": True}},
}
print(generate(inputs, limit_mm_per_prompt={"video": 1, "audio": 1}))

Video without audio

video_path = "https://raw.githubusercontent.com/TaoLiveAIGC/TLive-Omni/main/data/silence_video.mp4"
messages = [{
    "role": "user",
    "content": [
        {"type": "video", "video": video_path},
        {"type": "text", "text": "Describe the visual events in this video."},
    ],
}]
inputs = {
    "prompt": build_prompt(messages),
    "multi_modal_data": {"video": [video_path]},
    "mm_processor_kwargs": {"videos_kwargs": {"fps": 1.0, "use_audio_in_video": False, "return_metadata": True}},
}
print(generate(inputs, limit_mm_per_prompt={"video": 1}))

πŸ“– Citation

If you find our work helpful, please consider citing our paper:

@article{hu2026tliveomni,
  title={TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming},
  author={Hu, Yibo and Qian, Yu and Gu, Mao and Tao, Yingfan and Chen, Yuhao and Luo, Yongdong and Liu, Zhuoqun and Jin, Meiguang and Ma, Junfeng},
  journal={arXiv preprint arXiv:2608.20958},
  year={2026}
}

πŸ“„ License

This project is released under the Apache License 2.0.

Contributors

GM
GMadeus

1 commits

Leon1207

1 commits