TaoLiveAIGC/TLive-Omni

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

118

stars

4

commits

Python

primary language

Aug 24, 2026

updated

audio-understanding
large-language-models
multimodal-large-language-models
omni
omni-modal
omni-modal-video-understanding
video-understanding
Browse cluster: Vision-Language Models & Multimodal AI

README

🌟 TLive-Omni 🌟: An Omni-Modal Understanding Model for E-Commerce Live Streaming

logo

Technical Report Hugging Face 4B Model Hugging Face 9B Model License

📰 News

  • 📄 [2026-08-24] Technical ReportTLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming is now available on arXiv!
  • 🔥 [2026-08-20] Model Release — TLive-Omni is now available on Hugging Face! With TLive-Omni-4B and TLive-Omni-9B.

📋 Overview

TLive-Omni is an omni-modal understanding model for e-commerce live-stream, mapping image, video, audio, and text into a unified text-output interface. Built on a Qwen3.5 backbone with a grafted AuT audio encoder, it supports 256K tokens of context, trained via a three-stage SFT recipe followed by Faithful-RFT reinforcement fine-tuning.

✨ Highlights

  • Timestamped Per-vGrid layout — Audio and video tokens are organized into timestamped grid with explicit boundaries, keeping audio segments adjacent to their corresponding visual content for fine-grained temporal alignment over long streams.
  • Three-stage SFT recipe — Progressive training from audio-language alignment to full multimodal SFT, developing live-commerce understanding from omni-modal perception to instruction-following responses.
  • Faithful-RFT — A reinforcement fine-tuning stage for faithful and real-time live-stream demands, suppressing explicit reasoning traces and directly optimizing answer quality for live-commerce tasks.
  • Rich atomic capabilities — A scenario-oriented taxonomy covering speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc, supported by a compact data production engine.
  • Strong live-commerce performance with competitive generalization — 4B and 9B variants demonstrate strong results across live-commerce audio, image, and video tasks, together with excellent generalization on general benchmarks.

🏗️ Architecture

architecture

TLive-Omni is built on a Qwen3.5 backbone and extends it with a audio encoder through a lightweight MLP aligner, forming a unified text-output omni-modal understanding model. For video inputs with audio, each temporal grid is organized into a timestamped grid that interleaves video and audio token blocks, keeping audio segments adjacent to their corresponding visual content. The model supports up to 256K tokens of context at inference.

📊 Benchmark Results

We evaluate TLive-Omni-4B and TLive-Omni-9B on both live-commerce tasks and general benchmarks. Dash (-) denotes an unreported result or undisclosed parameter count. The Best results among open-source models are marked in bold, while the second-best results are in underlined.

Live-Commerce Audio Evaluation

ModelParamsLive ASR
CER ↓
Spk. ASR
cpWER ↓
Audio DescriptionAudio QA
Acc. ↑
Acc. ↑Hal. ↓
Closed‑source Omni models
Gemini 2.5 Flash16.3017.1465.2126.1976.28
Gemini 2.5 Pro11.4812.1781.1014.1682.85
Gemini 3 Flash15.1819.0468.2726.1774.68
Gemini 3 Pro12.0911.6785.0710.9288.62
Gemini 3.5 Flash13.0911.9979.9714.3687.99
Qwen3.5‑Omni Flash6.8113.2362.8227.8178.04
Open‑source Audio models
MiMo‑Audio7B12.7164.2626.0170.97
Fun‑Audio‑Chat8B14.5561.3532.2169.71
Step‑Audio‑R1.132B10.2175.0820.5069.80
Open‑source Omni models
OmniVinci9B39.9047.3666.51
Nemotron 3 Nano Omni30B‑A3B12.1017.6533.0139.7764.90
Ming‑Lite‑Omni v1.520B‑A3B10.0645.9944.0540.54
MiniCPM‑o 2.68B13.8849.8441.7639.74
MiniCPM‑o 4.59B10.7018.8947.5952.4142.47
Qwen2.5‑Omni7B7.8647.9236.9261.38
Qwen3‑Omni30B‑A3B6.7527.8461.0630.2276.76
Ours
TLive‑Omni4B6.6612.8876.1220.9772.60
TLive‑Omni9B6.4612.2775.9621.0076.28

Notes: Live ASR and Spk. ASR denote live-commerce ASR and speaker-attributed ASR.

Live-Commerce Image Evaluation

ModelParamsVisual GroundingText Understanding
Live AP ↑Prod AP ↑Loc. F1 ↑Rec. NED ↓Cls. Acc. ↑
Closed‑source Omni models
Gemini 2.5 Flash61.0828.8120.5243.2851.21
Gemini 2.5 Pro51.9832.6331.6027.8261.86
Gemini 3 Flash80.3865.6761.1116.2569.11
Gemini 3 Pro73.8058.8368.609.7276.86
Gemini 3.5 Flash84.1574.8964.4416.6469.76
Qwen3.5‑Omni Flash79.9660.4474.0712.4853.25
Open‑source Omni models
OmniVinci9B34.868.9350.2532.7757.29
Nemotron 3 Nano Omni30B‑A3B73.0848.6252.9129.4237.86
Ming‑Lite‑Omni v1.520B‑A3B52.4640.7313.2759.1632.94
MiniCPM‑o 2.68B3.821.775.7477.5815.92
MiniCPM‑o 4.59B23.9053.635.4371.6511.62
Qwen2.5‑Omni7B75.6122.8542.6437.7951.25
Qwen3‑Omni30B‑A3B79.2268.8830.4614.8369.46
Ours
TLive‑Omni4B82.8591.4586.994.7279.06
TLive‑Omni9B82.3389.9687.594.2479.85

Live-Commerce Video Evaluation

ModelParamsTG
mIoU ↑
Dense CaptionVideo QA
Acc. ↑
Shot Understanding
Acc. ↑Hal. ↓Layout ↑Shot Size ↑Camera ↑Content ↑
Closed‑source Omni models
Gemini 2.5 Flash76.5054.6010.9788.2180.0046.8084.2068.60
Gemini 2.5 Pro76.2241.9516.8892.6285.2050.8076.0070.40
Gemini 3 Flash77.4332.2120.7689.6476.8045.7078.5071.60
Gemini 3 Pro77.9037.8020.9984.3680.4043.4075.7074.80
Gemini 3.5 Flash77.9033.8017.3086.9083.4044.2075.5070.20
Qwen3.5‑Omni Flash62.1032.9420.9187.2884.4048.9085.5066.40
Open‑source Omni models
OmniVinci9B13.1018.5927.1372.5173.6052.7068.1049.20
Nemotron 3 Nano Omni30B‑A3B23.3917.9616.6282.5679.2034.0080.2058.60
Ming‑Lite‑Omni v1.520B‑A3B14.3413.8139.3364.5174.6041.7072.8051.60
MiniCPM‑o 2.68B14.5610.5326.9360.3066.6038.3068.3049.00
MiniCPM‑o 4.59B43.2021.0628.6184.6278.2042.8079.2066.60
Qwen2.5‑Omni7B30.8316.5136.4475.4874.4038.1081.2068.00
Qwen3‑Omni30B‑A3B39.2221.4425.8281.6282.2037.4076.1063.60
Ours
TLive‑Omni4B77.6369.239.5792.3178.4051.2080.9069.80
TLive‑Omni9B81.4974.638.7693.2377.0051.0082.0071.00

General Benchmark: Image Reasoning & QA

ModelParamsMMMUMathVistaDynaMathVABMMBenchRWQAMMStarSimpleVQA
Closed‑source models
Gemini 2.5 Flash76.375.369.775.986.675.775.859.2
Gemini 2.5 Pro80.977.778.578.588.476.078.566.9
Gemini 3 Pro87.287.985.193.783.383.173.2
GPT‑4o70.763.854.486.0
GPT‑5 (minimal)74.450.974.053.481.377.365.256.7
Qwen3.5‑Omni Flash76.982.979.388.877.575.754.4
Open‑source VLM models
MiMo‑VL‑SFT7B64.681.846.978.084.5
SAIL‑VL28B55.476.417.876.370.7
Valley2.58B62.174.432.785.570.567.3
LLaVA‑OneVision‑28B85.769.764.8
InternVL3.54B66.677.135.780.366.365.0
InternVL3.58B73.478.437.779.567.569.3
Qwen3‑VL4B67.473.765.371.983.970.969.848.0
Qwen3‑VL8B69.677.267.774.084.571.570.950.2
Qwen3.54B72.181.069.662.386.372.574.844.6
Qwen3.59B74.282.274.671.887.772.976.348.9
Open‑source Omni models
InteractiveOmni4B61.161.778.962.6
InteractiveOmni8B66.968.081.466.8
VITA‑1.57B52.166.276.759.9
Valley38B69.3
OmniVinci9B49.763.567.5
Nemotron 3 Nano Omni30B‑A3B55.271.9
Ming‑Lite‑Omni v1.520B‑A3B54.372.065.1
MiniCPM‑o 2.68B50.471.980.564.0
MiniCPM‑o 4.59B67.687.673.1
Qwen2.5‑Omni7B59.267.981.870.364.0
Qwen3‑Omni30B‑A3B69.175.968.5
Ours
TLive‑Omni4B70.979.972.571.887.077.773.947.6
TLive‑Omni9B73.481.973.375.588.976.675.150.0

Notes: MMBench results are reported on the EN-DEV-v1.1 split. VAB and RWQA denote VLMsAreBlind and RealWorldQA.

General Benchmark: Hallucination, OCR, Grounding & Spatial Reasoning

ModelParamsHallusionAI2DOCRBenchCC‑OCRCharXivRefCOCOERQAEmbSpatial
Closed‑source models
Gemini 2.5 Flash59.187.786.474.860.1
Gemini 2.5 Pro60.990.087.276.862.950.373.3
Gemini 3 Pro68.694.190.479.081.484.170.561.2
GPT‑4o82.684.3
GPT‑5 (minimal)53.784.178.766.157.842.075.1
Qwen3.5‑Omni Flash89.089.180.864.492.650.082.7
Open‑source VLM models
MiMo‑VL‑SFT7B83.287.654.485.7
SAIL‑VL28B55.187.791.374.0
Valley2.58B56.384.487.0
LLaVA‑OneVision‑28B84.378.243.378.1
InternVL3.54B44.882.682.239.689.438.5
InternVL3.58B54.584.084.044.489.741.073.2
Qwen3‑VL4B57.684.188.176.239.789.041.379.6
Qwen3‑VL8B61.185.789.679.946.489.145.878.5
Qwen3.54B76.987.185.971.162.987.646.876.6
Qwen3.59B76.088.088.573.467.590.047.378.7
Open‑source Omni models
InteractiveOmni4B52.283.880.0
InteractiveOmni8B61.384.383.7
VITA‑1.57B44.979.373.2
Valley38B55.9
Nemotron 3 Nano Omni30B‑A3B88.588.349.180.6
Ming‑Lite‑Omni v1.520B‑A3B54.684.988.987.8
MiniCPM‑o 2.68B51.985.889.7
MiniCPM‑o 4.59B63.287.687.6
Qwen2.5‑Omni7B83.287.7
Qwen3‑Omni30B‑A3B59.785.286.061.1
Ours
TLive‑Omni4B77.786.686.680.561.387.442.379.3
TLive‑Omni9B76.088.690.381.363.190.048.080.4

Notes: CharXiv results are on the RQ split.

General Benchmark: Video Understanding

ModelParamsMVBenchMLVUV‑MMELongVBLVBenchMMVUV‑MMMU
Closed‑source models
Gemini 2.5 Flash77.875.662.268.265.2
Gemini 2.5 Pro65.881.280.669.072.279.4
Gemini 3 Pro74.183.087.776.776.277.587.6
GPT‑4o71.9
GPT‑5 (minimal)64.678.377.368.161.6
Qwen3.5‑Omni Flash70.881.977.062.7
Open‑source VLM models
MiMo‑VL‑SFT7B66.953.1
SAIL‑VL28B62.758.3
LLaVA‑OneVision‑28B66.276.671.966.955.556.2
LLaVA‑Video7B58.670.863.358.244.247.136.1
InternVL3.54B71.270.465.460.843.247.657.6
InternVL3.58B72.170.266.062.146.760.2
MiniCPM‑V 4.58B75.167.963.950.458.957.1
LongVU7B66.965.460.6
LongVILA7B67.160.157.1
Mage‑VL4B65.168.764.061.341.8
Molmo24B75.163.069.668.053.951.250.7
Molmo28B75.960.269.967.552.8
NVILA8B68.170.164.257.7
Kangaroo8B61.161.056.054.839.4
Video‑XL28B74.866.661.048.450.039.9
VideoChat34B70.156.756.457.4
VideoLLaMA 37B69.773.066.259.845.344.134.6
Qwen3‑VL4B68.975.369.356.250.556.2
Qwen3‑VL8B68.778.171.458.058.765.3
Qwen3.54B66.675.171.665.155.357.869.8
Qwen3.59B75.779.766.967.960.963.770.3
Open‑source Omni models
InteractiveOmni4B68.063.357.0
InteractiveOmni8B71.666.059.1
VITA‑1.57B55.456.1
Valley38B55.661.2
OmniVinci9B70.668.261.3
Nemotron 3 Nano Omni30B‑A3B70.8
Ming‑Lite‑Omni v1.520B‑A3B69.467.159.5
MiniCPM‑o 2.68B63.9
MiniCPM‑o 4.59B76.570.466.0
Qwen2.5‑Omni7B70.364.3
Qwen3‑Omni30B‑A3B75.270.5
Ours
TLive‑Omni4B69.076.171.366.157.159.973.9
TLive‑Omni9B72.580.975.669.960.867.172.8

Notes: LongVB denotes LongVideoBench. V-MME and V-MMMU denote Video-MME and VideoMMMU.

General Benchmark: Video Temporal Grounding (TimeLens-Bench)

ModelParamsCharades‑TLActivityNet‑TLQVHighlights‑TL
Closed‑source models
Gemini 2.5 Flash48.652.564.3
Gemini 2.5 Pro52.858.170.4
GPT‑4o41.840.452.1
GPT‑5 (minimal)40.542.956.8
Open‑source VLM models
MiMo‑VL‑SFT7B39.635.541.5
LLaVA‑OneVision‑28B53.553.866.4
LLaVA‑Video7B15.214.610.4
InternVL3.54B16.014.917.7
InternVL3.58B27.831.331.3
MiniCPM‑V 4.58B31.932.346.1
Mage‑VL4B50.745.457.4
Molmo24B33.339.858.7
Video‑XL‑28B38.930.046.2
VideoChat34B56.154.667.0
VideoLLaMA 37B39.829.836.9
Qwen3‑VL4B46.448.258.7
Qwen3‑VL8B48.346.859.4
Qwen3.54B48.751.655.0
Qwen3.59B52.054.057.2
Ours
TLive‑Omni4B57.058.269.2
TLive‑Omni9B56.355.464.1

General Benchmark: Omni-Modal Perception & Reasoning

ModelParamsAVUTWorldSenseV‑HolmesDailyOmniOmniVBFutureOmni
Closed‑source models
Gemini 2.5 Flash65.450.955.6
Gemini 3.1 Pro85.665.582.7
Qwen3.5‑Omni Flash81.457.981.8
Open‑source Omni models
video‑SALMONN 2+3B66.248.342.267.7
video‑SALMONN 2+7B69.550.946.971.8
OmniVinci9B48.266.536.752.8
Nemotron 3 Nano Omni30B‑A3B55.274.5
MiniCPM‑o 4.59B78.655.764.380.241.156.1
Qwen2.5‑Omni7B45.462.436.548.9
Qwen3‑Omni30B‑A3B74.254.050.471.943.853.4
Ours
TLive‑Omni4B78.654.057.578.641.657.2
TLive‑Omni9B80.056.059.380.543.258.5

Notes: OmniVB denotes OmniVideoBench. V-Holmes denotes VideoHolmes.

🎯 Qualitative Examples

Live-Commerce Capabilities

Live-Commerce Qualitative Examples
Six representative live-commerce cases:
  1. Live-Commerce Video QA — Links audio cues about a hidden 4-cm height boost with visual evidence.
  2. Temporal Grounding — Localizes repeated appearances of a queried product badge.
  3. Product Visual Grounding — Predicted vs. ground-truth bounding boxes for a queried product.
  4. OCR — Text extraction with commerce-oriented semantic labels.
  5. Dense Video Captioning — Temporally segmented descriptions of a clothing demonstration.
  6. Multi-Dimensional Shot Tagging — Structured labels for shot size, camera, layout and content category.

General-Capability Examples

General-Capability Qualitative Examples
Six general-capability cases:
  1. Dense Video Captioning — Time-aware segment descriptions with visual frames.
  2. Temporal Grounding — Localizes a specific action moment with matched intervals.
  3. Visual Grounding — Bounding box prediction for queried regions.
  4. OCR — Text extraction from nutrition labels at line-level precision.
  5. Omni-Modal QA — Cross-modal reasoning combining audio cues with visual evidence.
  6. Multi-Dimensional Shot Tagging — Structured labels for shot size, camera, layout and content category.

📦 Model Zoo

ModelStageAvailability
TLive-Omni-4BSFT + Faithful-RFTModel Weights
TLive-Omni-9BSFT + Faithful-RFTModel Weights

⚙️ Installation

This release targets Python 3.10 on Linux x86_64 with CUDA 12.8 and PyTorch 2.10.0.

conda create -n tlive python=3.10 -y
conda activate tlive
pip install -r environments/requirements.txt

environments/requirements.txt includes custom wheels for the supported environment and model. If any wheel does not match your hardware, CUDA version, or Python version, replace it with a compatible build for your setup.

🚀 Quick Start

# Text
python examples/inference_five_modes.py \
  --model /path/to/model --mode text

# Image
python examples/inference_five_modes.py \
  --model /path/to/model --mode image --image data/image.jpg

# Standalone audio
python examples/inference_five_modes.py \
  --model /path/to/model --mode audio --audio data/audio.mp3

# Video frames + audio track
python examples/inference_five_modes.py \
  --model /path/to/model --mode vocal-video --video data/vocal_video.mp4

# Video frames only
python examples/inference_five_modes.py \
  --model /path/to/model --mode silence-video --video data/silence_video.mp4

For temporal localization outputs, we recommend the MM:SS - MM:SS interval format, for example 01:23 - 01:35.

For direct processor calls, set use_audio_in_video=True to use the video's audio track, or False for visual-only video:

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    use_audio_in_video=True,
)

⚡ vLLM

Installation

First install the pre-built wheel (Python 3.10 + CUDA 12.8 + Linux x86_64), built and tested on NVIDIA H20 GPUs (Hopper, sm_90):

pip install https://github.com/TaoLiveAIGC/TLive-Omni/releases/download/v1.0.0-rc1/vllm-0.19.0+cu128-cp310-cp310-linux_x86_64.whl

If your GPU, driver, or CUDA setup is not compatible with this wheel, build vLLM from source using the customized code in vllm/.

Inference

# Text
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode text

# Image
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode image --image data/image.jpg

# Standalone audio
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode audio --audio data/audio.mp3

# Video frames + audio track
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode vocal-video --video data/vocal_video.mp4

# Video frames only
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode silence-video --video data/silence_video.mp4

🙏 Acknowledgement

TLive-Omni is built with reference to the following open-source projects: Qwen3.5, Qwen3-Omni, Transformers, ms-swift, DeepSpeed, and vLLM. We sincerely thank these projects and the Qwen team for their outstanding open-source models.

📖 Citation

If you find our work helpful, please consider citing our paper:

@article{hu2026tliveomni,
  title={TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming},
  author={Hu, Yibo and Qian, Yu and Gu, Mao and Tao, Yingfan and Chen, Yuhao and Luo, Yongdong and Liu, Zhuoqun and Jin, Meiguang and Ma, Junfeng},
  journal={arXiv preprint arXiv:2608.20958},
  year={2026}
}

📄 License

This project is released under the Apache License 2.0.

Contributors

Leon1207

3 commits

GMAmadeus

1 commits

TaoLiveAIGC/TLive-Omni

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

118

stars

4

commits

Python

primary language

Aug 24, 2026

updated

audio-understanding
large-language-models
multimodal-large-language-models
omni
omni-modal
omni-modal-video-understanding
video-understanding
Browse cluster: Vision-Language Models & Multimodal AI

README

🌟 TLive-Omni 🌟: An Omni-Modal Understanding Model for E-Commerce Live Streaming

logo

Technical Report Hugging Face 4B Model Hugging Face 9B Model License

📰 News

  • 📄 [2026-08-24] Technical ReportTLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming is now available on arXiv!
  • 🔥 [2026-08-20] Model Release — TLive-Omni is now available on Hugging Face! With TLive-Omni-4B and TLive-Omni-9B.

📋 Overview

TLive-Omni is an omni-modal understanding model for e-commerce live-stream, mapping image, video, audio, and text into a unified text-output interface. Built on a Qwen3.5 backbone with a grafted AuT audio encoder, it supports 256K tokens of context, trained via a three-stage SFT recipe followed by Faithful-RFT reinforcement fine-tuning.

✨ Highlights

  • Timestamped Per-vGrid layout — Audio and video tokens are organized into timestamped grid with explicit boundaries, keeping audio segments adjacent to their corresponding visual content for fine-grained temporal alignment over long streams.
  • Three-stage SFT recipe — Progressive training from audio-language alignment to full multimodal SFT, developing live-commerce understanding from omni-modal perception to instruction-following responses.
  • Faithful-RFT — A reinforcement fine-tuning stage for faithful and real-time live-stream demands, suppressing explicit reasoning traces and directly optimizing answer quality for live-commerce tasks.
  • Rich atomic capabilities — A scenario-oriented taxonomy covering speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc, supported by a compact data production engine.
  • Strong live-commerce performance with competitive generalization — 4B and 9B variants demonstrate strong results across live-commerce audio, image, and video tasks, together with excellent generalization on general benchmarks.

🏗️ Architecture

architecture

TLive-Omni is built on a Qwen3.5 backbone and extends it with a audio encoder through a lightweight MLP aligner, forming a unified text-output omni-modal understanding model. For video inputs with audio, each temporal grid is organized into a timestamped grid that interleaves video and audio token blocks, keeping audio segments adjacent to their corresponding visual content. The model supports up to 256K tokens of context at inference.

📊 Benchmark Results

We evaluate TLive-Omni-4B and TLive-Omni-9B on both live-commerce tasks and general benchmarks. Dash (-) denotes an unreported result or undisclosed parameter count. The Best results among open-source models are marked in bold, while the second-best results are in underlined.

Live-Commerce Audio Evaluation

ModelParamsLive ASR
CER ↓
Spk. ASR
cpWER ↓
Audio DescriptionAudio QA
Acc. ↑
Acc. ↑Hal. ↓
Closed‑source Omni models
Gemini 2.5 Flash16.3017.1465.2126.1976.28
Gemini 2.5 Pro11.4812.1781.1014.1682.85
Gemini 3 Flash15.1819.0468.2726.1774.68
Gemini 3 Pro12.0911.6785.0710.9288.62
Gemini 3.5 Flash13.0911.9979.9714.3687.99
Qwen3.5‑Omni Flash6.8113.2362.8227.8178.04
Open‑source Audio models
MiMo‑Audio7B12.7164.2626.0170.97
Fun‑Audio‑Chat8B14.5561.3532.2169.71
Step‑Audio‑R1.132B10.2175.0820.5069.80
Open‑source Omni models
OmniVinci9B39.9047.3666.51
Nemotron 3 Nano Omni30B‑A3B12.1017.6533.0139.7764.90
Ming‑Lite‑Omni v1.520B‑A3B10.0645.9944.0540.54
MiniCPM‑o 2.68B13.8849.8441.7639.74
MiniCPM‑o 4.59B10.7018.8947.5952.4142.47
Qwen2.5‑Omni7B7.8647.9236.9261.38
Qwen3‑Omni30B‑A3B6.7527.8461.0630.2276.76
Ours
TLive‑Omni4B6.6612.8876.1220.9772.60
TLive‑Omni9B6.4612.2775.9621.0076.28

Notes: Live ASR and Spk. ASR denote live-commerce ASR and speaker-attributed ASR.

Live-Commerce Image Evaluation

ModelParamsVisual GroundingText Understanding
Live AP ↑Prod AP ↑Loc. F1 ↑Rec. NED ↓Cls. Acc. ↑
Closed‑source Omni models
Gemini 2.5 Flash61.0828.8120.5243.2851.21
Gemini 2.5 Pro51.9832.6331.6027.8261.86
Gemini 3 Flash80.3865.6761.1116.2569.11
Gemini 3 Pro73.8058.8368.609.7276.86
Gemini 3.5 Flash84.1574.8964.4416.6469.76
Qwen3.5‑Omni Flash79.9660.4474.0712.4853.25
Open‑source Omni models
OmniVinci9B34.868.9350.2532.7757.29
Nemotron 3 Nano Omni30B‑A3B73.0848.6252.9129.4237.86
Ming‑Lite‑Omni v1.520B‑A3B52.4640.7313.2759.1632.94
MiniCPM‑o 2.68B3.821.775.7477.5815.92
MiniCPM‑o 4.59B23.9053.635.4371.6511.62
Qwen2.5‑Omni7B75.6122.8542.6437.7951.25
Qwen3‑Omni30B‑A3B79.2268.8830.4614.8369.46
Ours
TLive‑Omni4B82.8591.4586.994.7279.06
TLive‑Omni9B82.3389.9687.594.2479.85

Live-Commerce Video Evaluation

ModelParamsTG
mIoU ↑
Dense CaptionVideo QA
Acc. ↑
Shot Understanding
Acc. ↑Hal. ↓Layout ↑Shot Size ↑Camera ↑Content ↑
Closed‑source Omni models
Gemini 2.5 Flash76.5054.6010.9788.2180.0046.8084.2068.60
Gemini 2.5 Pro76.2241.9516.8892.6285.2050.8076.0070.40
Gemini 3 Flash77.4332.2120.7689.6476.8045.7078.5071.60
Gemini 3 Pro77.9037.8020.9984.3680.4043.4075.7074.80
Gemini 3.5 Flash77.9033.8017.3086.9083.4044.2075.5070.20
Qwen3.5‑Omni Flash62.1032.9420.9187.2884.4048.9085.5066.40
Open‑source Omni models
OmniVinci9B13.1018.5927.1372.5173.6052.7068.1049.20
Nemotron 3 Nano Omni30B‑A3B23.3917.9616.6282.5679.2034.0080.2058.60
Ming‑Lite‑Omni v1.520B‑A3B14.3413.8139.3364.5174.6041.7072.8051.60
MiniCPM‑o 2.68B14.5610.5326.9360.3066.6038.3068.3049.00
MiniCPM‑o 4.59B43.2021.0628.6184.6278.2042.8079.2066.60
Qwen2.5‑Omni7B30.8316.5136.4475.4874.4038.1081.2068.00
Qwen3‑Omni30B‑A3B39.2221.4425.8281.6282.2037.4076.1063.60
Ours
TLive‑Omni4B77.6369.239.5792.3178.4051.2080.9069.80
TLive‑Omni9B81.4974.638.7693.2377.0051.0082.0071.00

General Benchmark: Image Reasoning & QA

ModelParamsMMMUMathVistaDynaMathVABMMBenchRWQAMMStarSimpleVQA
Closed‑source models
Gemini 2.5 Flash76.375.369.775.986.675.775.859.2
Gemini 2.5 Pro80.977.778.578.588.476.078.566.9
Gemini 3 Pro87.287.985.193.783.383.173.2
GPT‑4o70.763.854.486.0
GPT‑5 (minimal)74.450.974.053.481.377.365.256.7
Qwen3.5‑Omni Flash76.982.979.388.877.575.754.4
Open‑source VLM models
MiMo‑VL‑SFT7B64.681.846.978.084.5
SAIL‑VL28B55.476.417.876.370.7
Valley2.58B62.174.432.785.570.567.3
LLaVA‑OneVision‑28B85.769.764.8
InternVL3.54B66.677.135.780.366.365.0
InternVL3.58B73.478.437.779.567.569.3
Qwen3‑VL4B67.473.765.371.983.970.969.848.0
Qwen3‑VL8B69.677.267.774.084.571.570.950.2
Qwen3.54B72.181.069.662.386.372.574.844.6
Qwen3.59B74.282.274.671.887.772.976.348.9
Open‑source Omni models
InteractiveOmni4B61.161.778.962.6
InteractiveOmni8B66.968.081.466.8
VITA‑1.57B52.166.276.759.9
Valley38B69.3
OmniVinci9B49.763.567.5
Nemotron 3 Nano Omni30B‑A3B55.271.9
Ming‑Lite‑Omni v1.520B‑A3B54.372.065.1
MiniCPM‑o 2.68B50.471.980.564.0
MiniCPM‑o 4.59B67.687.673.1
Qwen2.5‑Omni7B59.267.981.870.364.0
Qwen3‑Omni30B‑A3B69.175.968.5
Ours
TLive‑Omni4B70.979.972.571.887.077.773.947.6
TLive‑Omni9B73.481.973.375.588.976.675.150.0

Notes: MMBench results are reported on the EN-DEV-v1.1 split. VAB and RWQA denote VLMsAreBlind and RealWorldQA.

General Benchmark: Hallucination, OCR, Grounding & Spatial Reasoning

ModelParamsHallusionAI2DOCRBenchCC‑OCRCharXivRefCOCOERQAEmbSpatial
Closed‑source models
Gemini 2.5 Flash59.187.786.474.860.1
Gemini 2.5 Pro60.990.087.276.862.950.373.3
Gemini 3 Pro68.694.190.479.081.484.170.561.2
GPT‑4o82.684.3
GPT‑5 (minimal)53.784.178.766.157.842.075.1
Qwen3.5‑Omni Flash89.089.180.864.492.650.082.7
Open‑source VLM models
MiMo‑VL‑SFT7B83.287.654.485.7
SAIL‑VL28B55.187.791.374.0
Valley2.58B56.384.487.0
LLaVA‑OneVision‑28B84.378.243.378.1
InternVL3.54B44.882.682.239.689.438.5
InternVL3.58B54.584.084.044.489.741.073.2
Qwen3‑VL4B57.684.188.176.239.789.041.379.6
Qwen3‑VL8B61.185.789.679.946.489.145.878.5
Qwen3.54B76.987.185.971.162.987.646.876.6
Qwen3.59B76.088.088.573.467.590.047.378.7
Open‑source Omni models
InteractiveOmni4B52.283.880.0
InteractiveOmni8B61.384.383.7
VITA‑1.57B44.979.373.2
Valley38B55.9
Nemotron 3 Nano Omni30B‑A3B88.588.349.180.6
Ming‑Lite‑Omni v1.520B‑A3B54.684.988.987.8
MiniCPM‑o 2.68B51.985.889.7
MiniCPM‑o 4.59B63.287.687.6
Qwen2.5‑Omni7B83.287.7
Qwen3‑Omni30B‑A3B59.785.286.061.1
Ours
TLive‑Omni4B77.786.686.680.561.387.442.379.3
TLive‑Omni9B76.088.690.381.363.190.048.080.4

Notes: CharXiv results are on the RQ split.

General Benchmark: Video Understanding

ModelParamsMVBenchMLVUV‑MMELongVBLVBenchMMVUV‑MMMU
Closed‑source models
Gemini 2.5 Flash77.875.662.268.265.2
Gemini 2.5 Pro65.881.280.669.072.279.4
Gemini 3 Pro74.183.087.776.776.277.587.6
GPT‑4o71.9
GPT‑5 (minimal)64.678.377.368.161.6
Qwen3.5‑Omni Flash70.881.977.062.7
Open‑source VLM models
MiMo‑VL‑SFT7B66.953.1
SAIL‑VL28B62.758.3
LLaVA‑OneVision‑28B66.276.671.966.955.556.2
LLaVA‑Video7B58.670.863.358.244.247.136.1
InternVL3.54B71.270.465.460.843.247.657.6
InternVL3.58B72.170.266.062.146.760.2
MiniCPM‑V 4.58B75.167.963.950.458.957.1
LongVU7B66.965.460.6
LongVILA7B67.160.157.1
Mage‑VL4B65.168.764.061.341.8
Molmo24B75.163.069.668.053.951.250.7
Molmo28B75.960.269.967.552.8
NVILA8B68.170.164.257.7
Kangaroo8B61.161.056.054.839.4
Video‑XL28B74.866.661.048.450.039.9
VideoChat34B70.156.756.457.4
VideoLLaMA 37B69.773.066.259.845.344.134.6
Qwen3‑VL4B68.975.369.356.250.556.2
Qwen3‑VL8B68.778.171.458.058.765.3
Qwen3.54B66.675.171.665.155.357.869.8
Qwen3.59B75.779.766.967.960.963.770.3
Open‑source Omni models
InteractiveOmni4B68.063.357.0
InteractiveOmni8B71.666.059.1
VITA‑1.57B55.456.1
Valley38B55.661.2
OmniVinci9B70.668.261.3
Nemotron 3 Nano Omni30B‑A3B70.8
Ming‑Lite‑Omni v1.520B‑A3B69.467.159.5
MiniCPM‑o 2.68B63.9
MiniCPM‑o 4.59B76.570.466.0
Qwen2.5‑Omni7B70.364.3
Qwen3‑Omni30B‑A3B75.270.5
Ours
TLive‑Omni4B69.076.171.366.157.159.973.9
TLive‑Omni9B72.580.975.669.960.867.172.8

Notes: LongVB denotes LongVideoBench. V-MME and V-MMMU denote Video-MME and VideoMMMU.

General Benchmark: Video Temporal Grounding (TimeLens-Bench)

ModelParamsCharades‑TLActivityNet‑TLQVHighlights‑TL
Closed‑source models
Gemini 2.5 Flash48.652.564.3
Gemini 2.5 Pro52.858.170.4
GPT‑4o41.840.452.1
GPT‑5 (minimal)40.542.956.8
Open‑source VLM models
MiMo‑VL‑SFT7B39.635.541.5
LLaVA‑OneVision‑28B53.553.866.4
LLaVA‑Video7B15.214.610.4
InternVL3.54B16.014.917.7
InternVL3.58B27.831.331.3
MiniCPM‑V 4.58B31.932.346.1
Mage‑VL4B50.745.457.4
Molmo24B33.339.858.7
Video‑XL‑28B38.930.046.2
VideoChat34B56.154.667.0
VideoLLaMA 37B39.829.836.9
Qwen3‑VL4B46.448.258.7
Qwen3‑VL8B48.346.859.4
Qwen3.54B48.751.655.0
Qwen3.59B52.054.057.2
Ours
TLive‑Omni4B57.058.269.2
TLive‑Omni9B56.355.464.1

General Benchmark: Omni-Modal Perception & Reasoning

ModelParamsAVUTWorldSenseV‑HolmesDailyOmniOmniVBFutureOmni
Closed‑source models
Gemini 2.5 Flash65.450.955.6
Gemini 3.1 Pro85.665.582.7
Qwen3.5‑Omni Flash81.457.981.8
Open‑source Omni models
video‑SALMONN 2+3B66.248.342.267.7
video‑SALMONN 2+7B69.550.946.971.8
OmniVinci9B48.266.536.752.8
Nemotron 3 Nano Omni30B‑A3B55.274.5
MiniCPM‑o 4.59B78.655.764.380.241.156.1
Qwen2.5‑Omni7B45.462.436.548.9
Qwen3‑Omni30B‑A3B74.254.050.471.943.853.4
Ours
TLive‑Omni4B78.654.057.578.641.657.2
TLive‑Omni9B80.056.059.380.543.258.5

Notes: OmniVB denotes OmniVideoBench. V-Holmes denotes VideoHolmes.

🎯 Qualitative Examples

Live-Commerce Capabilities

Live-Commerce Qualitative Examples
Six representative live-commerce cases:
  1. Live-Commerce Video QA — Links audio cues about a hidden 4-cm height boost with visual evidence.
  2. Temporal Grounding — Localizes repeated appearances of a queried product badge.
  3. Product Visual Grounding — Predicted vs. ground-truth bounding boxes for a queried product.
  4. OCR — Text extraction with commerce-oriented semantic labels.
  5. Dense Video Captioning — Temporally segmented descriptions of a clothing demonstration.
  6. Multi-Dimensional Shot Tagging — Structured labels for shot size, camera, layout and content category.

General-Capability Examples

General-Capability Qualitative Examples
Six general-capability cases:
  1. Dense Video Captioning — Time-aware segment descriptions with visual frames.
  2. Temporal Grounding — Localizes a specific action moment with matched intervals.
  3. Visual Grounding — Bounding box prediction for queried regions.
  4. OCR — Text extraction from nutrition labels at line-level precision.
  5. Omni-Modal QA — Cross-modal reasoning combining audio cues with visual evidence.
  6. Multi-Dimensional Shot Tagging — Structured labels for shot size, camera, layout and content category.

📦 Model Zoo

ModelStageAvailability
TLive-Omni-4BSFT + Faithful-RFTModel Weights
TLive-Omni-9BSFT + Faithful-RFTModel Weights

⚙️ Installation

This release targets Python 3.10 on Linux x86_64 with CUDA 12.8 and PyTorch 2.10.0.

conda create -n tlive python=3.10 -y
conda activate tlive
pip install -r environments/requirements.txt

environments/requirements.txt includes custom wheels for the supported environment and model. If any wheel does not match your hardware, CUDA version, or Python version, replace it with a compatible build for your setup.

🚀 Quick Start

# Text
python examples/inference_five_modes.py \
  --model /path/to/model --mode text

# Image
python examples/inference_five_modes.py \
  --model /path/to/model --mode image --image data/image.jpg

# Standalone audio
python examples/inference_five_modes.py \
  --model /path/to/model --mode audio --audio data/audio.mp3

# Video frames + audio track
python examples/inference_five_modes.py \
  --model /path/to/model --mode vocal-video --video data/vocal_video.mp4

# Video frames only
python examples/inference_five_modes.py \
  --model /path/to/model --mode silence-video --video data/silence_video.mp4

For temporal localization outputs, we recommend the MM:SS - MM:SS interval format, for example 01:23 - 01:35.

For direct processor calls, set use_audio_in_video=True to use the video's audio track, or False for visual-only video:

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    use_audio_in_video=True,
)

⚡ vLLM

Installation

First install the pre-built wheel (Python 3.10 + CUDA 12.8 + Linux x86_64), built and tested on NVIDIA H20 GPUs (Hopper, sm_90):

pip install https://github.com/TaoLiveAIGC/TLive-Omni/releases/download/v1.0.0-rc1/vllm-0.19.0+cu128-cp310-cp310-linux_x86_64.whl

If your GPU, driver, or CUDA setup is not compatible with this wheel, build vLLM from source using the customized code in vllm/.

Inference

# Text
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode text

# Image
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode image --image data/image.jpg

# Standalone audio
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode audio --audio data/audio.mp3

# Video frames + audio track
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode vocal-video --video data/vocal_video.mp4

# Video frames only
python examples/inference_vllm_five_modes.py \
  --model /path/to/model --mode silence-video --video data/silence_video.mp4

🙏 Acknowledgement

TLive-Omni is built with reference to the following open-source projects: Qwen3.5, Qwen3-Omni, Transformers, ms-swift, DeepSpeed, and vLLM. We sincerely thank these projects and the Qwen team for their outstanding open-source models.

📖 Citation

If you find our work helpful, please consider citing our paper:

@article{hu2026tliveomni,
  title={TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming},
  author={Hu, Yibo and Qian, Yu and Gu, Mao and Tao, Yingfan and Chen, Yuhao and Luo, Yongdong and Liu, Zhuoqun and Jin, Meiguang and Ma, Junfeng},
  journal={arXiv preprint arXiv:2608.20958},
  year={2026}
}

📄 License

This project is released under the Apache License 2.0.

Contributors

Leon1207

3 commits

GMAmadeus

1 commits

Languages

Python

87.9%

Cuda

6.1%

C++

3.7%

Shell

1.1%