alanshaoTT/LAT-Audio-Repo

34

stars

3,066

commits

Python

primary language

Apr 28, 2026

updated

README

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

Mingchen Shao1, Hang Su2, Wenjie Tian1, Bingshen Mu1, Zhennan Lin1, Lichun Fan2, Zhenbo Luo2, Jian Luan2, Lei Xie1†,

1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University

πŸ“‘ Paper Β Β  | Β Β  πŸ™ GitHub Β Β  | Β Β  πŸ€— HuggingFace Β Β  | Β Β  πŸ’¬ Contact Us

πŸš€ Overview

Large Audio Language Models (LALMs) perform well on short audio but struggle with long-form audio due to temporal hallucination and timestamp drift.

We propose LAT-Audio, a global-to-local reasoning framework that enables precise temporal awareness through:

  • A global timeline for structured temporal understanding
  • A Think-With-Audio Chain-of-Thought (TWA-CoT) for iterative reasoning
  • A tool-augmented mechanism to retrieve local audio evidence

We further introduce:

  • LAT-Chronicle: a 1.2k-hour long-form audio dataset
  • LAT-Bench: the first human-verified long-form temporal benchmark

🧠 Framework

LAT-Audio follows a progressive global-to-local reasoning paradigm:

  1. Construct a global timeline as temporal-semantic anchors
  2. Perform multi-step reasoning via TWA-CoT
  3. Iteratively retrieve audio evidence through tool calls

πŸ“Š Dataset & Benchmark

ComponentDescription
LAT-ChronicleA 1.2k-hour long-form audio dataset (1k Chinese / 200h English) with fine-grained temporal annotations across six real-world scenarios.
LAT-BenchA human-verified benchmark for long-form temporal reasoning, covering three core tasks: Dense Audio Captioning (DAC), Temporal Audio Grounding (TAG), and Targeted Audio Captioning (TAC).

LAT-Chronicle is constructed via a human-in-the-loop pipeline with multi-track atomic annotations (speech, sound events, music, and environment), enabling precise temporal supervision for long-form audio understanding.

LAT-Bench supports audio up to 30 minutes, providing realistic evaluation settings. All annotations are carefully validated to ensure high temporal accuracy and consistency, covering diverse scenarios such as conversations, vlogs, and complex acoustic environments.

πŸ‘‰ Download:

πŸ€– Models

ModelReasoningTraining DataDescription
LAT-AudioYesLAT-ChronicleTool-augmented multi-step reasoning model with global-to-local temporal inference
LAT-Audio-BaseNoLAT-Chronicle + in-houseDirect modeling baseline fine-tuned from Qwen3-Omni with more in-house data, offering faster and simpler inference

πŸ‘‰ Download:

πŸ› οΈ Quick Start

Environment Setup

conda create -n lat-audio python=3.10 -y
conda activate lat-audio
pip install -r requirements.txt

Usage

Inference

LAT-Audio

cd examples/train/multimodal/
export QWEN3_OMNI_USE_SWIFT_FORWARD=true
export AUDIO_SEQ_COMP_FACTOR=2
python lat-audio-infer.py \
  --ckpt ./checkpoints/LAT-Audio \
  --task TAG_EN \
  --test_file ./data/TAG_test_EN.jsonl \
  --crop_dir ./tmp_crops \
  --output ./results/TAG_EN.jsonl

LAT-Audio-Base

cd examples/train/multimodal/
bash lat-audio-base-infer.sh

Evaluation Protocol

Dense audio caption

cd examples/train/multimodal/
python ./eval/eval-DAC.py \
  --input ./result/DAC/EN.jsonl \
  --output ./reports/dac_en_metrics.json \
  --pred-field response \
  --ref-field labels \
  --thresholds 0.3,0.5,0.7 \
  --device cuda \
  --sbert-model ./paraphrase-multilingual-MiniLM-L12-v2

Target audio caption

cd examples/train/multimodal/
python ./eval/eval-TAC.py \
  --input ./result/TAC/CN.jsonl \
  --output ./reports/tac_metrics.json \
  --pred-field response \
  --ref-field labels \
  --device cuda \
  --sbert-model ./paraphrase-multilingual-MiniLM-L12-v2

Temporal audio grounding

cd examples/train/multimodal/
python ./eval/eval-TAG.py \
  --input ./result/TAG/test.jsonl \
  --output ./reports/tag_metrics.json \
  --pred-field response \
  --ref-field labels \
  --thresholds 0.3,0.5,0.7

πŸ“œ Citation

@article{shao2026lataudio,
  title={Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding},
  author={Shao, Mingchen and Su, Hang and Tian, Wenjie and Mu, Bingshen and Lin, Zhennan and Fan, Lichun and Luo, Zhenbo and Luan, Jian and Xie, Lei},
  journal={arXiv preprint arXiv:2604.22245},
  year={2026}
}

Contact

If you are interested in leaving a message to our research team, feel free to email mcshao@mail.nwpu.edu.cn .

Contributors

(top 30 of 126)

Jintao-Huang

1,683 commits

tastelikefeet

572 commits

hjh0119

411 commits

alanshaoTT

58 commits

alanshaoTT/LAT-Audio-Repo

34

stars

3,066

commits

Python

primary language

Apr 28, 2026

updated

README

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

Mingchen Shao1, Hang Su2, Wenjie Tian1, Bingshen Mu1, Zhennan Lin1, Lichun Fan2, Zhenbo Luo2, Jian Luan2, Lei Xie1†,

1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University

πŸ“‘ Paper Β Β  | Β Β  πŸ™ GitHub Β Β  | Β Β  πŸ€— HuggingFace Β Β  | Β Β  πŸ’¬ Contact Us

πŸš€ Overview

Large Audio Language Models (LALMs) perform well on short audio but struggle with long-form audio due to temporal hallucination and timestamp drift.

We propose LAT-Audio, a global-to-local reasoning framework that enables precise temporal awareness through:

  • A global timeline for structured temporal understanding
  • A Think-With-Audio Chain-of-Thought (TWA-CoT) for iterative reasoning
  • A tool-augmented mechanism to retrieve local audio evidence

We further introduce:

  • LAT-Chronicle: a 1.2k-hour long-form audio dataset
  • LAT-Bench: the first human-verified long-form temporal benchmark

🧠 Framework

LAT-Audio follows a progressive global-to-local reasoning paradigm:

  1. Construct a global timeline as temporal-semantic anchors
  2. Perform multi-step reasoning via TWA-CoT
  3. Iteratively retrieve audio evidence through tool calls

πŸ“Š Dataset & Benchmark

ComponentDescription
LAT-ChronicleA 1.2k-hour long-form audio dataset (1k Chinese / 200h English) with fine-grained temporal annotations across six real-world scenarios.
LAT-BenchA human-verified benchmark for long-form temporal reasoning, covering three core tasks: Dense Audio Captioning (DAC), Temporal Audio Grounding (TAG), and Targeted Audio Captioning (TAC).

LAT-Chronicle is constructed via a human-in-the-loop pipeline with multi-track atomic annotations (speech, sound events, music, and environment), enabling precise temporal supervision for long-form audio understanding.

LAT-Bench supports audio up to 30 minutes, providing realistic evaluation settings. All annotations are carefully validated to ensure high temporal accuracy and consistency, covering diverse scenarios such as conversations, vlogs, and complex acoustic environments.

πŸ‘‰ Download:

πŸ€– Models

ModelReasoningTraining DataDescription
LAT-AudioYesLAT-ChronicleTool-augmented multi-step reasoning model with global-to-local temporal inference
LAT-Audio-BaseNoLAT-Chronicle + in-houseDirect modeling baseline fine-tuned from Qwen3-Omni with more in-house data, offering faster and simpler inference

πŸ‘‰ Download:

πŸ› οΈ Quick Start

Environment Setup

conda create -n lat-audio python=3.10 -y
conda activate lat-audio
pip install -r requirements.txt

Usage

Inference

LAT-Audio

cd examples/train/multimodal/
export QWEN3_OMNI_USE_SWIFT_FORWARD=true
export AUDIO_SEQ_COMP_FACTOR=2
python lat-audio-infer.py \
  --ckpt ./checkpoints/LAT-Audio \
  --task TAG_EN \
  --test_file ./data/TAG_test_EN.jsonl \
  --crop_dir ./tmp_crops \
  --output ./results/TAG_EN.jsonl

LAT-Audio-Base

cd examples/train/multimodal/
bash lat-audio-base-infer.sh

Evaluation Protocol

Dense audio caption

cd examples/train/multimodal/
python ./eval/eval-DAC.py \
  --input ./result/DAC/EN.jsonl \
  --output ./reports/dac_en_metrics.json \
  --pred-field response \
  --ref-field labels \
  --thresholds 0.3,0.5,0.7 \
  --device cuda \
  --sbert-model ./paraphrase-multilingual-MiniLM-L12-v2

Target audio caption

cd examples/train/multimodal/
python ./eval/eval-TAC.py \
  --input ./result/TAC/CN.jsonl \
  --output ./reports/tac_metrics.json \
  --pred-field response \
  --ref-field labels \
  --device cuda \
  --sbert-model ./paraphrase-multilingual-MiniLM-L12-v2

Temporal audio grounding

cd examples/train/multimodal/
python ./eval/eval-TAG.py \
  --input ./result/TAG/test.jsonl \
  --output ./reports/tag_metrics.json \
  --pred-field response \
  --ref-field labels \
  --thresholds 0.3,0.5,0.7

πŸ“œ Citation

@article{shao2026lataudio,
  title={Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding},
  author={Shao, Mingchen and Su, Hang and Tian, Wenjie and Mu, Bingshen and Lin, Zhennan and Fan, Lichun and Luo, Zhenbo and Luan, Jian and Xie, Lei},
  journal={arXiv preprint arXiv:2604.22245},
  year={2026}
}

Contact

If you are interested in leaving a message to our research team, feel free to email mcshao@mail.nwpu.edu.cn .

Contributors

(top 30 of 126)

Jintao-Huang

1,683 commits

tastelikefeet

572 commits

hjh0119

411 commits

alanshaoTT

58 commits

Languages

Python

99.8%