MERaLiON/MERaLiON-3-10B

Model

⚙️ vLLM coming soon

8

22 commits

1 linked in READMEs

updated Jun 17, 2026

See the code

README

🔥 MERaLiON-3 🔥

🚀 MERaLiON-3-10B

💻 Web Demo | ⚙️ vLLM coming soon

Introduction

We are pleased to announce the release of our flagship speech-text large language model, MERaLiON-3-10B. MERaLiON-3-10B demonstrates competitive performance across benchmark evaluations in Age Recognition, Gender Recognition, Spoken Question Answering (SQA), and Contextual Paralinguistic Question Answering (CPQA) in the Southeast Asian context as compared to the latest AudioLLMs, including Gemini 3 Flash and Qwen3 Omni Instruct. The benchmark contains speech and prompts in Malay, Indonesian, English, Chinese, Tamil, Thai and Vietnamese to better represent the Southeast Asian context. The following table presents task-specific evaluation scores, assessed using the LLM-as-a-Judge framework across multiple datasets. Higher scores indicate better performance. We will open-source the benchmark separately as part of a paper. See the Evaluation section for detailed benchmarking.

BenchmarkMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
Age (commonvoice-en, ta, th, vi, zh)76.8461.7770.3877.0068.90
Gender (Multi-dataset)92.7054.1995.3481.7240.25
Spoken Q&A (SQA)59.6156.7658.7459.7557.48
Contextual paralinguistic Q&A (CPQA)57.0248.3154.2154.0754.54

MERaLiON-3-10B also maintains its competitive performance in other tasks such as Multilingual Automatic Speech Recognition (ASR), Speech Translation (ST), Audio Scene Understanding and general speech comprehension vis-à-vis MERaLiON-2-10B.

Model Description:

MERaLiON stands for Multimodal Empathetic Reasoning and Learning in One Network, with models tailored for Singapore’s multilingual and multicultural landscape, as well as the wider Southeast Asian region.

MERaLiON-3-10B is finetuned on 150,000 hours of speech and audio data across 6 diverse tasks: Automatic Speech Recognition (ASR), SQA, Spoken Dialogue Summarization (SDS), Audio Captioning (AC), Audio-Scene Question Answering (ASQA) and CPQA.

  • Developed by: I2R, A*STAR, Singapore
  • Model type: Multimodal LLM
  • Language(s): Primarily English (Global and Singapore), Chinese, with support for audio of regional languages including Malay, Tamil, Indonesian, Thai, and Vietnamese.
  • Audio: Mono channel audio, 16000 hz, up to 300 seconds.
  • License: MERaLiON Public License
  • Demo: MERaLiON-AudioLLM Web Demo

Performance:

We benchmarked MERaLiON-3-10B against Qwen3 Omni, Gemini 3 Flash, GPT 4o Audio, and MERaLiON-2-10B, and it performed the best on 31 out of 59 benchmarks for tasks related to age recognition, gender recognition, SQA, and CPQA. MERaLiON-3-10B-preview maintains competitive performance vis-à-vis MERaLiON-2-10B on the Audiobench benchmarks.

Age recognition

Age recognition tasks categorise speakers as teens (10-19), adults (20-59), or seniors (60-100). The prompts are either in English, or in a Southeast Asian language. LLM-as-a-judge is used to evaluate the correctness of each response.

DatasetLangVarMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
Commonvoiceeneng64.8663.1064.2068.0065.00
sea64.8663.1064.2068.0065.00
taeng79.0064.6573.5079.0071.00
sea59.9047.9048.4078.0062.00
theng83.7257.8178.0677.0078.00
sea81.1642.1964.1384.0053.00
vieng92.3273.2384.3981.0086.00
sea90.4064.3577.6787.0081.00
zheng77.4572.4075.6075.0083.00
sea74.7069.0073.6073.0045.00
Average76.8461.7770.3877.0068.90

Gender recognition

The gender recognition benchmark consists of speech samples in Indonesian, Tamil, Thai, Vietnamese, Chinese, Malay, English, and Khmer. The text prompts are either in English, or in a Southeast Asian language. LLM-as-a-judge is used to evaluate the correctness of each response.

DatasetLangVarMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
commonvoiceideng97.1045.2096.8086.0046.00
sea96.9057.3096.1090.0053.93
taeng97.1053.0096.8065.0033.00
sea51.0040.4081.9071.0035.00
theng97.7250.0796.9287.0050.00
sea96.9223.9695.1882.0040.00
vieng98.6924.0598.8287.0026.00
sea98.5614.6496.8688.0035.00
zheng98.1053.7098.2089.0049.00
sea97.8035.5098.1082.0021.00
emotataeng100.0067.3199.8983.0025.00
sea63.6848.9397.6586.0033.00
fleurseneng99.6958.27100.0073.0078.00
sea99.6958.27100.0073.0078.00
kmeng100.0056.60100.0094.0062.00
sea97.3943.40100.0099.0015.00
indowavesentimentideng100.0071.67100.0084.0060.00
sea100.0060.67100.0088.0014.00
m3edzheng92.9084.3094.3073.0023.00
sea91.8070.7094.4072.0012.00
openslrtaeng100.0055.3099.0075.0047.00
sea67.5037.8087.9081.0036.00
sg streetseneng99.5989.63100.0087.0032.00
sea99.5989.63100.0087.0032.00
asr-smalduscmseng99.3052.4098.6097.0076.00
sea99.6044.0098.8099.0024.00
thai elderly speechtheng99.4068.1599.2977.0046.00
sea99.2926.9297.3976.0051.00
thai sertheng91.2063.4690.4785.0044.00
sea88.2761.7889.7476.0034.00
vietnam-celebvieng73.7065.8073.8062.0041.00
sea73.8061.4074.0061.0036.00
Average92.7054.1995.3481.7240.25

Spoken question and answer (SQA)

The benchmark consists of speech in English, Malay, Tamil, and Chinese, with text prompts in English containing questions related to the speech. As studies have found that LLM judges tend to favor longer, verbose answers even if they are not as clear, high-quality, or accurate as shorter alternatives, we have adjusted the judge's prompt to address verbosity bias.

DatasetMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
ytb_sqa_batch165.6065.8966.6663.2560.43
ytb_sqa_batch3_ms54.3550.4056.2557.7555.80
ytb_sqa_batch3_ta57.3453.6052.2559.4556.25
ytb_sqa_batch3_zh_en61.1557.1559.8058.5557.45
Average59.6156.7658.7459.7557.48

Contextual paralinguistic question and answer (CPQA)

The audio includes both speech and non-speech elements, and when no speech is present, LLMs are expected to reason solely based on acoustic or musical elements. The speech samples were in languages of Chinese, Malay, Tamil, English, a mix of any of the languages (codeswitch), or could include dialects such as Hokkien. To test for robustness in instruction following, the text prompts were designed to be diverse, and were written in any of the following languages: English, Malay, Tamil, Indonesian, Vietnamese, Chinese, or Thai. LLMs are expected to reply in the same language as the text prompt. Similar to SQA, we have adjusted the judge's prompt to address verbosity bias.

DatasetMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
yx_youtube_zh58.8850.1857.2754.6754.79
yx_youtube_codeswitch63.0447.3655.5659.4060.32
yx_youtube_dialect61.1247.7256.3655.3654.92
yx_youtube_ms62.0046.1653.8857.0056.36
yx_youtube_ta58.1238.8849.6056.6054.64
yx_youtube_en58.6451.6056.7653.5252.88
ytb_short_eval_cpqa_human151.6347.5753.9547.4249.97
ytb_short_eval_cpqa_llm157.1856.2556.0754.9452.44
ytb_long_eval_cpqa_llm159.0557.4857.4454.9456.32
ytb_long_eval_cpqa_human159.2251.3359.2156.3455.00
Emotional-YTB-MY_zh_30_test_CPQA_v151.2446.8151.2251.0753.41
Emotional-YTB-MY_ms_30_test_CPQA_v150.6344.8248.7949.1253.01
Emotional-YTB-MY_ta_test_CPQA_v150.5241.8848.6252.5654.96
Average57.0248.3154.2154.0754.54

Automatic Speech Recognition (ASR), instruction following and audio understanding

MERaLiON-3-10B continues to demonstrate competitive performance in ASR, instruction following and audio understanding as compared to MERaLiON-2-10B, with improvements on many metrics on Audiobench. Please visit AudioBench benchmark for dataset-level evaluation results.

BenchmarkMERaLiON-3-10BMERaLiON-2-10BMERaLiON-2-10B-ASRMERaLiON-2-3B
ASR (lower better)0.13250.14850.13320.1697
Speech Instruction75.6070.2013.4019.10
Audio Scene Question Answering58.3651.1449.5146.14
Spoken QA (Singlish)66.3866.5561.8559.70
Audio Captioning36.8635.6034.4733.24
Spoken Dialogue Summarisation53.7553.1055.8048.55
Spoken QA (English)82.0479.7473.9868.72
Music Understanding70.4363.9460.6655.60
Accent Recognition41.3941.8247.7960.05
Speech Translation27.7627.3928.5422.13

How to Use

[!WARNING] Out of Scope use: This model is not intended for use in tool calling, math, and coding tasks.

MERaLiON-3 requires transformers version 4.50.1

pip install transformers==4.50.1
pip install librosa

To run in GPU, MERaLiON-3 requires flash-attn.

pip install flash-attn --no-build-isolation

[!TIP] Should you face any difficulties installing the above packages, you can try installing within this Docker container instead: pytorch/pytorch:2.5.1-cuda12.1-cudnn9-devel, whose cuda and torch environments have been tested working.

Audio Input

  • For ASR tasks, the maximum audio length is suggested to be 30 seconds at 16,000 Hz.
  • For general speech & audio understanding tasks, the maximum audio length which we tested for was up to 300 seconds at 16,000 Hz sampling rate.

Text Prompt

MERaLiON-3 is trained with this prompt template:

Instruction: <TextHere> \nFollow the text instruction based on the following audio: <SpeechHere>

It is generally recommended to follow this template, i.e., replace <TextHere> with your text instruction while leaving the <SpeechHere> untouched. We list a few useful example prompts here:

Standard prompts for better accuracy

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"

transcription_prompt = prompt_template.format(query="Please transcribe this speech.")
translation_prompt = prompt_template.format(query="Please translate the speech into Malay")
summarization_prompt = prompt_template.format(query="Please summarize this speech")
audio_captioning_prompt_1 = prompt_template.format(query="Please describe the audio")
audio_captioning_prompt_2 = prompt_template.format(query="Please create a caption for the audio")
audio_scene_understanding_prompt = prompt_template.format(query="Are there people crying in the audio?")
speech_as_instruction_prompt = prompt_template.format(query="Please respond to the audio") # given a speech instruction is provided in the audio clip.
emotion_recognition_prompt_1 = prompt_template.format(query="What is the emotion of the speaker")
emotion_recognition_prompt_2 = prompt_template.format(query="Describe the paralinguistic features of the audio")
gender_recognition_prompt = prompt_template.format(query="What is the gender of the speaker")

More flexible prompts for enriched responses

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"

prompt_1 = prompt_template.format(query="describe the paralinguistics feature and return in json format.")
prompt_2 = prompt_template.format(query="Please summarize the content of the speech and analyse the paralinguistics features of this audio. Return in json format.")
prompt_3 = prompt_template.format(query="Please translate this speech to Singapore's 4 official languages.")

AI agent prompts (beyond the default prompt template)

prompt_1 = \
"""
You are MERaLiON-AudioLLM, an empathic AI assistant developed by A*STAR. MERaLiON stands for Multimodal Empathetic Reasoning and Learning in One Network.
You are a friendly and empathetic conversational partner, and is proficient in understanding human emotions, accents, and genders from paralinguistic features.
Maintain a tone that is warm, non-judgmental, and supportive while replying to user. 

User's voice:  <SpeechHere>
"""

Huggingface Inference with CPU

import librosa
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo_id = "MERaLiON/MERaLiON-3-10B"

processor = AutoProcessor.from_pretrained(
    repo_id, 
    trust_remote_code=True,
    )
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    repo_id,
    use_safetensors=True,
    trust_remote_code=True,
)

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"
transcribe_prompt = "Please transcribe this speech."
translate_prompt = "Can you please translate this speech into written Chinese?"

# batch inference of 2 samples
conversation = [
    [{"role": "user", "content": prompt_template.format(query=transcribe_prompt)}],
    [{"role": "user", "content": prompt_template.format(query=translate_prompt)}],
]

chat_prompt = processor.tokenizer.apply_chat_template(
    conversation=conversation,
    tokenize=False,
    add_generation_prompt=True
)

# Use audio at 16000hz.
audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
audio_array = [audio_array]*2
inputs = processor(text=chat_prompt, audios=audio_array)

# adjust the `max_new_tokens` based on your use case.
# Please note the inclusion of `no_repeat_ngram_size=6`.
outputs = model.generate(**inputs, max_new_tokens=256, no_repeat_ngram_size=6)
response = processor.batch_decode(outputs, skip_special_tokens=True)

Huggingface GPU Inference

import torch
import librosa
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo_id = "MERaLiON/MERaLiON-3-10B"
device = "cuda"

processor = AutoProcessor.from_pretrained(
    repo_id, 
    trust_remote_code=True,
    )
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    repo_id,
    use_safetensors=True,
    trust_remote_code=True,
    attn_implementation="flash_attention_2",
    torch_dtype=torch.bfloat16
).to(device)

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"
transcribe_prompt = "Please transcribe this speech."
translate_prompt = "Can you please translate this speech into written Chinese?"

# batch inference of 2 samples
conversation = [
    [{"role": "user", "content": prompt_template.format(query=transcribe_prompt)}],
    [{"role": "user", "content": prompt_template.format(query=translate_prompt)}],
]

chat_prompt = processor.tokenizer.apply_chat_template(
    conversation=conversation,
    tokenize=False,
    add_generation_prompt=True
)

# Use audio at 16000hz.
audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
audio_array = [audio_array]*2
inputs = processor(text=chat_prompt, audios=audio_array)

inputs = inputs.to(device, dtype=torch.bfloat16)

# adjust the `max_new_tokens` based on your use case.
# Please note the inclusion of `no_repeat_ngram_size=6`.
outputs = model.generate(**inputs, max_new_tokens=256, no_repeat_ngram_size=6)
response = processor.batch_decode(outputs, skip_special_tokens=True)

⚠️ Disclaimer

The current MERaLiON-3 has not been specifically aligned for safety and may generate content that is inappropriate, offensive, or harmful. Developers and users are responsible for performing their own safety fine-tuning and implementing necessary security measures. The authors shall not be held liable for any claims, damages, or other liabilities arising from the use of the released models, weights, or code.

Compute and Infrastructure

MERaLiON-3 was trained on the ASPIRE 2A+ Supercomputer Cluster, provided by National Supercomputing Centre (NSCC), Singapore. ASPIRE 2A+ cluster provides multiple H100 nodes, with each compute node equipped with 8 Nvidia H100 GPUs, 2 TB of RAM, and 30 TB of locally attached NVMe storage. These nodes are interconnected via a rail-optimised, full fat-tree topology, utilising 400 Gb/s NDR InfiniBand cables. Additionally, the cluster incorporates a 2.5 PB SSD-based Lustre file system, linked to the H100 nodes through high-speed InfiniBand connections.

With a global batch size of 768, we trained the current release of MERaLiON-3 for around 250k steps, which took around 2.5 days to complete using 16 nodes, 128 H100 GPUs.

📚 Citation

If you find our work useful, please cite our papers:

MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
AudioBench: A Universal Benchmark for Audio Large Language Models
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
Incorporating contextual paralinguistic understanding in large speech-language models
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish

@misc{he2024meralionaudiollmtechnicalreport,
      title={MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models}, 
      author={{MERaLiON Team}},
      year={2024},
      eprint={2412.09818},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2412.09818}, 
}
@article{wang2024audiobench,
    title={AudioBench: A Universal Benchmark for Audio Large Language Models},
    author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang, Wenyu and Liu, Zhengyuan and Aw, AiTi and Chen, Nancy F},
    journal={NAACL},
    year={2025}
    }
@inproceedings{wang2025benchmarking,
  title={Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data},
  author={Wang, Qiongqiong and Sailor, Hardik Bhupendra and Liu, Tianchi and Zhang, Wenyu and Huzaifah, Muhammad and Lertcheva, Nattadaporn and Sun, Shuo and Chen, Nancy F and Wu, Jinyang and Aw, AiTi},
  booktitle={Findings of EMNLP 2025},
  year={2025}
}
@inproceedings{cpqa_interspeech,
  title={Contextual Paralinguistic Data Creation for  Multi-Modal Speech-LLM: Data Condensation and Spoken {QA} Generation},
  author={Wang, Qiongqiong and Sailor, Hardik B and Liu, Tianchi and Aw, Ai Ti},
  booktitle={Proc. Interspeech},
  year={2025},
}
@inproceedings{cpqa_asru,
  title={Incorporating Contextual Paralinguistic
Understanding in Large Speech-Language Models},
  author={
      Wang, Qiongqiong and Sailor, Hardik B and Wong, Jeremy H. M. and Liu, Tianchi and Sun, Shuo and Zhang, Wenyu and Huzaifah, Muhammad and  Chen, Nancy and Aw, Ai Ti},
  booktitle={Proc. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)},
  year={2025},
}
@article{wang2025advancing,
    title={Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models},
    author={Wang, Bin and Zou, Xunlong and Sun, Shuo and Zhang, Wenyu and He, Yingxu and Liu, Zhuohan and Wei, Chengwei and Chen, Nancy F and Aw, AiTi},
    journal={arXiv preprint arXiv:2501.01034},
    year={2025}
    }
@article{zhang2024mowe,
    title={MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders},
    author={Zhang, Wenyu and Sun, Shuo and Wang, Bin and Zou, Xunlong and Liu, Zhuohan and He, Yingxu and Lin, Geyu and Chen, Nancy F and Aw, Ai Ti},
    journal={ICASSP},
    year={2025}
    }
@misc{huang2025meraliontextllmcrosslingualunderstandinglarge,
      title={MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish}, 
      author={Xin Huang and Tarun Kumar Vangani and Minh Duc Pham and Xunlong Zou and Bin Wang and Zhengyuan Liu and Ai Ti Aw},
      year={2025},
      eprint={2501.08335},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2501.08335}, 
}
automatic-speech-recognition
custom_code
meralion
meralion-3
meralion3
safetensors
transformers

Contributors

lewiswoncy

19 commits

sailorhb

2 commits

YingxuHe

1 commits

MERaLiON/MERaLiON-3-10B

Model

⚙️ vLLM coming soon

8

22 commits

1 linked in READMEs

updated Jun 17, 2026

See the code

README

🔥 MERaLiON-3 🔥

🚀 MERaLiON-3-10B

💻 Web Demo | ⚙️ vLLM coming soon

Introduction

We are pleased to announce the release of our flagship speech-text large language model, MERaLiON-3-10B. MERaLiON-3-10B demonstrates competitive performance across benchmark evaluations in Age Recognition, Gender Recognition, Spoken Question Answering (SQA), and Contextual Paralinguistic Question Answering (CPQA) in the Southeast Asian context as compared to the latest AudioLLMs, including Gemini 3 Flash and Qwen3 Omni Instruct. The benchmark contains speech and prompts in Malay, Indonesian, English, Chinese, Tamil, Thai and Vietnamese to better represent the Southeast Asian context. The following table presents task-specific evaluation scores, assessed using the LLM-as-a-Judge framework across multiple datasets. Higher scores indicate better performance. We will open-source the benchmark separately as part of a paper. See the Evaluation section for detailed benchmarking.

BenchmarkMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
Age (commonvoice-en, ta, th, vi, zh)76.8461.7770.3877.0068.90
Gender (Multi-dataset)92.7054.1995.3481.7240.25
Spoken Q&A (SQA)59.6156.7658.7459.7557.48
Contextual paralinguistic Q&A (CPQA)57.0248.3154.2154.0754.54

MERaLiON-3-10B also maintains its competitive performance in other tasks such as Multilingual Automatic Speech Recognition (ASR), Speech Translation (ST), Audio Scene Understanding and general speech comprehension vis-à-vis MERaLiON-2-10B.

Model Description:

MERaLiON stands for Multimodal Empathetic Reasoning and Learning in One Network, with models tailored for Singapore’s multilingual and multicultural landscape, as well as the wider Southeast Asian region.

MERaLiON-3-10B is finetuned on 150,000 hours of speech and audio data across 6 diverse tasks: Automatic Speech Recognition (ASR), SQA, Spoken Dialogue Summarization (SDS), Audio Captioning (AC), Audio-Scene Question Answering (ASQA) and CPQA.

  • Developed by: I2R, A*STAR, Singapore
  • Model type: Multimodal LLM
  • Language(s): Primarily English (Global and Singapore), Chinese, with support for audio of regional languages including Malay, Tamil, Indonesian, Thai, and Vietnamese.
  • Audio: Mono channel audio, 16000 hz, up to 300 seconds.
  • License: MERaLiON Public License
  • Demo: MERaLiON-AudioLLM Web Demo

Performance:

We benchmarked MERaLiON-3-10B against Qwen3 Omni, Gemini 3 Flash, GPT 4o Audio, and MERaLiON-2-10B, and it performed the best on 31 out of 59 benchmarks for tasks related to age recognition, gender recognition, SQA, and CPQA. MERaLiON-3-10B-preview maintains competitive performance vis-à-vis MERaLiON-2-10B on the Audiobench benchmarks.

Age recognition

Age recognition tasks categorise speakers as teens (10-19), adults (20-59), or seniors (60-100). The prompts are either in English, or in a Southeast Asian language. LLM-as-a-judge is used to evaluate the correctness of each response.

DatasetLangVarMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
Commonvoiceeneng64.8663.1064.2068.0065.00
sea64.8663.1064.2068.0065.00
taeng79.0064.6573.5079.0071.00
sea59.9047.9048.4078.0062.00
theng83.7257.8178.0677.0078.00
sea81.1642.1964.1384.0053.00
vieng92.3273.2384.3981.0086.00
sea90.4064.3577.6787.0081.00
zheng77.4572.4075.6075.0083.00
sea74.7069.0073.6073.0045.00
Average76.8461.7770.3877.0068.90

Gender recognition

The gender recognition benchmark consists of speech samples in Indonesian, Tamil, Thai, Vietnamese, Chinese, Malay, English, and Khmer. The text prompts are either in English, or in a Southeast Asian language. LLM-as-a-judge is used to evaluate the correctness of each response.

DatasetLangVarMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
commonvoiceideng97.1045.2096.8086.0046.00
sea96.9057.3096.1090.0053.93
taeng97.1053.0096.8065.0033.00
sea51.0040.4081.9071.0035.00
theng97.7250.0796.9287.0050.00
sea96.9223.9695.1882.0040.00
vieng98.6924.0598.8287.0026.00
sea98.5614.6496.8688.0035.00
zheng98.1053.7098.2089.0049.00
sea97.8035.5098.1082.0021.00
emotataeng100.0067.3199.8983.0025.00
sea63.6848.9397.6586.0033.00
fleurseneng99.6958.27100.0073.0078.00
sea99.6958.27100.0073.0078.00
kmeng100.0056.60100.0094.0062.00
sea97.3943.40100.0099.0015.00
indowavesentimentideng100.0071.67100.0084.0060.00
sea100.0060.67100.0088.0014.00
m3edzheng92.9084.3094.3073.0023.00
sea91.8070.7094.4072.0012.00
openslrtaeng100.0055.3099.0075.0047.00
sea67.5037.8087.9081.0036.00
sg streetseneng99.5989.63100.0087.0032.00
sea99.5989.63100.0087.0032.00
asr-smalduscmseng99.3052.4098.6097.0076.00
sea99.6044.0098.8099.0024.00
thai elderly speechtheng99.4068.1599.2977.0046.00
sea99.2926.9297.3976.0051.00
thai sertheng91.2063.4690.4785.0044.00
sea88.2761.7889.7476.0034.00
vietnam-celebvieng73.7065.8073.8062.0041.00
sea73.8061.4074.0061.0036.00
Average92.7054.1995.3481.7240.25

Spoken question and answer (SQA)

The benchmark consists of speech in English, Malay, Tamil, and Chinese, with text prompts in English containing questions related to the speech. As studies have found that LLM judges tend to favor longer, verbose answers even if they are not as clear, high-quality, or accurate as shorter alternatives, we have adjusted the judge's prompt to address verbosity bias.

DatasetMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
ytb_sqa_batch165.6065.8966.6663.2560.43
ytb_sqa_batch3_ms54.3550.4056.2557.7555.80
ytb_sqa_batch3_ta57.3453.6052.2559.4556.25
ytb_sqa_batch3_zh_en61.1557.1559.8058.5557.45
Average59.6156.7658.7459.7557.48

Contextual paralinguistic question and answer (CPQA)

The audio includes both speech and non-speech elements, and when no speech is present, LLMs are expected to reason solely based on acoustic or musical elements. The speech samples were in languages of Chinese, Malay, Tamil, English, a mix of any of the languages (codeswitch), or could include dialects such as Hokkien. To test for robustness in instruction following, the text prompts were designed to be diverse, and were written in any of the following languages: English, Malay, Tamil, Indonesian, Vietnamese, Chinese, or Thai. LLMs are expected to reply in the same language as the text prompt. Similar to SQA, we have adjusted the judge's prompt to address verbosity bias.

DatasetMERaLiON-3-10BMERaLiON-2-10BQwen3 OmniGemini 3 FlashGPT 4o Audio
yx_youtube_zh58.8850.1857.2754.6754.79
yx_youtube_codeswitch63.0447.3655.5659.4060.32
yx_youtube_dialect61.1247.7256.3655.3654.92
yx_youtube_ms62.0046.1653.8857.0056.36
yx_youtube_ta58.1238.8849.6056.6054.64
yx_youtube_en58.6451.6056.7653.5252.88
ytb_short_eval_cpqa_human151.6347.5753.9547.4249.97
ytb_short_eval_cpqa_llm157.1856.2556.0754.9452.44
ytb_long_eval_cpqa_llm159.0557.4857.4454.9456.32
ytb_long_eval_cpqa_human159.2251.3359.2156.3455.00
Emotional-YTB-MY_zh_30_test_CPQA_v151.2446.8151.2251.0753.41
Emotional-YTB-MY_ms_30_test_CPQA_v150.6344.8248.7949.1253.01
Emotional-YTB-MY_ta_test_CPQA_v150.5241.8848.6252.5654.96
Average57.0248.3154.2154.0754.54

Automatic Speech Recognition (ASR), instruction following and audio understanding

MERaLiON-3-10B continues to demonstrate competitive performance in ASR, instruction following and audio understanding as compared to MERaLiON-2-10B, with improvements on many metrics on Audiobench. Please visit AudioBench benchmark for dataset-level evaluation results.

BenchmarkMERaLiON-3-10BMERaLiON-2-10BMERaLiON-2-10B-ASRMERaLiON-2-3B
ASR (lower better)0.13250.14850.13320.1697
Speech Instruction75.6070.2013.4019.10
Audio Scene Question Answering58.3651.1449.5146.14
Spoken QA (Singlish)66.3866.5561.8559.70
Audio Captioning36.8635.6034.4733.24
Spoken Dialogue Summarisation53.7553.1055.8048.55
Spoken QA (English)82.0479.7473.9868.72
Music Understanding70.4363.9460.6655.60
Accent Recognition41.3941.8247.7960.05
Speech Translation27.7627.3928.5422.13

How to Use

[!WARNING] Out of Scope use: This model is not intended for use in tool calling, math, and coding tasks.

MERaLiON-3 requires transformers version 4.50.1

pip install transformers==4.50.1
pip install librosa

To run in GPU, MERaLiON-3 requires flash-attn.

pip install flash-attn --no-build-isolation

[!TIP] Should you face any difficulties installing the above packages, you can try installing within this Docker container instead: pytorch/pytorch:2.5.1-cuda12.1-cudnn9-devel, whose cuda and torch environments have been tested working.

Audio Input

  • For ASR tasks, the maximum audio length is suggested to be 30 seconds at 16,000 Hz.
  • For general speech & audio understanding tasks, the maximum audio length which we tested for was up to 300 seconds at 16,000 Hz sampling rate.

Text Prompt

MERaLiON-3 is trained with this prompt template:

Instruction: <TextHere> \nFollow the text instruction based on the following audio: <SpeechHere>

It is generally recommended to follow this template, i.e., replace <TextHere> with your text instruction while leaving the <SpeechHere> untouched. We list a few useful example prompts here:

Standard prompts for better accuracy

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"

transcription_prompt = prompt_template.format(query="Please transcribe this speech.")
translation_prompt = prompt_template.format(query="Please translate the speech into Malay")
summarization_prompt = prompt_template.format(query="Please summarize this speech")
audio_captioning_prompt_1 = prompt_template.format(query="Please describe the audio")
audio_captioning_prompt_2 = prompt_template.format(query="Please create a caption for the audio")
audio_scene_understanding_prompt = prompt_template.format(query="Are there people crying in the audio?")
speech_as_instruction_prompt = prompt_template.format(query="Please respond to the audio") # given a speech instruction is provided in the audio clip.
emotion_recognition_prompt_1 = prompt_template.format(query="What is the emotion of the speaker")
emotion_recognition_prompt_2 = prompt_template.format(query="Describe the paralinguistic features of the audio")
gender_recognition_prompt = prompt_template.format(query="What is the gender of the speaker")

More flexible prompts for enriched responses

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"

prompt_1 = prompt_template.format(query="describe the paralinguistics feature and return in json format.")
prompt_2 = prompt_template.format(query="Please summarize the content of the speech and analyse the paralinguistics features of this audio. Return in json format.")
prompt_3 = prompt_template.format(query="Please translate this speech to Singapore's 4 official languages.")

AI agent prompts (beyond the default prompt template)

prompt_1 = \
"""
You are MERaLiON-AudioLLM, an empathic AI assistant developed by A*STAR. MERaLiON stands for Multimodal Empathetic Reasoning and Learning in One Network.
You are a friendly and empathetic conversational partner, and is proficient in understanding human emotions, accents, and genders from paralinguistic features.
Maintain a tone that is warm, non-judgmental, and supportive while replying to user. 

User's voice:  <SpeechHere>
"""

Huggingface Inference with CPU

import librosa
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo_id = "MERaLiON/MERaLiON-3-10B"

processor = AutoProcessor.from_pretrained(
    repo_id, 
    trust_remote_code=True,
    )
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    repo_id,
    use_safetensors=True,
    trust_remote_code=True,
)

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"
transcribe_prompt = "Please transcribe this speech."
translate_prompt = "Can you please translate this speech into written Chinese?"

# batch inference of 2 samples
conversation = [
    [{"role": "user", "content": prompt_template.format(query=transcribe_prompt)}],
    [{"role": "user", "content": prompt_template.format(query=translate_prompt)}],
]

chat_prompt = processor.tokenizer.apply_chat_template(
    conversation=conversation,
    tokenize=False,
    add_generation_prompt=True
)

# Use audio at 16000hz.
audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
audio_array = [audio_array]*2
inputs = processor(text=chat_prompt, audios=audio_array)

# adjust the `max_new_tokens` based on your use case.
# Please note the inclusion of `no_repeat_ngram_size=6`.
outputs = model.generate(**inputs, max_new_tokens=256, no_repeat_ngram_size=6)
response = processor.batch_decode(outputs, skip_special_tokens=True)

Huggingface GPU Inference

import torch
import librosa
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo_id = "MERaLiON/MERaLiON-3-10B"
device = "cuda"

processor = AutoProcessor.from_pretrained(
    repo_id, 
    trust_remote_code=True,
    )
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    repo_id,
    use_safetensors=True,
    trust_remote_code=True,
    attn_implementation="flash_attention_2",
    torch_dtype=torch.bfloat16
).to(device)

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"
transcribe_prompt = "Please transcribe this speech."
translate_prompt = "Can you please translate this speech into written Chinese?"

# batch inference of 2 samples
conversation = [
    [{"role": "user", "content": prompt_template.format(query=transcribe_prompt)}],
    [{"role": "user", "content": prompt_template.format(query=translate_prompt)}],
]

chat_prompt = processor.tokenizer.apply_chat_template(
    conversation=conversation,
    tokenize=False,
    add_generation_prompt=True
)

# Use audio at 16000hz.
audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
audio_array = [audio_array]*2
inputs = processor(text=chat_prompt, audios=audio_array)

inputs = inputs.to(device, dtype=torch.bfloat16)

# adjust the `max_new_tokens` based on your use case.
# Please note the inclusion of `no_repeat_ngram_size=6`.
outputs = model.generate(**inputs, max_new_tokens=256, no_repeat_ngram_size=6)
response = processor.batch_decode(outputs, skip_special_tokens=True)

⚠️ Disclaimer

The current MERaLiON-3 has not been specifically aligned for safety and may generate content that is inappropriate, offensive, or harmful. Developers and users are responsible for performing their own safety fine-tuning and implementing necessary security measures. The authors shall not be held liable for any claims, damages, or other liabilities arising from the use of the released models, weights, or code.

Compute and Infrastructure

MERaLiON-3 was trained on the ASPIRE 2A+ Supercomputer Cluster, provided by National Supercomputing Centre (NSCC), Singapore. ASPIRE 2A+ cluster provides multiple H100 nodes, with each compute node equipped with 8 Nvidia H100 GPUs, 2 TB of RAM, and 30 TB of locally attached NVMe storage. These nodes are interconnected via a rail-optimised, full fat-tree topology, utilising 400 Gb/s NDR InfiniBand cables. Additionally, the cluster incorporates a 2.5 PB SSD-based Lustre file system, linked to the H100 nodes through high-speed InfiniBand connections.

With a global batch size of 768, we trained the current release of MERaLiON-3 for around 250k steps, which took around 2.5 days to complete using 16 nodes, 128 H100 GPUs.

📚 Citation

If you find our work useful, please cite our papers:

MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
AudioBench: A Universal Benchmark for Audio Large Language Models
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
Incorporating contextual paralinguistic understanding in large speech-language models
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish

@misc{he2024meralionaudiollmtechnicalreport,
      title={MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models}, 
      author={{MERaLiON Team}},
      year={2024},
      eprint={2412.09818},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2412.09818}, 
}
@article{wang2024audiobench,
    title={AudioBench: A Universal Benchmark for Audio Large Language Models},
    author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang, Wenyu and Liu, Zhengyuan and Aw, AiTi and Chen, Nancy F},
    journal={NAACL},
    year={2025}
    }
@inproceedings{wang2025benchmarking,
  title={Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data},
  author={Wang, Qiongqiong and Sailor, Hardik Bhupendra and Liu, Tianchi and Zhang, Wenyu and Huzaifah, Muhammad and Lertcheva, Nattadaporn and Sun, Shuo and Chen, Nancy F and Wu, Jinyang and Aw, AiTi},
  booktitle={Findings of EMNLP 2025},
  year={2025}
}
@inproceedings{cpqa_interspeech,
  title={Contextual Paralinguistic Data Creation for  Multi-Modal Speech-LLM: Data Condensation and Spoken {QA} Generation},
  author={Wang, Qiongqiong and Sailor, Hardik B and Liu, Tianchi and Aw, Ai Ti},
  booktitle={Proc. Interspeech},
  year={2025},
}
@inproceedings{cpqa_asru,
  title={Incorporating Contextual Paralinguistic
Understanding in Large Speech-Language Models},
  author={
      Wang, Qiongqiong and Sailor, Hardik B and Wong, Jeremy H. M. and Liu, Tianchi and Sun, Shuo and Zhang, Wenyu and Huzaifah, Muhammad and  Chen, Nancy and Aw, Ai Ti},
  booktitle={Proc. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)},
  year={2025},
}
@article{wang2025advancing,
    title={Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models},
    author={Wang, Bin and Zou, Xunlong and Sun, Shuo and Zhang, Wenyu and He, Yingxu and Liu, Zhuohan and Wei, Chengwei and Chen, Nancy F and Aw, AiTi},
    journal={arXiv preprint arXiv:2501.01034},
    year={2025}
    }
@article{zhang2024mowe,
    title={MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders},
    author={Zhang, Wenyu and Sun, Shuo and Wang, Bin and Zou, Xunlong and Liu, Zhuohan and He, Yingxu and Lin, Geyu and Chen, Nancy F and Aw, Ai Ti},
    journal={ICASSP},
    year={2025}
    }
@misc{huang2025meraliontextllmcrosslingualunderstandinglarge,
      title={MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish}, 
      author={Xin Huang and Tarun Kumar Vangani and Minh Duc Pham and Xunlong Zou and Bin Wang and Zhengyuan Liu and Ai Ti Aw},
      year={2025},
      eprint={2501.08335},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2501.08335}, 
}
automatic-speech-recognition
custom_code
meralion
meralion-3
meralion3
safetensors
transformers

Contributors

lewiswoncy

19 commits

sailorhb

2 commits

YingxuHe

1 commits