MERaLiON/MERaLiON-2-10B

Model

Introduction

13

107 commits

1 linked in READMEs

updated Mar 31, 2026

See the code

README

🔥 MERaLiON-2 🔥

🚀 MERaLiON-2-10B | 🚀 MERaLiON-2-10B-ASR | 🚀 MERaLiON-2-3B

💻 Web Demo | ⚙️ vLLM

Introduction

We are pleased to announce the release of MERaLiON2, the latest addition to the MERaLiON family of speech-text large language models. Our flagship model, MERaLiON-2-10B, demonstrates competitive performance across benchmark evaluations in tasks such as multilingual automatic speech recognition (ASR), speech translation (ST), audio scene understanding, emotion recognition, and general speech comprehension. These results are comparable to those achieved by other state-of-the-art open-source AudioLLMs, including Qwen2.5-Omni-7B and Phi-4-multimodal-instruct.

MERaLiON-2-10B is specifically designed to follow complex instructions with a nuanced understanding of Singapore’s multilingual and multicultural context. It integrates a localized Whisper-large-v3 speech encoder and Gemma-2-9b text decoder. The following graph presents task-specific evaluation scores, assessed using the LLM-as-a-Judge framework across multiple datasets. For the speech translation task, performance is measured using the BLEU metric, where higher scores indicate better translation quality.

model_capability

In addition, we introduce an ASR-optimized variant, MERaLiON-2-10B-ASR, which delivers a 5–30% performance improvement over OpenAI’s whisper-large-v3 on speech recognition tasks. This enhancement spans Singapore’s 4 official languages—English, Mandarin, Malay, and Tamil—as well as 3 South-East Asian languages: Indonesian, Thai, and Vietnamese. The model also demonstrates robust handling of code-switching scenarios and local colloquialisms, reflecting its adaptability to Singapore’s diverse linguistic landscape.

The following visualization illustrates the 1 - Word Error Rate (WER) metric across these seven languages, comparing MERaLiON-2-10B-ASR with other leading models. A higher value indicates better transcription accuracy.

model_capability

We also provide MERaLiON-2-3B that balances performance with reduced computational requirements, enabling broader accessibility and lightweight deployment.

  • Extended Audio Length: Support audio inputs up to 300 seconds (5 minutes) for audio & speech question answering tasks, 30s for a satisfactory performance for speech transcription (ASR) and speech translation (ST) tasks.

  • Expanded Language Coverage: In addition to English, Chinese, and Singlish, V2 introduces support for Malay, Tamil, and other South-East Asia languages including Indonesian, Thai, and Vietnamese.

  • Improved Performance: Achieves higher performance across a wide range of tasks. See the Evaluation section for detailed benchmarks.

  • Higher Quality Training Data: Trained on 120,000 hours of curated speech and audio data, filtered for quality and diversity, with an emphasis on local and multilingual audio sources.

  • Three Model Variants: Available in general-purpose (MERaLiON-2-10B), ASR-optimized (MERaLiON-2-10B-ASR) and light-weight (MERaLiON-2-3B) configurations to balance latency, compute efficiency, and task performance across different deployment needs.

Model Description:

MERaLiON stands for Multimodal Empathetic Reasoning and Learning in One Network.

MERaLiON-2 is a family of Speech-Text Large Language Models tailored for Singapore’s multilingual and multicultural landscape, as well as the wider Southeast Asian region. The 10B model integrates a localized Whisper-Large-V3 speech encoder with the Gemma2-9b-IT text decoder. The 3B model integrates a localized Whisper-Large-V3 speech encoder with the Gemma2-2b-IT text decoder.

MERaLiON-2-10B is finetuned on 120,000 hours of speech and audio data across 6 diverse tasks: Automatic Speech Recognition (ASR), Spoken Question Answering (SQA), Spoken Dialogue Summarization (SDS), Audio Captioning (AC), Audio-Scene Question Answering (ASQA) and Paralinguistic Question Answering (PQA). The model supports long-form audio inputs of up to 300 seconds (5 minutes) and is specifically adapted to handle the linguistic nuances, accents, and dialects commonly found across Singapore and neighboring countries.

  • Developed by: I2R, A*STAR, Singapore
  • Model type: Multimodal LLM
  • Language(s): Primarily English (Global and Singapore), Chinese, with support for audio of regional languages including Malay, Tamil, Indonesian, Thai, and Vietnamese.
  • Audio: Mono channel audio, 16000 hz, up to 300 seconds.
  • License: MERaLiON Public License
  • Demo: MERaLiON-AudioLLM Web Demo

MERaLiON-2 is an upgraded version of MERaLiON-AudioLLM.

Performance:

We benchmark MERaLiON-2 series models with extended AudioBench benchmark against several recently released open-source multimodal models — SALMONN-7B, Qwen2.5-Omni series and Phi-4-Multimodal — as well as two cascade model.

Better Automatic Speech Recognition (ASR) Accuracy

MERaLiON-2-10B-ASR and MERaLiON-2-10B demonstrate leading performance in Singlish, Mandarin, Malay, Tamil, and other Southeast Asian languages, while maintaining competitive results in English compared to Whisper-large-v3. The following table shows the average transcription Word Error Rate by language for the MERaLiON family and other leading AudioLLMs. The Private Dataset includes a collection of Singapore's locally accented speeches with code-switch. Please visit AudioBench benchmark for dataset-level evaluation results.

#T_0910c th { text-align: center; } #T_0910c_row0_col0, #T_0910c_row1_col0, #T_0910c_row2_col0, #T_0910c_row3_col0, #T_0910c_row4_col0, #T_0910c_row5_col0, #T_0910c_row6_col7, #T_0910c_row7_col0, #T_0910c_row8_col0 { font-weight: bold; text-decoration: underline; text-align: center; } #T_0910c_row0_col1, #T_0910c_row1_col1, #T_0910c_row2_col1, #T_0910c_row3_col1, #T_0910c_row4_col1, #T_0910c_row5_col1, #T_0910c_row6_col1, #T_0910c_row7_col1, #T_0910c_row8_col1 { text-align: center; } #T_0910c_row0_col2, #T_0910c_row0_col3, #T_0910c_row0_col4, #T_0910c_row0_col5, #T_0910c_row0_col6, #T_0910c_row0_col7, #T_0910c_row0_col8, #T_0910c_row0_col9, #T_0910c_row0_col10, #T_0910c_row0_col11, #T_0910c_row1_col2, #T_0910c_row1_col3, #T_0910c_row1_col4, #T_0910c_row1_col5, #T_0910c_row1_col6, #T_0910c_row1_col7, #T_0910c_row1_col8, #T_0910c_row1_col9, #T_0910c_row1_col10, #T_0910c_row1_col11, #T_0910c_row2_col2, #T_0910c_row2_col3, #T_0910c_row2_col4, #T_0910c_row2_col5, #T_0910c_row2_col6, #T_0910c_row2_col7, #T_0910c_row2_col8, #T_0910c_row2_col9, #T_0910c_row2_col10, #T_0910c_row2_col11, #T_0910c_row3_col2, #T_0910c_row3_col3, #T_0910c_row3_col4, #T_0910c_row3_col5, #T_0910c_row3_col6, #T_0910c_row3_col7, #T_0910c_row3_col8, #T_0910c_row3_col9, #T_0910c_row3_col10, #T_0910c_row3_col11, #T_0910c_row4_col2, #T_0910c_row4_col3, #T_0910c_row4_col4, #T_0910c_row4_col5, #T_0910c_row4_col6, #T_0910c_row4_col7, #T_0910c_row4_col8, #T_0910c_row4_col9, #T_0910c_row4_col10, #T_0910c_row4_col11, #T_0910c_row5_col2, #T_0910c_row5_col3, #T_0910c_row5_col4, #T_0910c_row5_col5, #T_0910c_row5_col6, #T_0910c_row5_col7, #T_0910c_row5_col8, #T_0910c_row5_col9, #T_0910c_row5_col10, #T_0910c_row5_col11, #T_0910c_row6_col0, #T_0910c_row6_col2, #T_0910c_row6_col3, #T_0910c_row6_col4, #T_0910c_row6_col5, #T_0910c_row6_col6, #T_0910c_row6_col8, #T_0910c_row6_col9, #T_0910c_row6_col10, #T_0910c_row6_col11, #T_0910c_row7_col2, #T_0910c_row7_col3, #T_0910c_row7_col4, #T_0910c_row7_col5, #T_0910c_row7_col6, #T_0910c_row7_col7, #T_0910c_row7_col8, #T_0910c_row7_col9, #T_0910c_row7_col10, #T_0910c_row7_col11, #T_0910c_row8_col2, #T_0910c_row8_col3, #T_0910c_row8_col4, #T_0910c_row8_col5, #T_0910c_row8_col6, #T_0910c_row8_col7, #T_0910c_row8_col8, #T_0910c_row8_col9, #T_0910c_row8_col10, #T_0910c_row8_col11 { text-align: center; }
 MERaLiON-2-10B-ASRMERaLiON-2-10BMERaLiON-2-3Bwhisper_large_v3cascade-whisper_large_v3-llama_3_8b_instructcascade-whisper_large_v2-gemma2_9b_cpt-sea_lionv3_instructMERaLiON-AudioLLM-Whisper-SEA-LIONQwen2.5-Omni-7BSeaLLMs-Audio-7BQwen2.5-Omni-3BSALMONN_7Bphi_4_multimodal_instruct
Thai0.0965260.1093650.1072790.1210730.1202570.1721050.9193300.1264970.1171520.1631501.1910991.510068
Tamil0.2712790.3270810.3440810.4414830.4752250.4923360.5613151.0249162.3254021.3151431.3066941.876722
Singlish0.1298300.1688130.1803950.2489450.2516080.2557170.1438000.4390710.7959900.3893930.4414900.448863
Malay0.1946380.2090740.2798910.2196920.3119210.3143780.2898951.4606640.7655652.9437501.0858673.762933
English0.0785440.0882590.1222950.0808410.0815680.1048300.1105670.1342160.1978240.1103530.1914920.098225
Indonesian0.1210200.1428130.1319500.1371020.1353900.1594760.2983650.1686590.2202270.2052161.6535023.565510
Mandarian0.1036940.1320250.1458780.1709800.1968670.2917330.2911830.1024190.3097820.1304290.9395450.238879
Vietnamese0.1186930.1348080.1551100.1484740.1360750.1640780.9520400.2054910.2220010.1867861.5211741.805643
Private Dataset0.1061500.1123600.1472580.1166300.1184340.1438120.1306670.2227700.4965400.1645560.2733040.229450

Better Instruction Following and Audio Understanding

MERaLiON-2-10B exhibits substantial advancements in speech and audio understanding, as well as paralinguistic tasks. Notably, it adeptly handles complex instructions and responds with enhanced flexibility, effectively preserving the pre-trained knowledge from Gemma during the audio fine-tuning process. This capability enables MERaLiON-2-10B to provide detailed explanations regarding speech content and the speaker's emotional state. Furthermore, with appropriate prompt adjustments, the model can assume various roles, such as a voice assistant, virtual caregiver, or an integral component of sophisticated multi-agent AI systems and software solutions. Please visit AudioBench benchmark for dataset-level evaluation results.

#T_b6ba8 th { text-align: center; } #T_b6ba8_row0_col0, #T_b6ba8_row2_col0, #T_b6ba8_row3_col0, #T_b6ba8_row5_col0, #T_b6ba8_row6_col0, #T_b6ba8_row8_col0, #T_b6ba8_row9_col0, #T_b6ba8_row10_col0 { text-align: center; } #T_b6ba8_row0_col1, #T_b6ba8_row0_col2, #T_b6ba8_row0_col3, #T_b6ba8_row0_col4, #T_b6ba8_row0_col5, #T_b6ba8_row0_col6, #T_b6ba8_row0_col7, #T_b6ba8_row0_col8, #T_b6ba8_row0_col9, #T_b6ba8_row0_col11, #T_b6ba8_row0_col12, #T_b6ba8_row0_col13, #T_b6ba8_row1_col1, #T_b6ba8_row1_col2, #T_b6ba8_row1_col3, #T_b6ba8_row1_col4, #T_b6ba8_row1_col5, #T_b6ba8_row1_col6, #T_b6ba8_row1_col7, #T_b6ba8_row1_col8, #T_b6ba8_row1_col9, #T_b6ba8_row1_col10, #T_b6ba8_row1_col11, #T_b6ba8_row1_col12, #T_b6ba8_row1_col13, #T_b6ba8_row2_col2, #T_b6ba8_row2_col3, #T_b6ba8_row2_col4, #T_b6ba8_row2_col5, #T_b6ba8_row2_col6, #T_b6ba8_row2_col7, #T_b6ba8_row2_col8, #T_b6ba8_row2_col9, #T_b6ba8_row2_col10, #T_b6ba8_row2_col11, #T_b6ba8_row2_col12, #T_b6ba8_row2_col13, #T_b6ba8_row3_col1, #T_b6ba8_row3_col3, #T_b6ba8_row3_col4, #T_b6ba8_row3_col5, #T_b6ba8_row3_col6, #T_b6ba8_row3_col7, #T_b6ba8_row3_col8, #T_b6ba8_row3_col9, #T_b6ba8_row3_col10, #T_b6ba8_row3_col11, #T_b6ba8_row3_col12, #T_b6ba8_row3_col13, #T_b6ba8_row4_col1, #T_b6ba8_row4_col2, #T_b6ba8_row4_col3, #T_b6ba8_row4_col4, #T_b6ba8_row4_col5, #T_b6ba8_row4_col6, #T_b6ba8_row4_col7, #T_b6ba8_row4_col8, #T_b6ba8_row4_col9, #T_b6ba8_row4_col10, #T_b6ba8_row4_col11, #T_b6ba8_row4_col12, #T_b6ba8_row4_col13, #T_b6ba8_row5_col1, #T_b6ba8_row5_col2, #T_b6ba8_row5_col3, #T_b6ba8_row5_col5, #T_b6ba8_row5_col6, #T_b6ba8_row5_col7, #T_b6ba8_row5_col8, #T_b6ba8_row5_col9, #T_b6ba8_row5_col10, #T_b6ba8_row5_col11, #T_b6ba8_row5_col12, #T_b6ba8_row5_col13, #T_b6ba8_row6_col1, #T_b6ba8_row6_col3, #T_b6ba8_row6_col4, #T_b6ba8_row6_col5, #T_b6ba8_row6_col6, #T_b6ba8_row6_col7, #T_b6ba8_row6_col8, #T_b6ba8_row6_col9, #T_b6ba8_row6_col10, #T_b6ba8_row6_col11, #T_b6ba8_row6_col12, #T_b6ba8_row6_col13, #T_b6ba8_row7_col1, #T_b6ba8_row7_col2, #T_b6ba8_row7_col3, #T_b6ba8_row7_col4, #T_b6ba8_row7_col5, #T_b6ba8_row7_col6, #T_b6ba8_row7_col7, #T_b6ba8_row7_col8, #T_b6ba8_row7_col9, #T_b6ba8_row7_col10, #T_b6ba8_row7_col11, #T_b6ba8_row7_col12, #T_b6ba8_row7_col13, #T_b6ba8_row8_col1, #T_b6ba8_row8_col2, #T_b6ba8_row8_col3, #T_b6ba8_row8_col4, #T_b6ba8_row8_col6, #T_b6ba8_row8_col7, #T_b6ba8_row8_col8, #T_b6ba8_row8_col9, #T_b6ba8_row8_col10, #T_b6ba8_row8_col11, #T_b6ba8_row8_col12, #T_b6ba8_row8_col13, #T_b6ba8_row9_col1, #T_b6ba8_row9_col2, #T_b6ba8_row9_col4, #T_b6ba8_row9_col5, #T_b6ba8_row9_col6, #T_b6ba8_row9_col7, #T_b6ba8_row9_col8, #T_b6ba8_row9_col9, #T_b6ba8_row9_col10, #T_b6ba8_row9_col11, #T_b6ba8_row9_col12, #T_b6ba8_row9_col13, #T_b6ba8_row10_col1, #T_b6ba8_row10_col3, #T_b6ba8_row10_col4, #T_b6ba8_row10_col5, #T_b6ba8_row10_col6, #T_b6ba8_row10_col7, #T_b6ba8_row10_col8, #T_b6ba8_row10_col9, #T_b6ba8_row10_col10, #T_b6ba8_row10_col11, #T_b6ba8_row10_col12, #T_b6ba8_row10_col13 { text-align: center; } #T_b6ba8_row0_col10, #T_b6ba8_row2_col1, #T_b6ba8_row3_col2, #T_b6ba8_row5_col4, #T_b6ba8_row6_col2, #T_b6ba8_row8_col5, #T_b6ba8_row9_col3, #T_b6ba8_row10_col2 { font-weight: bold; text-decoration: underline; text-align: center; } #T_b6ba8_row1_col0, #T_b6ba8_row4_col0, #T_b6ba8_row7_col0 { font-weight: bold; text-decoration: underline; text-align: center; }
 MERaLiON-2-10BMERaLiON-AudioLLM-Whisper-SEA-LIONMERaLiON-2-10B-ASRMERaLiON-2-3BSeaLLMs-Audio-7BQwen2-Audio-7B-InstructQwen2.5-Omni-3Bphi_4_multimodal_instructcascade-whisper_large_v3-llama_3_8b_instructQwen2.5-Omni-7Bcascade-whisper_large_v2-gemma2_9b_cpt-sea_lionv3_instructQwen-Audio-ChatSALMONN_7BWavLLM_fairseq
Speech Instruction70.20000070.80000013.40000019.10000066.90000048.70000065.00000036.20000066.10000058.30000072.90000010.20000012.90000020.400000
Emotion Recognition63.73626848.57731353.69329854.04079752.00757649.84654033.03783640.67780050.93757831.46939748.21496941.67155133.58486950.801545
Audio Scene Question Answering51.14037452.20775649.51188646.14135350.19373947.04802548.12322842.21714321.87694345.66915318.04368151.61862251.81695833.034083
Gender Recognition95.10942397.17739697.22033593.81026675.44939295.96326647.86721070.71804757.03940948.72471119.42113060.34934984.36509260.773275
Spoken QA (Singlish)66.55000058.90000061.85000059.70000051.35000046.70000060.50000061.95000059.35000058.40000053.75000042.30000043.20000051.200000
Audio Captioning35.60427036.97641934.46671033.24383945.08937237.27881039.20032830.8324092.91577831.8962433.14056839.98866328.8805706.200867
Spoken Dialogue Summarisation53.10000053.60000055.80000048.55000045.45000036.30000046.75000050.75000045.85000043.15000051.00000025.25000014.40000039.450000
Spoken QA (English)79.73504963.71148173.97583468.71517970.92051968.88856567.81854675.51315278.52656968.41513167.81453866.06904760.64907170.595242
Music Understanding63.94271351.34793660.65711955.60235963.68997571.60909959.30918355.26537556.69755747.59898950.46335359.05644549.70513944.313395
Accent Recognition41.81539643.79979947.78886460.05498110.14383610.9013970.4786943.09761521.3984820.58729325.92969317.55029411.57738114.294613
Speech Translation27.39111527.08636628.54035922.13025821.14321510.82666621.77662813.82711013.53627220.68824121.4379974.97318413.4860039.046791

How to Use

[!WARNING] Out of Scope use: This model is not intended for use in tool calling, math, and coding tasks.

MERaLiON-2 requires transformers version 4.50.1

pip install transformers==4.50.1
pip install librosa

Audio Input

  • For ASR tasks, the maximum audio length is suggested to be 30 seconds at 16,000 Hz.
  • For general speech & audio understanding tasks, the maximum audio length is suggested to be 300 seconds at 16,000 Hz sampling rate.

Text Prompt

MERaLiON-2 is trained with this prompt template:

Instruction: <TextHere> \nFollow the text instruction based on the following audio: <SpeechHere>

It is generally recommended to follow this template, i.e., replace <TextHere> with your text instruction while leaving the <SpeechHere> untouched. We list a few useful example prompts here:

Standard prompts for better accuracy

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"

transcription_prompt = prompt_template.format(query="Please transcribe this speech.")
translation_prompt = prompt_template.format(query="Please translate the speech into Malay")
summarization_prompt = prompt_template.format(query="Please summarize this speech")
audio_captioning_prompt_1 = prompt_template.format(query="Please describe the audio")
audio_captioning_prompt_2 = prompt_template.format(query="Please create a caption for the audio")
audio_scene_understanding_prompt = prompt_template.format(query="Is there people crying in the audio?")
speech_as_instruction_prompt = prompt_template.format(query="Please respond to the audio") # given an speech instruction is provided in the audio clip.
emotion_recognition_prompt_1 = prompt_template.format(query="What is the emotion of the speaker")
emotion_recognition_prompt_2 = prompt_template.format(query="Describe the paralinguistics feature of the audio")
gender_recognition_prompt = prompt_template.format(query="What is the gender of the speaker")

More flexible prompts for enriched responses

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"

prompt_1 = prompt_template.format(query="describe the paralinguistics feature and return in json format.")
prompt_2 = prompt_template.format(query="Please summarise the content of the speech and analyse the paralinguistics features of this audio. Return in json format.")
prompt_3 = prompt_template.format(query="Please translate this speech to Singapore's 4 official languages.")

AI agent prompts (beyond the default prompt template)

prompt_1 = \
"""
Your are MERaLiON-AudioLLM, an empathic AI assistant developed by A*STAR. MERaLiON stands for Multimodal Empathetic Reasoning and Learning in One Network.
You are a friendly and empathetic conversational partner, and is proficient in understanding human's emotion, accent, and gender from paralinguistic features.
Maintain a tone that is warm, non-judgmental, and supportive while replying to user. 

User's voice:  <SpeechHere>
"""

Huggingface Inference with CPU

import librosa
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo_id = "MERaLiON/MERaLiON-2-10B"

processor = AutoProcessor.from_pretrained(
    repo_id, 
    trust_remote_code=True,
    )
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    repo_id,
    use_safetensors=True,
    trust_remote_code=True,
)

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"
transcribe_prompt = "Please transcribe this speech."
translate_prompt = "Can you please translate this speech into written Chinese?"

# batch inference of 2 samples
conversation = [
    [{"role": "user", "content": prompt_template.format(query=transcribe_prompt)}],
    [{"role": "user", "content": prompt_template.format(query=translate_prompt)}],
]

chat_prompt = processor.tokenizer.apply_chat_template(
    conversation=conversation,
    tokenize=False,
    add_generation_prompt=True
)

# Use audio at 16000hz.
audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
audio_array = [audio_array]*2
inputs = processor(text=chat_prompt, audios=audio_array)

# adjust the `max_new_tokens` based on your use case.
outputs = model.generate(**inputs, max_new_tokens=256)
generated_ids = outputs[:, inputs['input_ids'].size(1):]
response = processor.batch_decode(generated_ids, skip_special_tokens=True)

Huggingface GPU Inference

import torch
import librosa
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo_id = "MERaLiON/MERaLiON-2-10B"
device = "cuda"

processor = AutoProcessor.from_pretrained(
    repo_id, 
    trust_remote_code=True,
    )
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    repo_id,
    use_safetensors=True,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16
).to(device)

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"
transcribe_prompt = "Please transcribe this speech."
translate_prompt = "Can you please translate this speech into written Chinese?"

# batch inference of 2 samples
conversation = [
    [{"role": "user", "content": prompt_template.format(query=transcribe_prompt)}],
    [{"role": "user", "content": prompt_template.format(query=translate_prompt)}],
]

chat_prompt = processor.tokenizer.apply_chat_template(
    conversation=conversation,
    tokenize=False,
    add_generation_prompt=True
)

# Use audio at 16000hz.
audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
audio_array = [audio_array]*2
inputs = processor(text=chat_prompt, audios=audio_array)

for key, value in inputs.items():
    if isinstance(value, torch.Tensor):
        inputs[key] = inputs[key].to(device)

        if value.dtype == torch.float32:
            inputs[key] = inputs[key].to(torch.bfloat16)

# adjust the `max_new_tokens` based on your use case.
outputs = model.generate(**inputs, max_new_tokens=256)
generated_ids = outputs[:, inputs['input_ids'].size(1):]
response = processor.batch_decode(generated_ids, skip_special_tokens=True)

⚠️ Disclaimer

The current MERaLiON-2 has not been specifically aligned for safety and may generate content that is inappropriate, offensive, or harmful. Developers and users are responsible for performing their own safety fine-tuning and implementing necessary security measures. The authors shall not be held liable for any claims, damages, or other liabilities arising from the use of the released models, weights, or code.

Compute and Infrastructure

MERaLiON-2 was trained on the ASPIRE 2A+ Supercomputer Cluster, provided by National Supercomputing Centre (NSCC), Singapore. ASPIRE 2A+ cluster provides multiple H100 nodes, with each compute node equipped with 8 Nvidia H100 GPUs, 2 TB of RAM, and 30 TB of locally attached NVMe storage. These nodes are interconnected via a rail-optimised, full fat-tree topology, utilising 400 Gb/s NDR InfiniBand cables. Additionally, the cluster incorporates a 2.5 PB SSD-based Lustre file system, linked to the H100 nodes through high-speed InfiniBand connections.

With a global batch size of 768, we trained the current release of MERaLiON-2 for around 200k steps, which took around 2 days to complete using 16 nodes, 128 H100 GPUs.

📚 Citation

If you find our work useful, please cite our papers:

MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
AudioBench: A Universal Benchmark for Audio Large Language Models
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish

@misc{he2024meralionaudiollmtechnicalreport,
      title={MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models}, 
      author={{MERaLiON Team}},
      year={2024},
      eprint={2412.09818},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2412.09818}, 
}
@article{wang2024audiobench,
    title={AudioBench: A Universal Benchmark for Audio Large Language Models},
    author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang, Wenyu and Liu, Zhengyuan and Aw, AiTi and Chen, Nancy F},
    journal={NAACL},
    year={2025}
    }
@article{wang2025advancing,
    title={Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models},
    author={Wang, Bin and Zou, Xunlong and Sun, Shuo and Zhang, Wenyu and He, Yingxu and Liu, Zhuohan and Wei, Chengwei and Chen, Nancy F and Aw, AiTi},
    journal={arXiv preprint arXiv:2501.01034},
    year={2025}
    }
@article{zhang2024mowe,
    title={MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders},
    author={Zhang, Wenyu and Sun, Shuo and Wang, Bin and Zou, Xunlong and Liu, Zhuohan and He, Yingxu and Lin, Geyu and Chen, Nancy F and Aw, Ai Ti},
    journal={ICASSP},
    year={2025}
    }
@misc{huang2025meraliontextllmcrosslingualunderstandinglarge,
      title={MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish}, 
      author={Xin Huang and Tarun Kumar Vangani and Minh Duc Pham and Xunlong Zou and Bin Wang and Zhengyuan Liu and Ai Ti Aw},
      year={2025},
      eprint={2501.08335},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2501.08335}, 
}
automatic-speech-recognition
custom_code
meralion
meralion-2
meralion2
safetensors
transformers

Contributors

YingxuHe

47 commits

ssfei81

26 commits

zxl

16 commits

zhuohan-7

12 commits

MERaLiON/MERaLiON-2-10B

Model

Introduction

13

107 commits

1 linked in READMEs

updated Mar 31, 2026

See the code

README

🔥 MERaLiON-2 🔥

🚀 MERaLiON-2-10B | 🚀 MERaLiON-2-10B-ASR | 🚀 MERaLiON-2-3B

💻 Web Demo | ⚙️ vLLM

Introduction

We are pleased to announce the release of MERaLiON2, the latest addition to the MERaLiON family of speech-text large language models. Our flagship model, MERaLiON-2-10B, demonstrates competitive performance across benchmark evaluations in tasks such as multilingual automatic speech recognition (ASR), speech translation (ST), audio scene understanding, emotion recognition, and general speech comprehension. These results are comparable to those achieved by other state-of-the-art open-source AudioLLMs, including Qwen2.5-Omni-7B and Phi-4-multimodal-instruct.

MERaLiON-2-10B is specifically designed to follow complex instructions with a nuanced understanding of Singapore’s multilingual and multicultural context. It integrates a localized Whisper-large-v3 speech encoder and Gemma-2-9b text decoder. The following graph presents task-specific evaluation scores, assessed using the LLM-as-a-Judge framework across multiple datasets. For the speech translation task, performance is measured using the BLEU metric, where higher scores indicate better translation quality.

model_capability

In addition, we introduce an ASR-optimized variant, MERaLiON-2-10B-ASR, which delivers a 5–30% performance improvement over OpenAI’s whisper-large-v3 on speech recognition tasks. This enhancement spans Singapore’s 4 official languages—English, Mandarin, Malay, and Tamil—as well as 3 South-East Asian languages: Indonesian, Thai, and Vietnamese. The model also demonstrates robust handling of code-switching scenarios and local colloquialisms, reflecting its adaptability to Singapore’s diverse linguistic landscape.

The following visualization illustrates the 1 - Word Error Rate (WER) metric across these seven languages, comparing MERaLiON-2-10B-ASR with other leading models. A higher value indicates better transcription accuracy.

model_capability

We also provide MERaLiON-2-3B that balances performance with reduced computational requirements, enabling broader accessibility and lightweight deployment.

  • Extended Audio Length: Support audio inputs up to 300 seconds (5 minutes) for audio & speech question answering tasks, 30s for a satisfactory performance for speech transcription (ASR) and speech translation (ST) tasks.

  • Expanded Language Coverage: In addition to English, Chinese, and Singlish, V2 introduces support for Malay, Tamil, and other South-East Asia languages including Indonesian, Thai, and Vietnamese.

  • Improved Performance: Achieves higher performance across a wide range of tasks. See the Evaluation section for detailed benchmarks.

  • Higher Quality Training Data: Trained on 120,000 hours of curated speech and audio data, filtered for quality and diversity, with an emphasis on local and multilingual audio sources.

  • Three Model Variants: Available in general-purpose (MERaLiON-2-10B), ASR-optimized (MERaLiON-2-10B-ASR) and light-weight (MERaLiON-2-3B) configurations to balance latency, compute efficiency, and task performance across different deployment needs.

Model Description:

MERaLiON stands for Multimodal Empathetic Reasoning and Learning in One Network.

MERaLiON-2 is a family of Speech-Text Large Language Models tailored for Singapore’s multilingual and multicultural landscape, as well as the wider Southeast Asian region. The 10B model integrates a localized Whisper-Large-V3 speech encoder with the Gemma2-9b-IT text decoder. The 3B model integrates a localized Whisper-Large-V3 speech encoder with the Gemma2-2b-IT text decoder.

MERaLiON-2-10B is finetuned on 120,000 hours of speech and audio data across 6 diverse tasks: Automatic Speech Recognition (ASR), Spoken Question Answering (SQA), Spoken Dialogue Summarization (SDS), Audio Captioning (AC), Audio-Scene Question Answering (ASQA) and Paralinguistic Question Answering (PQA). The model supports long-form audio inputs of up to 300 seconds (5 minutes) and is specifically adapted to handle the linguistic nuances, accents, and dialects commonly found across Singapore and neighboring countries.

  • Developed by: I2R, A*STAR, Singapore
  • Model type: Multimodal LLM
  • Language(s): Primarily English (Global and Singapore), Chinese, with support for audio of regional languages including Malay, Tamil, Indonesian, Thai, and Vietnamese.
  • Audio: Mono channel audio, 16000 hz, up to 300 seconds.
  • License: MERaLiON Public License
  • Demo: MERaLiON-AudioLLM Web Demo

MERaLiON-2 is an upgraded version of MERaLiON-AudioLLM.

Performance:

We benchmark MERaLiON-2 series models with extended AudioBench benchmark against several recently released open-source multimodal models — SALMONN-7B, Qwen2.5-Omni series and Phi-4-Multimodal — as well as two cascade model.

Better Automatic Speech Recognition (ASR) Accuracy

MERaLiON-2-10B-ASR and MERaLiON-2-10B demonstrate leading performance in Singlish, Mandarin, Malay, Tamil, and other Southeast Asian languages, while maintaining competitive results in English compared to Whisper-large-v3. The following table shows the average transcription Word Error Rate by language for the MERaLiON family and other leading AudioLLMs. The Private Dataset includes a collection of Singapore's locally accented speeches with code-switch. Please visit AudioBench benchmark for dataset-level evaluation results.

#T_0910c th { text-align: center; } #T_0910c_row0_col0, #T_0910c_row1_col0, #T_0910c_row2_col0, #T_0910c_row3_col0, #T_0910c_row4_col0, #T_0910c_row5_col0, #T_0910c_row6_col7, #T_0910c_row7_col0, #T_0910c_row8_col0 { font-weight: bold; text-decoration: underline; text-align: center; } #T_0910c_row0_col1, #T_0910c_row1_col1, #T_0910c_row2_col1, #T_0910c_row3_col1, #T_0910c_row4_col1, #T_0910c_row5_col1, #T_0910c_row6_col1, #T_0910c_row7_col1, #T_0910c_row8_col1 { text-align: center; } #T_0910c_row0_col2, #T_0910c_row0_col3, #T_0910c_row0_col4, #T_0910c_row0_col5, #T_0910c_row0_col6, #T_0910c_row0_col7, #T_0910c_row0_col8, #T_0910c_row0_col9, #T_0910c_row0_col10, #T_0910c_row0_col11, #T_0910c_row1_col2, #T_0910c_row1_col3, #T_0910c_row1_col4, #T_0910c_row1_col5, #T_0910c_row1_col6, #T_0910c_row1_col7, #T_0910c_row1_col8, #T_0910c_row1_col9, #T_0910c_row1_col10, #T_0910c_row1_col11, #T_0910c_row2_col2, #T_0910c_row2_col3, #T_0910c_row2_col4, #T_0910c_row2_col5, #T_0910c_row2_col6, #T_0910c_row2_col7, #T_0910c_row2_col8, #T_0910c_row2_col9, #T_0910c_row2_col10, #T_0910c_row2_col11, #T_0910c_row3_col2, #T_0910c_row3_col3, #T_0910c_row3_col4, #T_0910c_row3_col5, #T_0910c_row3_col6, #T_0910c_row3_col7, #T_0910c_row3_col8, #T_0910c_row3_col9, #T_0910c_row3_col10, #T_0910c_row3_col11, #T_0910c_row4_col2, #T_0910c_row4_col3, #T_0910c_row4_col4, #T_0910c_row4_col5, #T_0910c_row4_col6, #T_0910c_row4_col7, #T_0910c_row4_col8, #T_0910c_row4_col9, #T_0910c_row4_col10, #T_0910c_row4_col11, #T_0910c_row5_col2, #T_0910c_row5_col3, #T_0910c_row5_col4, #T_0910c_row5_col5, #T_0910c_row5_col6, #T_0910c_row5_col7, #T_0910c_row5_col8, #T_0910c_row5_col9, #T_0910c_row5_col10, #T_0910c_row5_col11, #T_0910c_row6_col0, #T_0910c_row6_col2, #T_0910c_row6_col3, #T_0910c_row6_col4, #T_0910c_row6_col5, #T_0910c_row6_col6, #T_0910c_row6_col8, #T_0910c_row6_col9, #T_0910c_row6_col10, #T_0910c_row6_col11, #T_0910c_row7_col2, #T_0910c_row7_col3, #T_0910c_row7_col4, #T_0910c_row7_col5, #T_0910c_row7_col6, #T_0910c_row7_col7, #T_0910c_row7_col8, #T_0910c_row7_col9, #T_0910c_row7_col10, #T_0910c_row7_col11, #T_0910c_row8_col2, #T_0910c_row8_col3, #T_0910c_row8_col4, #T_0910c_row8_col5, #T_0910c_row8_col6, #T_0910c_row8_col7, #T_0910c_row8_col8, #T_0910c_row8_col9, #T_0910c_row8_col10, #T_0910c_row8_col11 { text-align: center; }
 MERaLiON-2-10B-ASRMERaLiON-2-10BMERaLiON-2-3Bwhisper_large_v3cascade-whisper_large_v3-llama_3_8b_instructcascade-whisper_large_v2-gemma2_9b_cpt-sea_lionv3_instructMERaLiON-AudioLLM-Whisper-SEA-LIONQwen2.5-Omni-7BSeaLLMs-Audio-7BQwen2.5-Omni-3BSALMONN_7Bphi_4_multimodal_instruct
Thai0.0965260.1093650.1072790.1210730.1202570.1721050.9193300.1264970.1171520.1631501.1910991.510068
Tamil0.2712790.3270810.3440810.4414830.4752250.4923360.5613151.0249162.3254021.3151431.3066941.876722
Singlish0.1298300.1688130.1803950.2489450.2516080.2557170.1438000.4390710.7959900.3893930.4414900.448863
Malay0.1946380.2090740.2798910.2196920.3119210.3143780.2898951.4606640.7655652.9437501.0858673.762933
English0.0785440.0882590.1222950.0808410.0815680.1048300.1105670.1342160.1978240.1103530.1914920.098225
Indonesian0.1210200.1428130.1319500.1371020.1353900.1594760.2983650.1686590.2202270.2052161.6535023.565510
Mandarian0.1036940.1320250.1458780.1709800.1968670.2917330.2911830.1024190.3097820.1304290.9395450.238879
Vietnamese0.1186930.1348080.1551100.1484740.1360750.1640780.9520400.2054910.2220010.1867861.5211741.805643
Private Dataset0.1061500.1123600.1472580.1166300.1184340.1438120.1306670.2227700.4965400.1645560.2733040.229450

Better Instruction Following and Audio Understanding

MERaLiON-2-10B exhibits substantial advancements in speech and audio understanding, as well as paralinguistic tasks. Notably, it adeptly handles complex instructions and responds with enhanced flexibility, effectively preserving the pre-trained knowledge from Gemma during the audio fine-tuning process. This capability enables MERaLiON-2-10B to provide detailed explanations regarding speech content and the speaker's emotional state. Furthermore, with appropriate prompt adjustments, the model can assume various roles, such as a voice assistant, virtual caregiver, or an integral component of sophisticated multi-agent AI systems and software solutions. Please visit AudioBench benchmark for dataset-level evaluation results.

#T_b6ba8 th { text-align: center; } #T_b6ba8_row0_col0, #T_b6ba8_row2_col0, #T_b6ba8_row3_col0, #T_b6ba8_row5_col0, #T_b6ba8_row6_col0, #T_b6ba8_row8_col0, #T_b6ba8_row9_col0, #T_b6ba8_row10_col0 { text-align: center; } #T_b6ba8_row0_col1, #T_b6ba8_row0_col2, #T_b6ba8_row0_col3, #T_b6ba8_row0_col4, #T_b6ba8_row0_col5, #T_b6ba8_row0_col6, #T_b6ba8_row0_col7, #T_b6ba8_row0_col8, #T_b6ba8_row0_col9, #T_b6ba8_row0_col11, #T_b6ba8_row0_col12, #T_b6ba8_row0_col13, #T_b6ba8_row1_col1, #T_b6ba8_row1_col2, #T_b6ba8_row1_col3, #T_b6ba8_row1_col4, #T_b6ba8_row1_col5, #T_b6ba8_row1_col6, #T_b6ba8_row1_col7, #T_b6ba8_row1_col8, #T_b6ba8_row1_col9, #T_b6ba8_row1_col10, #T_b6ba8_row1_col11, #T_b6ba8_row1_col12, #T_b6ba8_row1_col13, #T_b6ba8_row2_col2, #T_b6ba8_row2_col3, #T_b6ba8_row2_col4, #T_b6ba8_row2_col5, #T_b6ba8_row2_col6, #T_b6ba8_row2_col7, #T_b6ba8_row2_col8, #T_b6ba8_row2_col9, #T_b6ba8_row2_col10, #T_b6ba8_row2_col11, #T_b6ba8_row2_col12, #T_b6ba8_row2_col13, #T_b6ba8_row3_col1, #T_b6ba8_row3_col3, #T_b6ba8_row3_col4, #T_b6ba8_row3_col5, #T_b6ba8_row3_col6, #T_b6ba8_row3_col7, #T_b6ba8_row3_col8, #T_b6ba8_row3_col9, #T_b6ba8_row3_col10, #T_b6ba8_row3_col11, #T_b6ba8_row3_col12, #T_b6ba8_row3_col13, #T_b6ba8_row4_col1, #T_b6ba8_row4_col2, #T_b6ba8_row4_col3, #T_b6ba8_row4_col4, #T_b6ba8_row4_col5, #T_b6ba8_row4_col6, #T_b6ba8_row4_col7, #T_b6ba8_row4_col8, #T_b6ba8_row4_col9, #T_b6ba8_row4_col10, #T_b6ba8_row4_col11, #T_b6ba8_row4_col12, #T_b6ba8_row4_col13, #T_b6ba8_row5_col1, #T_b6ba8_row5_col2, #T_b6ba8_row5_col3, #T_b6ba8_row5_col5, #T_b6ba8_row5_col6, #T_b6ba8_row5_col7, #T_b6ba8_row5_col8, #T_b6ba8_row5_col9, #T_b6ba8_row5_col10, #T_b6ba8_row5_col11, #T_b6ba8_row5_col12, #T_b6ba8_row5_col13, #T_b6ba8_row6_col1, #T_b6ba8_row6_col3, #T_b6ba8_row6_col4, #T_b6ba8_row6_col5, #T_b6ba8_row6_col6, #T_b6ba8_row6_col7, #T_b6ba8_row6_col8, #T_b6ba8_row6_col9, #T_b6ba8_row6_col10, #T_b6ba8_row6_col11, #T_b6ba8_row6_col12, #T_b6ba8_row6_col13, #T_b6ba8_row7_col1, #T_b6ba8_row7_col2, #T_b6ba8_row7_col3, #T_b6ba8_row7_col4, #T_b6ba8_row7_col5, #T_b6ba8_row7_col6, #T_b6ba8_row7_col7, #T_b6ba8_row7_col8, #T_b6ba8_row7_col9, #T_b6ba8_row7_col10, #T_b6ba8_row7_col11, #T_b6ba8_row7_col12, #T_b6ba8_row7_col13, #T_b6ba8_row8_col1, #T_b6ba8_row8_col2, #T_b6ba8_row8_col3, #T_b6ba8_row8_col4, #T_b6ba8_row8_col6, #T_b6ba8_row8_col7, #T_b6ba8_row8_col8, #T_b6ba8_row8_col9, #T_b6ba8_row8_col10, #T_b6ba8_row8_col11, #T_b6ba8_row8_col12, #T_b6ba8_row8_col13, #T_b6ba8_row9_col1, #T_b6ba8_row9_col2, #T_b6ba8_row9_col4, #T_b6ba8_row9_col5, #T_b6ba8_row9_col6, #T_b6ba8_row9_col7, #T_b6ba8_row9_col8, #T_b6ba8_row9_col9, #T_b6ba8_row9_col10, #T_b6ba8_row9_col11, #T_b6ba8_row9_col12, #T_b6ba8_row9_col13, #T_b6ba8_row10_col1, #T_b6ba8_row10_col3, #T_b6ba8_row10_col4, #T_b6ba8_row10_col5, #T_b6ba8_row10_col6, #T_b6ba8_row10_col7, #T_b6ba8_row10_col8, #T_b6ba8_row10_col9, #T_b6ba8_row10_col10, #T_b6ba8_row10_col11, #T_b6ba8_row10_col12, #T_b6ba8_row10_col13 { text-align: center; } #T_b6ba8_row0_col10, #T_b6ba8_row2_col1, #T_b6ba8_row3_col2, #T_b6ba8_row5_col4, #T_b6ba8_row6_col2, #T_b6ba8_row8_col5, #T_b6ba8_row9_col3, #T_b6ba8_row10_col2 { font-weight: bold; text-decoration: underline; text-align: center; } #T_b6ba8_row1_col0, #T_b6ba8_row4_col0, #T_b6ba8_row7_col0 { font-weight: bold; text-decoration: underline; text-align: center; }
 MERaLiON-2-10BMERaLiON-AudioLLM-Whisper-SEA-LIONMERaLiON-2-10B-ASRMERaLiON-2-3BSeaLLMs-Audio-7BQwen2-Audio-7B-InstructQwen2.5-Omni-3Bphi_4_multimodal_instructcascade-whisper_large_v3-llama_3_8b_instructQwen2.5-Omni-7Bcascade-whisper_large_v2-gemma2_9b_cpt-sea_lionv3_instructQwen-Audio-ChatSALMONN_7BWavLLM_fairseq
Speech Instruction70.20000070.80000013.40000019.10000066.90000048.70000065.00000036.20000066.10000058.30000072.90000010.20000012.90000020.400000
Emotion Recognition63.73626848.57731353.69329854.04079752.00757649.84654033.03783640.67780050.93757831.46939748.21496941.67155133.58486950.801545
Audio Scene Question Answering51.14037452.20775649.51188646.14135350.19373947.04802548.12322842.21714321.87694345.66915318.04368151.61862251.81695833.034083
Gender Recognition95.10942397.17739697.22033593.81026675.44939295.96326647.86721070.71804757.03940948.72471119.42113060.34934984.36509260.773275
Spoken QA (Singlish)66.55000058.90000061.85000059.70000051.35000046.70000060.50000061.95000059.35000058.40000053.75000042.30000043.20000051.200000
Audio Captioning35.60427036.97641934.46671033.24383945.08937237.27881039.20032830.8324092.91577831.8962433.14056839.98866328.8805706.200867
Spoken Dialogue Summarisation53.10000053.60000055.80000048.55000045.45000036.30000046.75000050.75000045.85000043.15000051.00000025.25000014.40000039.450000
Spoken QA (English)79.73504963.71148173.97583468.71517970.92051968.88856567.81854675.51315278.52656968.41513167.81453866.06904760.64907170.595242
Music Understanding63.94271351.34793660.65711955.60235963.68997571.60909959.30918355.26537556.69755747.59898950.46335359.05644549.70513944.313395
Accent Recognition41.81539643.79979947.78886460.05498110.14383610.9013970.4786943.09761521.3984820.58729325.92969317.55029411.57738114.294613
Speech Translation27.39111527.08636628.54035922.13025821.14321510.82666621.77662813.82711013.53627220.68824121.4379974.97318413.4860039.046791

How to Use

[!WARNING] Out of Scope use: This model is not intended for use in tool calling, math, and coding tasks.

MERaLiON-2 requires transformers version 4.50.1

pip install transformers==4.50.1
pip install librosa

Audio Input

  • For ASR tasks, the maximum audio length is suggested to be 30 seconds at 16,000 Hz.
  • For general speech & audio understanding tasks, the maximum audio length is suggested to be 300 seconds at 16,000 Hz sampling rate.

Text Prompt

MERaLiON-2 is trained with this prompt template:

Instruction: <TextHere> \nFollow the text instruction based on the following audio: <SpeechHere>

It is generally recommended to follow this template, i.e., replace <TextHere> with your text instruction while leaving the <SpeechHere> untouched. We list a few useful example prompts here:

Standard prompts for better accuracy

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"

transcription_prompt = prompt_template.format(query="Please transcribe this speech.")
translation_prompt = prompt_template.format(query="Please translate the speech into Malay")
summarization_prompt = prompt_template.format(query="Please summarize this speech")
audio_captioning_prompt_1 = prompt_template.format(query="Please describe the audio")
audio_captioning_prompt_2 = prompt_template.format(query="Please create a caption for the audio")
audio_scene_understanding_prompt = prompt_template.format(query="Is there people crying in the audio?")
speech_as_instruction_prompt = prompt_template.format(query="Please respond to the audio") # given an speech instruction is provided in the audio clip.
emotion_recognition_prompt_1 = prompt_template.format(query="What is the emotion of the speaker")
emotion_recognition_prompt_2 = prompt_template.format(query="Describe the paralinguistics feature of the audio")
gender_recognition_prompt = prompt_template.format(query="What is the gender of the speaker")

More flexible prompts for enriched responses

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"

prompt_1 = prompt_template.format(query="describe the paralinguistics feature and return in json format.")
prompt_2 = prompt_template.format(query="Please summarise the content of the speech and analyse the paralinguistics features of this audio. Return in json format.")
prompt_3 = prompt_template.format(query="Please translate this speech to Singapore's 4 official languages.")

AI agent prompts (beyond the default prompt template)

prompt_1 = \
"""
Your are MERaLiON-AudioLLM, an empathic AI assistant developed by A*STAR. MERaLiON stands for Multimodal Empathetic Reasoning and Learning in One Network.
You are a friendly and empathetic conversational partner, and is proficient in understanding human's emotion, accent, and gender from paralinguistic features.
Maintain a tone that is warm, non-judgmental, and supportive while replying to user. 

User's voice:  <SpeechHere>
"""

Huggingface Inference with CPU

import librosa
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo_id = "MERaLiON/MERaLiON-2-10B"

processor = AutoProcessor.from_pretrained(
    repo_id, 
    trust_remote_code=True,
    )
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    repo_id,
    use_safetensors=True,
    trust_remote_code=True,
)

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"
transcribe_prompt = "Please transcribe this speech."
translate_prompt = "Can you please translate this speech into written Chinese?"

# batch inference of 2 samples
conversation = [
    [{"role": "user", "content": prompt_template.format(query=transcribe_prompt)}],
    [{"role": "user", "content": prompt_template.format(query=translate_prompt)}],
]

chat_prompt = processor.tokenizer.apply_chat_template(
    conversation=conversation,
    tokenize=False,
    add_generation_prompt=True
)

# Use audio at 16000hz.
audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
audio_array = [audio_array]*2
inputs = processor(text=chat_prompt, audios=audio_array)

# adjust the `max_new_tokens` based on your use case.
outputs = model.generate(**inputs, max_new_tokens=256)
generated_ids = outputs[:, inputs['input_ids'].size(1):]
response = processor.batch_decode(generated_ids, skip_special_tokens=True)

Huggingface GPU Inference

import torch
import librosa
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor

repo_id = "MERaLiON/MERaLiON-2-10B"
device = "cuda"

processor = AutoProcessor.from_pretrained(
    repo_id, 
    trust_remote_code=True,
    )
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    repo_id,
    use_safetensors=True,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16
).to(device)

prompt_template = "Instruction: {query} \nFollow the text instruction based on the following audio: <SpeechHere>"
transcribe_prompt = "Please transcribe this speech."
translate_prompt = "Can you please translate this speech into written Chinese?"

# batch inference of 2 samples
conversation = [
    [{"role": "user", "content": prompt_template.format(query=transcribe_prompt)}],
    [{"role": "user", "content": prompt_template.format(query=translate_prompt)}],
]

chat_prompt = processor.tokenizer.apply_chat_template(
    conversation=conversation,
    tokenize=False,
    add_generation_prompt=True
)

# Use audio at 16000hz.
audio_array, sample_rate = librosa.load("/path/to/your/audio/file", sr=16000)
audio_array = [audio_array]*2
inputs = processor(text=chat_prompt, audios=audio_array)

for key, value in inputs.items():
    if isinstance(value, torch.Tensor):
        inputs[key] = inputs[key].to(device)

        if value.dtype == torch.float32:
            inputs[key] = inputs[key].to(torch.bfloat16)

# adjust the `max_new_tokens` based on your use case.
outputs = model.generate(**inputs, max_new_tokens=256)
generated_ids = outputs[:, inputs['input_ids'].size(1):]
response = processor.batch_decode(generated_ids, skip_special_tokens=True)

⚠️ Disclaimer

The current MERaLiON-2 has not been specifically aligned for safety and may generate content that is inappropriate, offensive, or harmful. Developers and users are responsible for performing their own safety fine-tuning and implementing necessary security measures. The authors shall not be held liable for any claims, damages, or other liabilities arising from the use of the released models, weights, or code.

Compute and Infrastructure

MERaLiON-2 was trained on the ASPIRE 2A+ Supercomputer Cluster, provided by National Supercomputing Centre (NSCC), Singapore. ASPIRE 2A+ cluster provides multiple H100 nodes, with each compute node equipped with 8 Nvidia H100 GPUs, 2 TB of RAM, and 30 TB of locally attached NVMe storage. These nodes are interconnected via a rail-optimised, full fat-tree topology, utilising 400 Gb/s NDR InfiniBand cables. Additionally, the cluster incorporates a 2.5 PB SSD-based Lustre file system, linked to the H100 nodes through high-speed InfiniBand connections.

With a global batch size of 768, we trained the current release of MERaLiON-2 for around 200k steps, which took around 2 days to complete using 16 nodes, 128 H100 GPUs.

📚 Citation

If you find our work useful, please cite our papers:

MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
AudioBench: A Universal Benchmark for Audio Large Language Models
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish

@misc{he2024meralionaudiollmtechnicalreport,
      title={MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models}, 
      author={{MERaLiON Team}},
      year={2024},
      eprint={2412.09818},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2412.09818}, 
}
@article{wang2024audiobench,
    title={AudioBench: A Universal Benchmark for Audio Large Language Models},
    author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang, Wenyu and Liu, Zhengyuan and Aw, AiTi and Chen, Nancy F},
    journal={NAACL},
    year={2025}
    }
@article{wang2025advancing,
    title={Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models},
    author={Wang, Bin and Zou, Xunlong and Sun, Shuo and Zhang, Wenyu and He, Yingxu and Liu, Zhuohan and Wei, Chengwei and Chen, Nancy F and Aw, AiTi},
    journal={arXiv preprint arXiv:2501.01034},
    year={2025}
    }
@article{zhang2024mowe,
    title={MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders},
    author={Zhang, Wenyu and Sun, Shuo and Wang, Bin and Zou, Xunlong and Liu, Zhuohan and He, Yingxu and Lin, Geyu and Chen, Nancy F and Aw, Ai Ti},
    journal={ICASSP},
    year={2025}
    }
@misc{huang2025meraliontextllmcrosslingualunderstandinglarge,
      title={MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish}, 
      author={Xin Huang and Tarun Kumar Vangani and Minh Duc Pham and Xunlong Zou and Bin Wang and Zhengyuan Liu and Ai Ti Aw},
      year={2025},
      eprint={2501.08335},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2501.08335}, 
}
automatic-speech-recognition
custom_code
meralion
meralion-2
meralion2
safetensors
transformers

Contributors

YingxuHe

47 commits

ssfei81

26 commits

zxl

16 commits

zhuohan-7

12 commits