MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long. Conditioned on lyrics and a detailed music description, it generates structurally coherent songs with expressive vocals, evolving arrangements, and stable long-form audio quality.
MiniMax Music 3 combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE. The model produces 32 kHz, 16-bit stereo WAV audio.
Explore music generation examples on the MiniMax Music 3 Demo.
MiniMax Music 3 natively supports full-song generation up to five minutes. The model maintains musical themes, rhythm, vocal identity, and arrangement progression across long sequences, enabling complete structures such as intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro.
The model accepts two complementary inputs:
[Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Instrumental], [Solo], and [Outro].For precise control, we recommend using a Structured Caption with three sections:
This representation allows the model to follow not only a global style, but also the musical development of the song over time.
MiniMax Music 3 uses a hierarchical autoregressive architecture that separates global musical modeling from local acoustic modeling.
The Global LLM is initialized from Qwen3-8B. During training, its embedding and output layers are first adapted to semantic music tokens. The Global and Local LLMs are then jointly trained to model all RVQ codebooks.
Instead of decoding only from discrete RVQ tokens, the synthesis module fuses the final hidden states of the Global and Local LLMs. These continuous representations preserve richer acoustic information for vocal articulation, instrumental texture, and temporal continuity.
The synthesis path is:
Global and Local LLM hidden states
↓
Hidden-state fusion
↓
Flow Matching (2.4B)
↓
Flow-VAE latent
↓
Flow-VAE Decoder (123M)
↓
32 kHz stereo audio
The Flow-VAE architecture is adapted from MiniMax Speech and retrained for the dynamic range and spectral characteristics of music.
The training tokenizer uses eight layers of Residual Vector Quantization (RVQ):
Training first optimizes the semantic codebook, then jointly trains all eight codebooks. At inference time, waveform synthesis uses the fused LLM hidden states and does not require the discrete tokenizer decoder.
MiniMax Music 3 is supported by SGLang-Omni. Follow the official installation guide to prepare the runtime environment.
hf download MiniMaxAI/MiniMax-Music3 --local-dir /path/to/minimax_ttm
We recommend the following inference frameworks to serve the model:
diffusers - see diffusers docs
sgl-omni serve --model-path MiniMaxAI/MiniMax-Music3 --port 8000
The service uses the shared speech API. Put the lyrics in input and the music description in instructions. Put lyric structure tags such as [Verse] and [Chorus] on their own lines.
curl http://127.0.0.1:8000/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "MiniMaxAI/MiniMax-Music3",
"input": "[Verse]\nMorning light filtering through the pine\n[Chorus]\nSoftly the world begins to breathe",
"instructions": "A warm acoustic pop song with intimate female vocals, fingerpicked guitar, soft piano, and a gradual emotional build into a wide final chorus.",
"response_format": "wav",
"seed": 7,
"max_new_tokens": 750,
"stream": false
}' \
--output minimax_music3.wav
max_new_tokens sets the maximum number of audio frames at 25 frames per second. Generation may finish before this limit when the model emits an end-of-audio token. The response is a 32 kHz, 16-bit stereo WAV file.
The following end-to-end example contains the complete lyrics, music description, and generation parameters used to produce the reference audio.
| Use case | Request | Result |
|---|---|---|
| Text-to-music | View script | minimax_ttm.wav |
MiniMax Music 3 is available as a diffusers modular pipeline. Until huggingface/diffusers#14456 is merged, install diffusers from the PR commit:
The snippet below fits 24GB+ VRAM GPUs
pip install git+https://github.com/huggingface/diffusers@dafe3733fcfdbf3c48915fe77be3aef65b5d6a2d transformers accelerate soundfile
import soundfile as sf
import torch
from diffusers import ModularPipeline
pipe = ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-Music3")
pipe.load_components(dtype=torch.bfloat16)
pipe.to("cuda")
lyrics = """[verse]
Morning light filtering through the pine
Every quiet street is yours and mine
[chorus]
Softly the world begins to breathe"""
prompt = (
"Genre: acoustic pop. BPM: 96. Key: C major. Warm and intimate, building gently into the chorus. "
"Vocals: soft female lead, close and breathy, light stacked harmonies in the chorus. "
"Arrangement: fingerpicked guitar and soft piano; brushed drums and upright bass enter in the chorus."
)
audio = pipe(
prompt=prompt,
lyrics=lyrics,
audio_duration=60.0,
generator=torch.Generator("cuda").manual_seed(7),
output="audios",
)[0]
sf.write("song.wav", audio.T.float().cpu().numpy(), pipe.sampling_rate)
The full precision fits under 24GB of VRAM. With automatic CPU offloading, generation takes in ~22 GB; additionally streaming the language model layer by layer makes it fit even 8 GB video cards:
import torch
from diffusers import ComponentsManager, ModularPipeline
from diffusers.hooks import apply_group_offloading
manager = ComponentsManager()
manager.enable_auto_cpu_offload(device="cuda")
pipe = ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-Music3", components_manager=manager)
pipe.load_components(dtype=torch.bfloat16)
# Only needed below ~22 GB of VRAM — slower, but fits in 8 GB.
apply_group_offloading(
pipe.language_model, onload_device=torch.device("cuda"), offload_type="leaf_level", use_stream=True
)
A concise natural-language description can be used directly. For richer prompts and more precise control, use the provided music-caption-rewriter skill to expand it into a Structured Caption containing Global Metadata, Vocal Details, and Arrangement. The skill preserves musical instructions attached to lyric section tags in the arrangement description while keeping the lyric text in the lyrics input.
npx skills add MiniMax-AI/MiniMax-Music3 --skill music-caption-rewriter
Contact us at model@minimax.io.
MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long. Conditioned on lyrics and a detailed music description, it generates structurally coherent songs with expressive vocals, evolving arrangements, and stable long-form audio quality.
MiniMax Music 3 combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE. The model produces 32 kHz, 16-bit stereo WAV audio.
Explore music generation examples on the MiniMax Music 3 Demo.
MiniMax Music 3 natively supports full-song generation up to five minutes. The model maintains musical themes, rhythm, vocal identity, and arrangement progression across long sequences, enabling complete structures such as intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro.
The model accepts two complementary inputs:
[Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Instrumental], [Solo], and [Outro].For precise control, we recommend using a Structured Caption with three sections:
This representation allows the model to follow not only a global style, but also the musical development of the song over time.
MiniMax Music 3 uses a hierarchical autoregressive architecture that separates global musical modeling from local acoustic modeling.
The Global LLM is initialized from Qwen3-8B. During training, its embedding and output layers are first adapted to semantic music tokens. The Global and Local LLMs are then jointly trained to model all RVQ codebooks.
Instead of decoding only from discrete RVQ tokens, the synthesis module fuses the final hidden states of the Global and Local LLMs. These continuous representations preserve richer acoustic information for vocal articulation, instrumental texture, and temporal continuity.
The synthesis path is:
Global and Local LLM hidden states
↓
Hidden-state fusion
↓
Flow Matching (2.4B)
↓
Flow-VAE latent
↓
Flow-VAE Decoder (123M)
↓
32 kHz stereo audio
The Flow-VAE architecture is adapted from MiniMax Speech and retrained for the dynamic range and spectral characteristics of music.
The training tokenizer uses eight layers of Residual Vector Quantization (RVQ):
Training first optimizes the semantic codebook, then jointly trains all eight codebooks. At inference time, waveform synthesis uses the fused LLM hidden states and does not require the discrete tokenizer decoder.
MiniMax Music 3 is supported by SGLang-Omni. Follow the official installation guide to prepare the runtime environment.
hf download MiniMaxAI/MiniMax-Music3 --local-dir /path/to/minimax_ttm
We recommend the following inference frameworks to serve the model:
diffusers - see diffusers docs
sgl-omni serve --model-path MiniMaxAI/MiniMax-Music3 --port 8000
The service uses the shared speech API. Put the lyrics in input and the music description in instructions. Put lyric structure tags such as [Verse] and [Chorus] on their own lines.
curl http://127.0.0.1:8000/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "MiniMaxAI/MiniMax-Music3",
"input": "[Verse]\nMorning light filtering through the pine\n[Chorus]\nSoftly the world begins to breathe",
"instructions": "A warm acoustic pop song with intimate female vocals, fingerpicked guitar, soft piano, and a gradual emotional build into a wide final chorus.",
"response_format": "wav",
"seed": 7,
"max_new_tokens": 750,
"stream": false
}' \
--output minimax_music3.wav
max_new_tokens sets the maximum number of audio frames at 25 frames per second. Generation may finish before this limit when the model emits an end-of-audio token. The response is a 32 kHz, 16-bit stereo WAV file.
The following end-to-end example contains the complete lyrics, music description, and generation parameters used to produce the reference audio.
| Use case | Request | Result |
|---|---|---|
| Text-to-music | View script | minimax_ttm.wav |
MiniMax Music 3 is available as a diffusers modular pipeline. Until huggingface/diffusers#14456 is merged, install diffusers from the PR commit:
The snippet below fits 24GB+ VRAM GPUs
pip install git+https://github.com/huggingface/diffusers@dafe3733fcfdbf3c48915fe77be3aef65b5d6a2d transformers accelerate soundfile
import soundfile as sf
import torch
from diffusers import ModularPipeline
pipe = ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-Music3")
pipe.load_components(dtype=torch.bfloat16)
pipe.to("cuda")
lyrics = """[verse]
Morning light filtering through the pine
Every quiet street is yours and mine
[chorus]
Softly the world begins to breathe"""
prompt = (
"Genre: acoustic pop. BPM: 96. Key: C major. Warm and intimate, building gently into the chorus. "
"Vocals: soft female lead, close and breathy, light stacked harmonies in the chorus. "
"Arrangement: fingerpicked guitar and soft piano; brushed drums and upright bass enter in the chorus."
)
audio = pipe(
prompt=prompt,
lyrics=lyrics,
audio_duration=60.0,
generator=torch.Generator("cuda").manual_seed(7),
output="audios",
)[0]
sf.write("song.wav", audio.T.float().cpu().numpy(), pipe.sampling_rate)
The full precision fits under 24GB of VRAM. With automatic CPU offloading, generation takes in ~22 GB; additionally streaming the language model layer by layer makes it fit even 8 GB video cards:
import torch
from diffusers import ComponentsManager, ModularPipeline
from diffusers.hooks import apply_group_offloading
manager = ComponentsManager()
manager.enable_auto_cpu_offload(device="cuda")
pipe = ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-Music3", components_manager=manager)
pipe.load_components(dtype=torch.bfloat16)
# Only needed below ~22 GB of VRAM — slower, but fits in 8 GB.
apply_group_offloading(
pipe.language_model, onload_device=torch.device("cuda"), offload_type="leaf_level", use_stream=True
)
A concise natural-language description can be used directly. For richer prompts and more precise control, use the provided music-caption-rewriter skill to expand it into a Structured Caption containing Global Metadata, Vocal Details, and Arrangement. The skill preserves musical instructions attached to lyric section tags in the arrangement description while keeping the lyric text in the lyrics input.
npx skills add MiniMax-AI/MiniMax-Music3 --skill music-caption-rewriter
Contact us at model@minimax.io.