bosonai/higgs-tts-3-4b

Model

755

stars

22

commits

5

repos using this model

3

linked in READMEs

Sep 4, 2026

updated

ast
ceb
ckb
controllable-tts
endpoints_compatible
expressive-speech
higgs_multimodal_qwen3
kab
kam
kea
luo
mhr
multilingual-tts
nso
safetensors
speech-generation
text-generation
text-to-speech
transformers
umb
voice-agent

README

Higgs TTS 3

Higgs TTS 3 is built for voice chat: it speaks, not just reads. It turns model responses into expressive conversational speech across 100+ languages, with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.

[!TIP] Released for research and non-commercial use under the Boson Higgs TTS 3 Research and Non-Commercial License. Production, hosted APIs, embedding in a product/service, or reselling the model requires a separate commercial license. Prohibited: voice cloning without consent, impersonation, fraud, election deception, biometric surveillance, or any unlawful use.

[!TIP] Free for digital creators โ€” including monetized content. Under the license's Creator Use Grant, creators may use Higgs TTS 3 to make and monetize podcasts, videos, and social posts for free. The one requirement is to credit Boson AI's Higgs Audio โ€” either in the audio or prominently in the accompanying text (e.g., the video description or show notes). Suggested credit: "This audio was created with Boson AI's Higgs Audio โ€” https://www.boson.ai/higgs-audio". See the Creator Use section below.

Higgs TTS 3 Architecture

Higgs autoregressive decoder consumes interleaved text and audio tokens. Audio is encoded by the Higgs Tokenizer into 8 codebooks at 25 fps, staggered via a delay pattern, then mapped to backbone hidden states through a multi-codebook fused embedding. Output codes pass through a multi-codebook fused head, are de-delayed, and decoded back to waveform.

ComponentSpec
Backbone~4B autoregressive decoder (36 L, hidden=2560, GQA 32/8)
Multi-codebook embedding / headFused single-tensor, tied with text embedding
Context length8,192 tokens (training sequence length)
Audio tokens8 codebooks ร— 1026 vocab, delay pattern
Sample rate24 kHz
Frame rate25 fps (40 ms / frame)

Supported Languages

The model reaches single-digit WER/CER on 102 languages, which split into two tiers.

WER/CER under 5 โ€” polished, production-quality (85)

๐Ÿ‡ฟ๐Ÿ‡ฆ Afrikaans ยท ๐Ÿ‡ธ๐Ÿ‡ฆ๐Ÿ‡ช๐Ÿ‡ฌ Arabic ยท ๐Ÿ‡ฆ๐Ÿ‡ฒ Armenian ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Assamese ยท ๐Ÿ‡ช๐Ÿ‡ธ Asturian ยท ๐Ÿ‡ฆ๐Ÿ‡ฟ Azerbaijani ยท ๐Ÿ‡ท๐Ÿ‡บ Bashkir ยท ๐Ÿ‡ช๐Ÿ‡ธ Basque ยท ๐Ÿ‡ง๐Ÿ‡พ Belarusian ยท ๐Ÿ‡ง๐Ÿ‡ฉ๐Ÿ‡ฎ๐Ÿ‡ณ Bengali ยท ๐Ÿ‡ง๐Ÿ‡ฆ Bosnian ยท ๐Ÿ‡ง๐Ÿ‡ฌ Bulgarian ยท ๐Ÿ‡ช๐Ÿ‡ธ Catalan ยท ๐Ÿ‡ต๐Ÿ‡ญ Cebuano ยท ๐Ÿ‡ฎ๐Ÿ‡ถ Central Kurdish ยท ๐Ÿ‡จ๐Ÿ‡ณ Chinese ยท ๐Ÿ‡ญ๐Ÿ‡ท Croatian ยท ๐Ÿ‡จ๐Ÿ‡ฟ Czech ยท ๐Ÿ‡ฉ๐Ÿ‡ฐ Danish ยท ๐Ÿ‡ณ๐Ÿ‡ฑ๐Ÿ‡ง๐Ÿ‡ช Dutch ยท ๐Ÿ‡ท๐Ÿ‡บ Eastern Mari ยท ๐Ÿ‡บ๐Ÿ‡ธ๐Ÿ‡ฌ๐Ÿ‡ง๐Ÿ‡ฆ๐Ÿ‡บ English ยท ๐ŸŒ Esperanto ยท ๐Ÿ‡ช๐Ÿ‡ช Estonian ยท ๐Ÿ‡ซ๐Ÿ‡ฎ Finnish ยท ๐Ÿ‡ซ๐Ÿ‡ท๐Ÿ‡จ๐Ÿ‡ฆ French ยท ๐Ÿ‡ช๐Ÿ‡ธ Galician ยท ๐Ÿ‡ฌ๐Ÿ‡ช Georgian ยท ๐Ÿ‡ฉ๐Ÿ‡ช๐Ÿ‡ฆ๐Ÿ‡น German ยท ๐Ÿ‡ฌ๐Ÿ‡ท Greek ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Gujarati ยท ๐Ÿ‡ญ๐Ÿ‡น Haitian Creole ยท ๐Ÿ‡ณ๐Ÿ‡ฌ Hausa ยท ๐Ÿ‡ฎ๐Ÿ‡ฑ Hebrew ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Hindi ยท ๐Ÿ‡ญ๐Ÿ‡บ Hungarian ยท ๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesian ยท ๐Ÿ‡ฎ๐Ÿ‡น Italian ยท ๐Ÿ‡ฏ๐Ÿ‡ต Japanese ยท ๐Ÿ‡ฎ๐Ÿ‡ฉ Javanese ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Kannada ยท ๐Ÿ‡ฐ๐Ÿ‡ฟ Kazakh ยท ๐Ÿ‡ฐ๐Ÿ‡ท Korean ยท ๐Ÿ‡ท๐Ÿ‡ผ Kinyarwanda ยท ๐Ÿ‡ฐ๐Ÿ‡ฌ Kyrgyz ยท ๐Ÿ‡ฑ๐Ÿ‡ป Latvian ยท ๐Ÿ‡จ๐Ÿ‡ฉ Lingala ยท ๐Ÿ‡ฑ๐Ÿ‡น Lithuanian ยท ๐Ÿ‡ฐ๐Ÿ‡ช Luo ยท ๐Ÿ‡ฒ๐Ÿ‡ฐ Macedonian ยท ๐Ÿ‡ฒ๐Ÿ‡พ๐Ÿ‡ฎ๐Ÿ‡ฉ Malay ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Malayalam ยท ๐Ÿ‡ฒ๐Ÿ‡น Maltese ยท ๐Ÿ‡ณ๐Ÿ‡ฟ Mฤori ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Marathi ยท ๐Ÿ‡ฒ๐Ÿ‡ณ Mongolian ยท ๐Ÿ‡ณ๐Ÿ‡ต Nepali ยท ๐Ÿ‡ณ๐Ÿ‡ด Norwegian ยท ๐Ÿ‡ซ๐Ÿ‡ท Occitan ยท ๐Ÿ‡ฎ๐Ÿ‡ท๐Ÿ‡ฆ๐Ÿ‡ซ Persian ยท ๐Ÿ‡ต๐Ÿ‡ฑ Polish ยท ๐Ÿ‡ต๐Ÿ‡น๐Ÿ‡ง๐Ÿ‡ท Portuguese ยท ๐Ÿ‡ท๐Ÿ‡ด Romanian ยท ๐Ÿ‡ท๐Ÿ‡บ Russian ยท ๐Ÿ‡ฟ๐Ÿ‡ฆ Sepedi ยท ๐Ÿ‡ท๐Ÿ‡ธ Serbian ยท ๐Ÿ‡ฟ๐Ÿ‡ผ Shona ยท ๐Ÿ‡ธ๐Ÿ‡ฐ Slovak ยท ๐Ÿ‡ธ๐Ÿ‡ฎ Slovene ยท ๐Ÿ‡ช๐Ÿ‡ธ๐Ÿ‡ฒ๐Ÿ‡ฝ Spanish ยท ๐Ÿ‡น๐Ÿ‡ฟ๐Ÿ‡ฐ๐Ÿ‡ช Swahili ยท ๐Ÿ‡ธ๐Ÿ‡ช Swedish ยท ๐Ÿ‡ต๐Ÿ‡ญ Tagalog ยท ๐Ÿ‡น๐Ÿ‡ฏ Tajik ยท ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ฑ๐Ÿ‡ฐ Tamil ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Telugu ยท ๐Ÿ‡น๐Ÿ‡ญ Thai ยท ๐Ÿ‡น๐Ÿ‡ท Turkish ยท ๐Ÿ‡บ๐Ÿ‡ฆ Ukrainian ยท ๐Ÿ‡ต๐Ÿ‡ฐ๐Ÿ‡ฎ๐Ÿ‡ณ Urdu ยท ๐Ÿ‡จ๐Ÿ‡ณ Uyghur ยท ๐Ÿ‡บ๐Ÿ‡ฟ Uzbek ยท ๐Ÿ‡ป๐Ÿ‡ณ Vietnamese ยท ๐Ÿ‡ฟ๐Ÿ‡ฆ Xhosa ยท ๐Ÿ‡ฟ๐Ÿ‡ฆ Zulu

WER/CER between 5 and 10 โ€” usable, but less polished (17)

๐Ÿ‡ฆ๐Ÿ‡ฑ Albanian ยท ๐Ÿ‡ฒ๐Ÿ‡ผ๐Ÿ‡ฟ๐Ÿ‡ฒ Chichewa/Nyanja ยท ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ต๐Ÿ‡ฐ Eastern Punjabi ยท ๐Ÿ‡บ๐Ÿ‡ฌ Ganda ยท ๐Ÿ‡ฎ๐Ÿ‡ธ Icelandic ยท ๐Ÿ‡ฎ๐Ÿ‡ช Irish ยท ๐Ÿ‡ฉ๐Ÿ‡ฟ Kabyle ยท ๐Ÿ‡จ๐Ÿ‡ป Kabuverdianu ยท ๐Ÿ‡ฐ๐Ÿ‡ช Kamba ยท ๐Ÿ‡ป๐Ÿ‡ฆ Latin ยท ๐Ÿ‡ฑ๐Ÿ‡บ Luxembourgish ยท ๐Ÿ‡ช๐Ÿ‡น๐Ÿ‡ฐ๐Ÿ‡ช Oromo ยท ๐Ÿ‡ฆ๐Ÿ‡ซ๐Ÿ‡ต๐Ÿ‡ฐ Pashto ยท ๐Ÿ‡ต๐Ÿ‡ฐ๐Ÿ‡ฎ๐Ÿ‡ณ Sindhi ยท ๐Ÿ‡ธ๐Ÿ‡ด Somali ยท ๐Ÿ‡ฆ๐Ÿ‡ด Umbundu ยท ๐Ÿ‡ฌ๐Ÿ‡ง Welsh

Control Tokens

All tags follow <|category:value|> syntax and can be inserted mid-utterance.

For how to place these tags when writing the target text (sentence-level vs. inline, sfx formatting, stacking, worked examples), see PROMPTING.md.

  • Emotion โ€” elation, amusement, enthusiasm, determination, pride, contentment, affection, relief, contemplation, confusion, surprise, awe, longing, arousal, anger, fear, disgust, bitterness, sadness, shame, helplessness
TokenDescription
<|emotion:elation|>Elation / joy
<|emotion:amusement|>Amusement / playful laughter
<|emotion:enthusiasm|>Enthusiasm / excitement
<|emotion:determination|>Determination / firmness
<|emotion:pride|>Pride / confidence
<|emotion:contentment|>Calm satisfaction
<|emotion:affection|>Warmth / affection
<|emotion:relief|>Relief
<|emotion:contemplation|>Thoughtful / reflective
<|emotion:confusion|>Confused
<|emotion:surprise|>Surprised
<|emotion:awe|>Awe / wonder
<|emotion:longing|>Longing / yearning
<|emotion:arousal|>Heightened desire
<|emotion:anger|>Anger
<|emotion:fear|>Fear
<|emotion:disgust|>Disgust
<|emotion:bitterness|>Bitterness
<|emotion:sadness|>Sadness
<|emotion:shame|>Shame
<|emotion:helplessness|>Helplessness
  • Style โ€” singing, shouting, whispering
TokenDescription
<|style:singing|>Singing
<|style:shouting|>Shouting / projected voice
<|style:whispering|>Whisper
  • Sound effects โ€” cough, laughter, crying, screaming, burping, humming, sigh, sniff, sneeze

Pair each token with the matching onomatopoeia immediately after it.

TokenDescriptionSuggested onomatopoeia
<|sfx:cough|>CoughAhem
<|sfx:laughter|>LaughterHaha / Hehe
<|sfx:crying|>CryingBoohoo / Sob
<|sfx:screaming|>ScreamingAhh / Aaah
<|sfx:burping|>BurpingBurp
<|sfx:humming|>HummingHmm / Mmm
<|sfx:sigh|>SighUh / Ahh
<|sfx:sniff|>SniffSff
<|sfx:sneeze|>SneezeAchoo
  • Prosody
    • Speed โ€” speed_very_slow, speed_slow, speed_fast, speed_very_fast
    • Pauses โ€” pause, long_pause
    • Pitch โ€” pitch_low, pitch_high
    • Delivery โ€” expressive_high, expressive_low
TokenEffect
<|prosody:speed_very_slow|>โ‰ˆ0.65ร— speed
<|prosody:speed_slow|>โ‰ˆ0.85ร— speed
<|prosody:speed_fast|>โ‰ˆ1.2ร— speed
<|prosody:speed_very_fast|>โ‰ˆ1.4ร— speed
<|prosody:pitch_low|>โ‰ˆโˆ’3 semitones
<|prosody:pitch_high|>โ‰ˆ+2.5 semitones
<|prosody:pause|>โ‰ˆ400โ€“700 ms pause
<|prosody:long_pause|>โ‰ˆ700โ€“1500 ms pause
<|prosody:expressive_high|>More expressive delivery
<|prosody:expressive_low|>Flatter delivery

Evaluation Benchmarks

Multilingual Voice Clone

We evaluate Higgs TTS 3 on public multilingual TTS suites and our internal 111-language Higgs-Multilingual set, covering both common and lower-resource languages.

WER / CER (โ†“, ร—100) macro-averaged across each benchmark's language set. Lower is better; bold marks the best per row. All numbers are reproducible end-to-end with original metrics and normalization.

BenchmarkHiggs TTS v2Higgs TTS 3Fish Audio S2 ProQwen3-TTS-1.7BVibeVoice-7BIndexTTS-2MiMo-Audio-7B-InstructMOSS-TTS-v1.5OmniVoiceChatterBoxFireRedTTS-2
SeedTTS2.101.111.311.303.591.633.701.731.2117.001.72
CV321.194.414.607.7311.66129.2671.556.114.9232.6219.20
MiniMax-Multilingual49.862.745.1527.418.21112.9185.673.782.9849.3012.52
Higgs-Multilingual52.243.618.6897.0913.7457.7159.6121.283.6357.5233.69

Emergent TTS

Win-rate (โ†‘) per category โ€” judge preference vs the BASELINE row; bold marks the highest win-rate per column. For a fair comparison, every model shares the same reference audio per prompt, and we run the benchmark text verbatim โ€” no inline control tags inserted.

ModelOverall โ†‘Emotions โ†‘Foreign Words โ†‘Paralinguistics โ†‘Complex Pronunciation โ†‘Questions โ†‘Syntactic Complexity โ†‘
Higgs TTS 353.65%53.75%48.75%68.57%25.10%61.43%60.71%
Fish Audio S2 Pro43.80%53.04%33.93%53.75%18.16%55.00%45.71%
Qwen3-TTS-1.7B38.84%45.54%24.64%44.29%30.00%53.39%34.11%
IndexTTS-231.12%39.29%5.36%42.50%12.45%45.89%38.93%
MOSS-TTS-v1.543.89%60.54%35.18%51.43%11.63%53.21%47.32%
OmniVoice40.82%61.07%28.75%52.68%13.67%45.00%40.36%

Usage

SGLang Usage

Pair the weights in this repo with SGLang-Omni โ€” a production serving stack with continuous batching for multi-codebook decoding and the same inline tag controls. The Higgs TTS cookbook walks you through installation, server launch, request examples, and the full API reference.

See the Higgs TTS cookbook for the full details.

Install and Serve

docker pull lmsysorg/sglang-omni:dev
docker run -it --gpus all --shm-size 32g --ipc host --network host --privileged \
  lmsysorg/sglang-omni:dev /bin/zsh

git clone git@github.com:sgl-project/sglang-omni.git && cd sglang-omni
uv venv .venv -p 3.12 && source .venv/bin/activate
uv pip install -v -e .
export HF_TOKEN=hf_xxxxxxxxxxxxxxxx
hf download bosonai/higgs-tts-3-4b

sgl-omni serve \
  --model-path bosonai/higgs-tts-3-4b \
  --port 8000

Zero-shot synthesis

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"input": "Hello, how are you?"}' \
  --output output.wav

Voice cloning

Supplying the reference transcript (text) materially improves cloning fidelity.

import requests

resp = requests.post(
    "http://localhost:8000/v1/audio/speech",
    json={
        "input": "Have a nice day and enjoy south california sunshine.",
        "references": [{
            "audio_path": "ref.wav",
            "text": "Hey, Adam here. Let's create something that feels real, sounds human, and connects every time.",
        }],
        "temperature": 0.8, "top_k": 50, "max_new_tokens": 1024,
    },
)
with open("output.wav", "wb") as f:
    f.write(resp.content)

Streaming (Server-Sent Events)

Set "stream": true to receive base64-encoded WAV chunks as the vocoder emits them โ€” sub-second time-to-first-audio. Each event carries audio.data (base64 WAV bytes); the terminal event has finish_reason: "stop" plus usage metadata.

import requests, base64, json

with requests.post(
    "http://localhost:8000/v1/audio/speech",
    json={"input": "Get the trust fund to the bank early.", "stream": True},
    stream=True,
) as resp, open("output.wav", "wb") as f:
    for line in resp.iter_lines():
        if not line or not line.startswith(b"data: ") or line == b"data: [DONE]":
            continue
        event = json.loads(line[6:])
        if event.get("finish_reason") == "stop":
            break
        audio = event.get("audio") or {}
        if audio.get("data"):
            f.write(base64.b64decode(audio["data"]))

Inline control tokens

Embed <|emotion:โ€ฆ|>, <|style:โ€ฆ|>, <|prosody:โ€ฆ|>, and <|sfx:โ€ฆ|> tokens directly in input. Two rules:

  1. Delivery tokens first. Emotion, style, and the prosody speed / pitch / expressive tokens shape the whole turn โ€” put them at the start of input. Positional tokens (<|prosody:pause|>, <|prosody:long_pause|>, <|sfx:โ€ฆ|>) go inline exactly where they fire.
  2. Pair every <|sfx:โ€ฆ|> with its onomatopoeia. E.g. <|sfx:laughter|>Haha, <|sfx:sigh|>Uh, <|sfx:sneeze|>Achoo. The written sound gives the model the acoustic cue to realize the effect.

Example โ€” amusement + laughter:

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"input": "<|emotion:amusement|><|prosody:expressive_high|>Wait, wait, that was kind of hilarious. <|sfx:laughter|>Hehe, no, seriously, I was not ready for that."}' \
  --output output.wav

Throughput

Throughput on Seed-TTS EN (full set, N=1088 per run). Client --max-concurrency sweep against a Higgs server (max_running_requests=16, bf16, CUDA Graph on). Each row is the mean of 3 runs. Hardware: 1ร— H100.

ConcurrencyThroughput (req/s)Mean latencyRTF (per-req)audio_s/s
11.62617 ms0.1476.89
22.70742 ms0.18011.37
45.45733 ms0.17722.84
88.91898 ms0.21737.38
1614.741079 ms0.26261.84
  • Concurrency โ€” Maximum number of in-flight client requests (--max-concurrency).
  • Throughput (req/s) โ€” Completed requests divided by total benchmark wall-clock time.
  • Mean latency โ€” Average end-to-end time per request (send to full response received).
  • RTF (per-req) โ€” Average ratio of processing time to generated audio duration per request (<1 is faster than real time).
  • audio_s/s โ€” Total seconds of audio produced divided by total benchmark wall-clock time.

To reproduce the results, follow the instructions in this script.

vLLM-Omni Usage

You can also serve these weights with vLLM-Omni, which exposes the same OpenAI-compatible /v1/audio/speech API with zero-shot voice cloning.

hf download bosonai/higgs-tts-3-4b

vllm-omni serve bosonai/higgs-tts-3-4b \
  --host 0.0.0.0 --port 8095 \
  --trust-remote-code --omni
curl -X POST http://localhost:8095/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model": "bosonai/higgs-tts-3-4b", "input": "Hello, how are you?"}' \
  --output output.wav

Plain-text TTS, voice-clone, and benchmark recipes are in the vLLM-Omni Higgs TTS 3 recipe.

API Usage

For zero-ops deployment, use the Boson AI API.

Citation

@misc{bosonai_higgs_audio_tts_v3_2026,
  title  = {Higgs TTS 3: Conversational Speech for Voice AI from Boson AI},
  author = {Boson AI},
  year   = {2026},
  howpublished = {https://huggingface.co/bosonai/higgs-tts-3-4b},
}

Creator Use (Free for Digital Creators)

In addition to research and non-commercial use, the license includes a Creator Use Grant that lets digital creators use Higgs TTS 3 to produce creative content for free โ€” including monetized content.

What's covered

  • Podcasts, videos, audiobooks, social media posts, and similar creative works
  • Personal and commercial/monetized creator channels (ad-supported, sponsored, subscription, etc.)

The one requirement: acknowledge Boson AI's Higgs Audio. The acknowledgment must appear in at least one of the following ways:

  • In the audio โ€” e.g. "This audio was created with Boson AI's Higgs Audio."
  • In the accompanying text, displayed prominently โ€” e.g. in the post body, video description, or show notes. It must be clearly visible and not hidden at the bottom of the credits or annotations.

Suggested credit string:

This audio was created with Boson AI's Higgs Audio โ€” https://www.boson.ai/higgs-audio

Still requires a separate commercial license. The Creator Use Grant covers creating content with the model. It does not cover hosting the model behind an API or as a service, redistributing/reselling/fine-tuning the model for resale, or embedding the model in a product or application. For these uses, contact us for a commercial license.

The Creator Use Grant does not change the use restrictions โ€” no non-consensual voice cloning or impersonation, no fraud or deception, and AI-generated audio must be disclosed where required. Full terms are in Section II-A of the LICENSE.

License

Boson Higgs TTS 3 Research and Non-Commercial License โ€” see LICENSE. Includes a Creator Use Grant (free monetized creator use with attribution; see Creator Use above).

Have a use case that isnโ€™t covered by the current license? Weโ€™d still love to hear from you. Reach out to contact@boson.ai โ€” weโ€™re open to discussing your use case and alternative licensing arrangements.

Contributors

SilinMeng0510

12 commits

muli

4 commits

bytebecky

3 commits

linyueqian

1 commits

bosonai/higgs-tts-3-4b

Model

755

stars

22

commits

5

repos using this model

3

linked in READMEs

Sep 4, 2026

updated

ast
ceb
ckb
controllable-tts
endpoints_compatible
expressive-speech
higgs_multimodal_qwen3
kab
kam
kea
luo
mhr
multilingual-tts
nso
safetensors
speech-generation
text-generation
text-to-speech
transformers
umb
voice-agent

README

Higgs TTS 3

Higgs TTS 3 is built for voice chat: it speaks, not just reads. It turns model responses into expressive conversational speech across 100+ languages, with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.

[!TIP] Released for research and non-commercial use under the Boson Higgs TTS 3 Research and Non-Commercial License. Production, hosted APIs, embedding in a product/service, or reselling the model requires a separate commercial license. Prohibited: voice cloning without consent, impersonation, fraud, election deception, biometric surveillance, or any unlawful use.

[!TIP] Free for digital creators โ€” including monetized content. Under the license's Creator Use Grant, creators may use Higgs TTS 3 to make and monetize podcasts, videos, and social posts for free. The one requirement is to credit Boson AI's Higgs Audio โ€” either in the audio or prominently in the accompanying text (e.g., the video description or show notes). Suggested credit: "This audio was created with Boson AI's Higgs Audio โ€” https://www.boson.ai/higgs-audio". See the Creator Use section below.

Higgs TTS 3 Architecture

Higgs autoregressive decoder consumes interleaved text and audio tokens. Audio is encoded by the Higgs Tokenizer into 8 codebooks at 25 fps, staggered via a delay pattern, then mapped to backbone hidden states through a multi-codebook fused embedding. Output codes pass through a multi-codebook fused head, are de-delayed, and decoded back to waveform.

ComponentSpec
Backbone~4B autoregressive decoder (36 L, hidden=2560, GQA 32/8)
Multi-codebook embedding / headFused single-tensor, tied with text embedding
Context length8,192 tokens (training sequence length)
Audio tokens8 codebooks ร— 1026 vocab, delay pattern
Sample rate24 kHz
Frame rate25 fps (40 ms / frame)

Supported Languages

The model reaches single-digit WER/CER on 102 languages, which split into two tiers.

WER/CER under 5 โ€” polished, production-quality (85)

๐Ÿ‡ฟ๐Ÿ‡ฆ Afrikaans ยท ๐Ÿ‡ธ๐Ÿ‡ฆ๐Ÿ‡ช๐Ÿ‡ฌ Arabic ยท ๐Ÿ‡ฆ๐Ÿ‡ฒ Armenian ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Assamese ยท ๐Ÿ‡ช๐Ÿ‡ธ Asturian ยท ๐Ÿ‡ฆ๐Ÿ‡ฟ Azerbaijani ยท ๐Ÿ‡ท๐Ÿ‡บ Bashkir ยท ๐Ÿ‡ช๐Ÿ‡ธ Basque ยท ๐Ÿ‡ง๐Ÿ‡พ Belarusian ยท ๐Ÿ‡ง๐Ÿ‡ฉ๐Ÿ‡ฎ๐Ÿ‡ณ Bengali ยท ๐Ÿ‡ง๐Ÿ‡ฆ Bosnian ยท ๐Ÿ‡ง๐Ÿ‡ฌ Bulgarian ยท ๐Ÿ‡ช๐Ÿ‡ธ Catalan ยท ๐Ÿ‡ต๐Ÿ‡ญ Cebuano ยท ๐Ÿ‡ฎ๐Ÿ‡ถ Central Kurdish ยท ๐Ÿ‡จ๐Ÿ‡ณ Chinese ยท ๐Ÿ‡ญ๐Ÿ‡ท Croatian ยท ๐Ÿ‡จ๐Ÿ‡ฟ Czech ยท ๐Ÿ‡ฉ๐Ÿ‡ฐ Danish ยท ๐Ÿ‡ณ๐Ÿ‡ฑ๐Ÿ‡ง๐Ÿ‡ช Dutch ยท ๐Ÿ‡ท๐Ÿ‡บ Eastern Mari ยท ๐Ÿ‡บ๐Ÿ‡ธ๐Ÿ‡ฌ๐Ÿ‡ง๐Ÿ‡ฆ๐Ÿ‡บ English ยท ๐ŸŒ Esperanto ยท ๐Ÿ‡ช๐Ÿ‡ช Estonian ยท ๐Ÿ‡ซ๐Ÿ‡ฎ Finnish ยท ๐Ÿ‡ซ๐Ÿ‡ท๐Ÿ‡จ๐Ÿ‡ฆ French ยท ๐Ÿ‡ช๐Ÿ‡ธ Galician ยท ๐Ÿ‡ฌ๐Ÿ‡ช Georgian ยท ๐Ÿ‡ฉ๐Ÿ‡ช๐Ÿ‡ฆ๐Ÿ‡น German ยท ๐Ÿ‡ฌ๐Ÿ‡ท Greek ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Gujarati ยท ๐Ÿ‡ญ๐Ÿ‡น Haitian Creole ยท ๐Ÿ‡ณ๐Ÿ‡ฌ Hausa ยท ๐Ÿ‡ฎ๐Ÿ‡ฑ Hebrew ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Hindi ยท ๐Ÿ‡ญ๐Ÿ‡บ Hungarian ยท ๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesian ยท ๐Ÿ‡ฎ๐Ÿ‡น Italian ยท ๐Ÿ‡ฏ๐Ÿ‡ต Japanese ยท ๐Ÿ‡ฎ๐Ÿ‡ฉ Javanese ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Kannada ยท ๐Ÿ‡ฐ๐Ÿ‡ฟ Kazakh ยท ๐Ÿ‡ฐ๐Ÿ‡ท Korean ยท ๐Ÿ‡ท๐Ÿ‡ผ Kinyarwanda ยท ๐Ÿ‡ฐ๐Ÿ‡ฌ Kyrgyz ยท ๐Ÿ‡ฑ๐Ÿ‡ป Latvian ยท ๐Ÿ‡จ๐Ÿ‡ฉ Lingala ยท ๐Ÿ‡ฑ๐Ÿ‡น Lithuanian ยท ๐Ÿ‡ฐ๐Ÿ‡ช Luo ยท ๐Ÿ‡ฒ๐Ÿ‡ฐ Macedonian ยท ๐Ÿ‡ฒ๐Ÿ‡พ๐Ÿ‡ฎ๐Ÿ‡ฉ Malay ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Malayalam ยท ๐Ÿ‡ฒ๐Ÿ‡น Maltese ยท ๐Ÿ‡ณ๐Ÿ‡ฟ Mฤori ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Marathi ยท ๐Ÿ‡ฒ๐Ÿ‡ณ Mongolian ยท ๐Ÿ‡ณ๐Ÿ‡ต Nepali ยท ๐Ÿ‡ณ๐Ÿ‡ด Norwegian ยท ๐Ÿ‡ซ๐Ÿ‡ท Occitan ยท ๐Ÿ‡ฎ๐Ÿ‡ท๐Ÿ‡ฆ๐Ÿ‡ซ Persian ยท ๐Ÿ‡ต๐Ÿ‡ฑ Polish ยท ๐Ÿ‡ต๐Ÿ‡น๐Ÿ‡ง๐Ÿ‡ท Portuguese ยท ๐Ÿ‡ท๐Ÿ‡ด Romanian ยท ๐Ÿ‡ท๐Ÿ‡บ Russian ยท ๐Ÿ‡ฟ๐Ÿ‡ฆ Sepedi ยท ๐Ÿ‡ท๐Ÿ‡ธ Serbian ยท ๐Ÿ‡ฟ๐Ÿ‡ผ Shona ยท ๐Ÿ‡ธ๐Ÿ‡ฐ Slovak ยท ๐Ÿ‡ธ๐Ÿ‡ฎ Slovene ยท ๐Ÿ‡ช๐Ÿ‡ธ๐Ÿ‡ฒ๐Ÿ‡ฝ Spanish ยท ๐Ÿ‡น๐Ÿ‡ฟ๐Ÿ‡ฐ๐Ÿ‡ช Swahili ยท ๐Ÿ‡ธ๐Ÿ‡ช Swedish ยท ๐Ÿ‡ต๐Ÿ‡ญ Tagalog ยท ๐Ÿ‡น๐Ÿ‡ฏ Tajik ยท ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ฑ๐Ÿ‡ฐ Tamil ยท ๐Ÿ‡ฎ๐Ÿ‡ณ Telugu ยท ๐Ÿ‡น๐Ÿ‡ญ Thai ยท ๐Ÿ‡น๐Ÿ‡ท Turkish ยท ๐Ÿ‡บ๐Ÿ‡ฆ Ukrainian ยท ๐Ÿ‡ต๐Ÿ‡ฐ๐Ÿ‡ฎ๐Ÿ‡ณ Urdu ยท ๐Ÿ‡จ๐Ÿ‡ณ Uyghur ยท ๐Ÿ‡บ๐Ÿ‡ฟ Uzbek ยท ๐Ÿ‡ป๐Ÿ‡ณ Vietnamese ยท ๐Ÿ‡ฟ๐Ÿ‡ฆ Xhosa ยท ๐Ÿ‡ฟ๐Ÿ‡ฆ Zulu

WER/CER between 5 and 10 โ€” usable, but less polished (17)

๐Ÿ‡ฆ๐Ÿ‡ฑ Albanian ยท ๐Ÿ‡ฒ๐Ÿ‡ผ๐Ÿ‡ฟ๐Ÿ‡ฒ Chichewa/Nyanja ยท ๐Ÿ‡ฎ๐Ÿ‡ณ๐Ÿ‡ต๐Ÿ‡ฐ Eastern Punjabi ยท ๐Ÿ‡บ๐Ÿ‡ฌ Ganda ยท ๐Ÿ‡ฎ๐Ÿ‡ธ Icelandic ยท ๐Ÿ‡ฎ๐Ÿ‡ช Irish ยท ๐Ÿ‡ฉ๐Ÿ‡ฟ Kabyle ยท ๐Ÿ‡จ๐Ÿ‡ป Kabuverdianu ยท ๐Ÿ‡ฐ๐Ÿ‡ช Kamba ยท ๐Ÿ‡ป๐Ÿ‡ฆ Latin ยท ๐Ÿ‡ฑ๐Ÿ‡บ Luxembourgish ยท ๐Ÿ‡ช๐Ÿ‡น๐Ÿ‡ฐ๐Ÿ‡ช Oromo ยท ๐Ÿ‡ฆ๐Ÿ‡ซ๐Ÿ‡ต๐Ÿ‡ฐ Pashto ยท ๐Ÿ‡ต๐Ÿ‡ฐ๐Ÿ‡ฎ๐Ÿ‡ณ Sindhi ยท ๐Ÿ‡ธ๐Ÿ‡ด Somali ยท ๐Ÿ‡ฆ๐Ÿ‡ด Umbundu ยท ๐Ÿ‡ฌ๐Ÿ‡ง Welsh

Control Tokens

All tags follow <|category:value|> syntax and can be inserted mid-utterance.

For how to place these tags when writing the target text (sentence-level vs. inline, sfx formatting, stacking, worked examples), see PROMPTING.md.

  • Emotion โ€” elation, amusement, enthusiasm, determination, pride, contentment, affection, relief, contemplation, confusion, surprise, awe, longing, arousal, anger, fear, disgust, bitterness, sadness, shame, helplessness
TokenDescription
<|emotion:elation|>Elation / joy
<|emotion:amusement|>Amusement / playful laughter
<|emotion:enthusiasm|>Enthusiasm / excitement
<|emotion:determination|>Determination / firmness
<|emotion:pride|>Pride / confidence
<|emotion:contentment|>Calm satisfaction
<|emotion:affection|>Warmth / affection
<|emotion:relief|>Relief
<|emotion:contemplation|>Thoughtful / reflective
<|emotion:confusion|>Confused
<|emotion:surprise|>Surprised
<|emotion:awe|>Awe / wonder
<|emotion:longing|>Longing / yearning
<|emotion:arousal|>Heightened desire
<|emotion:anger|>Anger
<|emotion:fear|>Fear
<|emotion:disgust|>Disgust
<|emotion:bitterness|>Bitterness
<|emotion:sadness|>Sadness
<|emotion:shame|>Shame
<|emotion:helplessness|>Helplessness
  • Style โ€” singing, shouting, whispering
TokenDescription
<|style:singing|>Singing
<|style:shouting|>Shouting / projected voice
<|style:whispering|>Whisper
  • Sound effects โ€” cough, laughter, crying, screaming, burping, humming, sigh, sniff, sneeze

Pair each token with the matching onomatopoeia immediately after it.

TokenDescriptionSuggested onomatopoeia
<|sfx:cough|>CoughAhem
<|sfx:laughter|>LaughterHaha / Hehe
<|sfx:crying|>CryingBoohoo / Sob
<|sfx:screaming|>ScreamingAhh / Aaah
<|sfx:burping|>BurpingBurp
<|sfx:humming|>HummingHmm / Mmm
<|sfx:sigh|>SighUh / Ahh
<|sfx:sniff|>SniffSff
<|sfx:sneeze|>SneezeAchoo
  • Prosody
    • Speed โ€” speed_very_slow, speed_slow, speed_fast, speed_very_fast
    • Pauses โ€” pause, long_pause
    • Pitch โ€” pitch_low, pitch_high
    • Delivery โ€” expressive_high, expressive_low
TokenEffect
<|prosody:speed_very_slow|>โ‰ˆ0.65ร— speed
<|prosody:speed_slow|>โ‰ˆ0.85ร— speed
<|prosody:speed_fast|>โ‰ˆ1.2ร— speed
<|prosody:speed_very_fast|>โ‰ˆ1.4ร— speed
<|prosody:pitch_low|>โ‰ˆโˆ’3 semitones
<|prosody:pitch_high|>โ‰ˆ+2.5 semitones
<|prosody:pause|>โ‰ˆ400โ€“700 ms pause
<|prosody:long_pause|>โ‰ˆ700โ€“1500 ms pause
<|prosody:expressive_high|>More expressive delivery
<|prosody:expressive_low|>Flatter delivery

Evaluation Benchmarks

Multilingual Voice Clone

We evaluate Higgs TTS 3 on public multilingual TTS suites and our internal 111-language Higgs-Multilingual set, covering both common and lower-resource languages.

WER / CER (โ†“, ร—100) macro-averaged across each benchmark's language set. Lower is better; bold marks the best per row. All numbers are reproducible end-to-end with original metrics and normalization.

BenchmarkHiggs TTS v2Higgs TTS 3Fish Audio S2 ProQwen3-TTS-1.7BVibeVoice-7BIndexTTS-2MiMo-Audio-7B-InstructMOSS-TTS-v1.5OmniVoiceChatterBoxFireRedTTS-2
SeedTTS2.101.111.311.303.591.633.701.731.2117.001.72
CV321.194.414.607.7311.66129.2671.556.114.9232.6219.20
MiniMax-Multilingual49.862.745.1527.418.21112.9185.673.782.9849.3012.52
Higgs-Multilingual52.243.618.6897.0913.7457.7159.6121.283.6357.5233.69

Emergent TTS

Win-rate (โ†‘) per category โ€” judge preference vs the BASELINE row; bold marks the highest win-rate per column. For a fair comparison, every model shares the same reference audio per prompt, and we run the benchmark text verbatim โ€” no inline control tags inserted.

ModelOverall โ†‘Emotions โ†‘Foreign Words โ†‘Paralinguistics โ†‘Complex Pronunciation โ†‘Questions โ†‘Syntactic Complexity โ†‘
Higgs TTS 353.65%53.75%48.75%68.57%25.10%61.43%60.71%
Fish Audio S2 Pro43.80%53.04%33.93%53.75%18.16%55.00%45.71%
Qwen3-TTS-1.7B38.84%45.54%24.64%44.29%30.00%53.39%34.11%
IndexTTS-231.12%39.29%5.36%42.50%12.45%45.89%38.93%
MOSS-TTS-v1.543.89%60.54%35.18%51.43%11.63%53.21%47.32%
OmniVoice40.82%61.07%28.75%52.68%13.67%45.00%40.36%

Usage

SGLang Usage

Pair the weights in this repo with SGLang-Omni โ€” a production serving stack with continuous batching for multi-codebook decoding and the same inline tag controls. The Higgs TTS cookbook walks you through installation, server launch, request examples, and the full API reference.

See the Higgs TTS cookbook for the full details.

Install and Serve

docker pull lmsysorg/sglang-omni:dev
docker run -it --gpus all --shm-size 32g --ipc host --network host --privileged \
  lmsysorg/sglang-omni:dev /bin/zsh

git clone git@github.com:sgl-project/sglang-omni.git && cd sglang-omni
uv venv .venv -p 3.12 && source .venv/bin/activate
uv pip install -v -e .
export HF_TOKEN=hf_xxxxxxxxxxxxxxxx
hf download bosonai/higgs-tts-3-4b

sgl-omni serve \
  --model-path bosonai/higgs-tts-3-4b \
  --port 8000

Zero-shot synthesis

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"input": "Hello, how are you?"}' \
  --output output.wav

Voice cloning

Supplying the reference transcript (text) materially improves cloning fidelity.

import requests

resp = requests.post(
    "http://localhost:8000/v1/audio/speech",
    json={
        "input": "Have a nice day and enjoy south california sunshine.",
        "references": [{
            "audio_path": "ref.wav",
            "text": "Hey, Adam here. Let's create something that feels real, sounds human, and connects every time.",
        }],
        "temperature": 0.8, "top_k": 50, "max_new_tokens": 1024,
    },
)
with open("output.wav", "wb") as f:
    f.write(resp.content)

Streaming (Server-Sent Events)

Set "stream": true to receive base64-encoded WAV chunks as the vocoder emits them โ€” sub-second time-to-first-audio. Each event carries audio.data (base64 WAV bytes); the terminal event has finish_reason: "stop" plus usage metadata.

import requests, base64, json

with requests.post(
    "http://localhost:8000/v1/audio/speech",
    json={"input": "Get the trust fund to the bank early.", "stream": True},
    stream=True,
) as resp, open("output.wav", "wb") as f:
    for line in resp.iter_lines():
        if not line or not line.startswith(b"data: ") or line == b"data: [DONE]":
            continue
        event = json.loads(line[6:])
        if event.get("finish_reason") == "stop":
            break
        audio = event.get("audio") or {}
        if audio.get("data"):
            f.write(base64.b64decode(audio["data"]))

Inline control tokens

Embed <|emotion:โ€ฆ|>, <|style:โ€ฆ|>, <|prosody:โ€ฆ|>, and <|sfx:โ€ฆ|> tokens directly in input. Two rules:

  1. Delivery tokens first. Emotion, style, and the prosody speed / pitch / expressive tokens shape the whole turn โ€” put them at the start of input. Positional tokens (<|prosody:pause|>, <|prosody:long_pause|>, <|sfx:โ€ฆ|>) go inline exactly where they fire.
  2. Pair every <|sfx:โ€ฆ|> with its onomatopoeia. E.g. <|sfx:laughter|>Haha, <|sfx:sigh|>Uh, <|sfx:sneeze|>Achoo. The written sound gives the model the acoustic cue to realize the effect.

Example โ€” amusement + laughter:

curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"input": "<|emotion:amusement|><|prosody:expressive_high|>Wait, wait, that was kind of hilarious. <|sfx:laughter|>Hehe, no, seriously, I was not ready for that."}' \
  --output output.wav

Throughput

Throughput on Seed-TTS EN (full set, N=1088 per run). Client --max-concurrency sweep against a Higgs server (max_running_requests=16, bf16, CUDA Graph on). Each row is the mean of 3 runs. Hardware: 1ร— H100.

ConcurrencyThroughput (req/s)Mean latencyRTF (per-req)audio_s/s
11.62617 ms0.1476.89
22.70742 ms0.18011.37
45.45733 ms0.17722.84
88.91898 ms0.21737.38
1614.741079 ms0.26261.84
  • Concurrency โ€” Maximum number of in-flight client requests (--max-concurrency).
  • Throughput (req/s) โ€” Completed requests divided by total benchmark wall-clock time.
  • Mean latency โ€” Average end-to-end time per request (send to full response received).
  • RTF (per-req) โ€” Average ratio of processing time to generated audio duration per request (<1 is faster than real time).
  • audio_s/s โ€” Total seconds of audio produced divided by total benchmark wall-clock time.

To reproduce the results, follow the instructions in this script.

vLLM-Omni Usage

You can also serve these weights with vLLM-Omni, which exposes the same OpenAI-compatible /v1/audio/speech API with zero-shot voice cloning.

hf download bosonai/higgs-tts-3-4b

vllm-omni serve bosonai/higgs-tts-3-4b \
  --host 0.0.0.0 --port 8095 \
  --trust-remote-code --omni
curl -X POST http://localhost:8095/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model": "bosonai/higgs-tts-3-4b", "input": "Hello, how are you?"}' \
  --output output.wav

Plain-text TTS, voice-clone, and benchmark recipes are in the vLLM-Omni Higgs TTS 3 recipe.

API Usage

For zero-ops deployment, use the Boson AI API.

Citation

@misc{bosonai_higgs_audio_tts_v3_2026,
  title  = {Higgs TTS 3: Conversational Speech for Voice AI from Boson AI},
  author = {Boson AI},
  year   = {2026},
  howpublished = {https://huggingface.co/bosonai/higgs-tts-3-4b},
}

Creator Use (Free for Digital Creators)

In addition to research and non-commercial use, the license includes a Creator Use Grant that lets digital creators use Higgs TTS 3 to produce creative content for free โ€” including monetized content.

What's covered

  • Podcasts, videos, audiobooks, social media posts, and similar creative works
  • Personal and commercial/monetized creator channels (ad-supported, sponsored, subscription, etc.)

The one requirement: acknowledge Boson AI's Higgs Audio. The acknowledgment must appear in at least one of the following ways:

  • In the audio โ€” e.g. "This audio was created with Boson AI's Higgs Audio."
  • In the accompanying text, displayed prominently โ€” e.g. in the post body, video description, or show notes. It must be clearly visible and not hidden at the bottom of the credits or annotations.

Suggested credit string:

This audio was created with Boson AI's Higgs Audio โ€” https://www.boson.ai/higgs-audio

Still requires a separate commercial license. The Creator Use Grant covers creating content with the model. It does not cover hosting the model behind an API or as a service, redistributing/reselling/fine-tuning the model for resale, or embedding the model in a product or application. For these uses, contact us for a commercial license.

The Creator Use Grant does not change the use restrictions โ€” no non-consensual voice cloning or impersonation, no fraud or deception, and AI-generated audio must be disclosed where required. Full terms are in Section II-A of the LICENSE.

License

Boson Higgs TTS 3 Research and Non-Commercial License โ€” see LICENSE. Includes a Creator Use Grant (free monetized creator use with attribution; see Creator Use above).

Have a use case that isnโ€™t covered by the current license? Weโ€™d still love to hear from you. Reach out to contact@boson.ai โ€” weโ€™re open to discussing your use case and alternative licensing arrangements.

Contributors

SilinMeng0510

12 commits

muli

4 commits

bytebecky

3 commits

linyueqian

1 commits