k2-fsa/OmniVoice

Model

1,366

stars

21

commits

8

repos using this model

9

linked in READMEs

Jul 3, 2026

updated

aae
aal
aao
abb
abn
abr
abs
abv
acm
acw
acx
adf
adx
ady
aeb
aec
afb
afo
ahl
ahs
ajg
aju
ala
aln
alo
amu
anc
ank
anp
anw
aom
apc
apd
arb
arq
ars
ary
arz
ast
avl
awo
ayl
ayp
bag
bas
bax
bba
bbj
bbl
bbu
bce
bci
bcs
bcy
bda
bde
bdm
beb
bew
bfd
bft
bgp
bhb
bhh
bho
bhp
bhr
bjj
bjk
bjn
bjt
bkh
bkm
bky
bmm
bmq
bnm
bnn
bns
bou
bqg
bra
brh
bri
brx
bsh
bsj
bsk
btm
btv
bug
bum
buo
bux
bwr
bxf
byc
bys
byv
byx
bzc
bzw
ccg
ceb
cen
cfa
cgg
chq
cjk
ckb
ckl
ckr
cky
cnh
cpy
cte
ctl
cut
cux
dag
dar
dav
dbd
dcc
deg
dgh
dgo
dje
dmk
dml
dru
dty
dua
dyu
dzg
ebr
ebu
ego
eiv
eko
ekr
elm
esu
eto
ets
etu
ewo
ext
eyo
fan
fat
ffm
fia
fil
fip
fkk
fmp
fub
fuc
fue
fuf
fuh
fui
fuq
fuv
gbm
gbr
gby
gcc
gdf
gej
ges
ggg
gid
gig
giz
gjk
gju
glw
gol
gom
gsl
gui
gur
guz
gwc
gwe
gwt
gya
gyz
hah
hao
haw
haz
hbb
hem
hia
hkk
hla
hno
hoj
hsb
hue
hul
hux
hwo
ibb
ida
idu
ijc
ijn
ikw
ish
iso
its
itw
itz
jal
jax
jgo
jmx
jns
jqr
juk
juo
kab
kai
kaj
kam
kbd
kbl
kbt
kcq
kdh
kea
keu
kfe
kfk
kfp
khg
khw
kjc
kjk
kln
kls
kmr
kmy
kna
knn
kol
koo
kpo
kqo
ksd
ksf
kto
kuh
kvx
kwm
kxp
kyx
lag
lcm
ldb
lij
lir
lkb
lla
lnu
loa
lrk
lss
ltg
lto
lua
luo
lus
lwg
mab
maf
mai
mau
max
mbo
mcf
mcn
mcx
mdd
mde
mdf
mek
mer
meu
mfm
mfn
mfo
mfv
mgg
mgi
mhk
mhr
mig
miu
mkf
mki
mlq
mne
mni
mqy
mrj
mrr
mrt
mse
msh
msw
mtr
mtu
mtx
mua
mug
mui
multilingual
mve
mvy
mxs
mxu
mxy
myv
mzl
nal
nan
nap
nbh
ncf
nco
ncx
ndi
ngi
nhg
nhi
nhn
nhq
nja
nla
nlv
nmg
nmz
nnh
noe
npi
nso
nyu
odk
odu
ogo
omnivoice
orc
oru
ory
pbs
pbt
pbu
pcm
pex
phl
phr
pip
piy
pko
plk
plt
pmq
pms
pmy
pnb
poc
poe
pow
prq
pst
pua
pwn
qug
qum
qup
qur
qus
quv
qux
quy
qva
qvi
qvj
qvl
qwa
qws
qxa
qxp
qxt
qxu
qxw
rag
rob
rof
roo
rth
rup
safetensors
sah
sat
sau
say
sbn
scl
scn
sei
shu
sip
siw
sjr
skg
skr
snc
snk
sol
sps
src
sro
ssi
ste
sua
sva
szy
tan
tar
tay
tbf
tcf
tcy
tdn
tdx
text-to-speech
tgc
the
thq
thr
thv
tig
tio
tkg
tkt
tli
tlp
tok
tpl
tpz
tqp
trp
trq
trv
trw
ttj
ttr
ttu
tui
tul
tuq
tuv
tuy
tvo
tvu
twu
txs
txy
udl
uki
umb
ush
uzn
vai
var
ver
vmc
vmj
vmm
vmp
vmz
voice-cloning
voice-design
vot
vro
wbl
wci
weo
wes
wja
wji
wof
xhe
xka
xmf
xmv
xmw
xpe
xti
xtu
yaq
yav
yay
ydd
ydg
yer
yes
yue
zero-shot
zga
zgh
zoc
zoh
zor
zpv
zpy
ztg
ztn
ztp
zts
ztu
zza
Browse cluster: Multilingual Text-to-Speech Systems

README

OmniVoice 🌍

OmniVoice

Hugging Face Model   Hugging Face Space     GitHub Code     Open In Colab

OmniVoice is a massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model-style architecture, it delivers high-quality speech with superior inference speed, supporting voice cloning and voice design.

Key Features

  • 600+ Languages Supported: The broadest language coverage among zero-shot TTS models.
  • Voice Cloning: State-of-the-art voice cloning quality from a short reference audio.
  • Voice Design: Control voices via assigned speaker attributes (gender, age, pitch, dialect/accent, whisper, etc.).
  • Fine-grained Control: Non-verbal symbols (e.g., [laughter]) and pronunciation correction via pinyin or phonemes.
  • Fast Inference: RTF as low as 0.025 (40x faster than real-time).
  • Diffusion Language Model-style Architecture: A clean, streamlined, and scalable design that delivers both quality and speed.

Usage

To get started, install the omnivoice library:

We recommend using a fresh virtual environment (e.g., conda, venv, etc.) to avoid conflicts.

Step 1: Install PyTorch

NVIDIA GPU
# Install pytorch with your CUDA version, e.g.
pip install torch==2.8.0+cu128 torchaudio==2.8.0+cu128 --extra-index-url https://download.pytorch.org/whl/cu128

See PyTorch official site for other versions installation.

Apple Silicon
pip install torch==2.8.0 torchaudio==2.8.0

Step 2: Install OmniVoice

pip install omnivoice

Python API

You can use OmniVoice for zero-shot voice cloning as follows:

from omnivoice import OmniVoice
import soundfile as sf
import torch

# Load the model
model = OmniVoice.from_pretrained(
    "k2-fsa/OmniVoice",
    device_map="cuda:0",
    dtype=torch.float16
)

# Generate audio
audio = model.generate(
    text="Hello, this is a test of zero-shot voice cloning.",
    ref_audio="ref.wav",
    ref_text="Transcription of the reference audio.",
) # audio is a list of `np.ndarray` with shape (T,) at 24 kHz.

sf.write("out.wav", audio[0], 24000)

For more generation modes (e.g., voice design), functions (e.g., non-verbal symbols, pronunciation correction) and comprehensive usage instructions, see our GitHub Repository.

Discussion & Communication

You can directly discuss on GitHub Issues.

You can also scan the QR code to join our wechat group or follow our wechat official account.

Wechat GroupWechat Official Account
wechatwechat

Citation

@article{zhu2026omnivoice,
      title={OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models},
      author={Zhu, Han and Ye, Lingxuan and Kang, Wei and Yao, Zengwei and Guo, Liyong and Kuang, Fangjun and Han, Zhifeng and Zhuang, Weiji and Lin, Long and Povey, Daniel},
      journal={arXiv preprint arXiv:2604.00688},
      year={2026}
}

License

Our code is released under the Apache 2.0 License. The pre-trained model is licensed under the CC-BY-NC due to constraints from its training data (e.g., Emilia).

Disclaimer

Users are strictly prohibited from using this model for unauthorized voice cloning, voice impersonation, fraud, scams, or any other illegal or unethical activities. All users shall ensure full compliance with applicable local laws, regulations, and ethical standards. The developers assume no liability for any misuse of this model and advocate for responsible AI development and use, encouraging the community to uphold safety and ethical principles in AI research and applications.

Contributors

zhu-han

18 commits

MihaiPopa-1

1 commits

nielsr

1 commits

TahirC

1 commits

k2-fsa/OmniVoice

Model

1,366

stars

21

commits

8

repos using this model

9

linked in READMEs

Jul 3, 2026

updated

aae
aal
aao
abb
abn
abr
abs
abv
acm
acw
acx
adf
adx
ady
aeb
aec
afb
afo
ahl
ahs
ajg
aju
ala
aln
alo
amu
anc
ank
anp
anw
aom
apc
apd
arb
arq
ars
ary
arz
ast
avl
awo
ayl
ayp
bag
bas
bax
bba
bbj
bbl
bbu
bce
bci
bcs
bcy
bda
bde
bdm
beb
bew
bfd
bft
bgp
bhb
bhh
bho
bhp
bhr
bjj
bjk
bjn
bjt
bkh
bkm
bky
bmm
bmq
bnm
bnn
bns
bou
bqg
bra
brh
bri
brx
bsh
bsj
bsk
btm
btv
bug
bum
buo
bux
bwr
bxf
byc
bys
byv
byx
bzc
bzw
ccg
ceb
cen
cfa
cgg
chq
cjk
ckb
ckl
ckr
cky
cnh
cpy
cte
ctl
cut
cux
dag
dar
dav
dbd
dcc
deg
dgh
dgo
dje
dmk
dml
dru
dty
dua
dyu
dzg
ebr
ebu
ego
eiv
eko
ekr
elm
esu
eto
ets
etu
ewo
ext
eyo
fan
fat
ffm
fia
fil
fip
fkk
fmp
fub
fuc
fue
fuf
fuh
fui
fuq
fuv
gbm
gbr
gby
gcc
gdf
gej
ges
ggg
gid
gig
giz
gjk
gju
glw
gol
gom
gsl
gui
gur
guz
gwc
gwe
gwt
gya
gyz
hah
hao
haw
haz
hbb
hem
hia
hkk
hla
hno
hoj
hsb
hue
hul
hux
hwo
ibb
ida
idu
ijc
ijn
ikw
ish
iso
its
itw
itz
jal
jax
jgo
jmx
jns
jqr
juk
juo
kab
kai
kaj
kam
kbd
kbl
kbt
kcq
kdh
kea
keu
kfe
kfk
kfp
khg
khw
kjc
kjk
kln
kls
kmr
kmy
kna
knn
kol
koo
kpo
kqo
ksd
ksf
kto
kuh
kvx
kwm
kxp
kyx
lag
lcm
ldb
lij
lir
lkb
lla
lnu
loa
lrk
lss
ltg
lto
lua
luo
lus
lwg
mab
maf
mai
mau
max
mbo
mcf
mcn
mcx
mdd
mde
mdf
mek
mer
meu
mfm
mfn
mfo
mfv
mgg
mgi
mhk
mhr
mig
miu
mkf
mki
mlq
mne
mni
mqy
mrj
mrr
mrt
mse
msh
msw
mtr
mtu
mtx
mua
mug
mui
multilingual
mve
mvy
mxs
mxu
mxy
myv
mzl
nal
nan
nap
nbh
ncf
nco
ncx
ndi
ngi
nhg
nhi
nhn
nhq
nja
nla
nlv
nmg
nmz
nnh
noe
npi
nso
nyu
odk
odu
ogo
omnivoice
orc
oru
ory
pbs
pbt
pbu
pcm
pex
phl
phr
pip
piy
pko
plk
plt
pmq
pms
pmy
pnb
poc
poe
pow
prq
pst
pua
pwn
qug
qum
qup
qur
qus
quv
qux
quy
qva
qvi
qvj
qvl
qwa
qws
qxa
qxp
qxt
qxu
qxw
rag
rob
rof
roo
rth
rup
safetensors
sah
sat
sau
say
sbn
scl
scn
sei
shu
sip
siw
sjr
skg
skr
snc
snk
sol
sps
src
sro
ssi
ste
sua
sva
szy
tan
tar
tay
tbf
tcf
tcy
tdn
tdx
text-to-speech
tgc
the
thq
thr
thv
tig
tio
tkg
tkt
tli
tlp
tok
tpl
tpz
tqp
trp
trq
trv
trw
ttj
ttr
ttu
tui
tul
tuq
tuv
tuy
tvo
tvu
twu
txs
txy
udl
uki
umb
ush
uzn
vai
var
ver
vmc
vmj
vmm
vmp
vmz
voice-cloning
voice-design
vot
vro
wbl
wci
weo
wes
wja
wji
wof
xhe
xka
xmf
xmv
xmw
xpe
xti
xtu
yaq
yav
yay
ydd
ydg
yer
yes
yue
zero-shot
zga
zgh
zoc
zoh
zor
zpv
zpy
ztg
ztn
ztp
zts
ztu
zza
Browse cluster: Multilingual Text-to-Speech Systems

README

OmniVoice 🌍

OmniVoice

Hugging Face Model   Hugging Face Space     GitHub Code     Open In Colab

OmniVoice is a massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model-style architecture, it delivers high-quality speech with superior inference speed, supporting voice cloning and voice design.

Key Features

  • 600+ Languages Supported: The broadest language coverage among zero-shot TTS models.
  • Voice Cloning: State-of-the-art voice cloning quality from a short reference audio.
  • Voice Design: Control voices via assigned speaker attributes (gender, age, pitch, dialect/accent, whisper, etc.).
  • Fine-grained Control: Non-verbal symbols (e.g., [laughter]) and pronunciation correction via pinyin or phonemes.
  • Fast Inference: RTF as low as 0.025 (40x faster than real-time).
  • Diffusion Language Model-style Architecture: A clean, streamlined, and scalable design that delivers both quality and speed.

Usage

To get started, install the omnivoice library:

We recommend using a fresh virtual environment (e.g., conda, venv, etc.) to avoid conflicts.

Step 1: Install PyTorch

NVIDIA GPU
# Install pytorch with your CUDA version, e.g.
pip install torch==2.8.0+cu128 torchaudio==2.8.0+cu128 --extra-index-url https://download.pytorch.org/whl/cu128

See PyTorch official site for other versions installation.

Apple Silicon
pip install torch==2.8.0 torchaudio==2.8.0

Step 2: Install OmniVoice

pip install omnivoice

Python API

You can use OmniVoice for zero-shot voice cloning as follows:

from omnivoice import OmniVoice
import soundfile as sf
import torch

# Load the model
model = OmniVoice.from_pretrained(
    "k2-fsa/OmniVoice",
    device_map="cuda:0",
    dtype=torch.float16
)

# Generate audio
audio = model.generate(
    text="Hello, this is a test of zero-shot voice cloning.",
    ref_audio="ref.wav",
    ref_text="Transcription of the reference audio.",
) # audio is a list of `np.ndarray` with shape (T,) at 24 kHz.

sf.write("out.wav", audio[0], 24000)

For more generation modes (e.g., voice design), functions (e.g., non-verbal symbols, pronunciation correction) and comprehensive usage instructions, see our GitHub Repository.

Discussion & Communication

You can directly discuss on GitHub Issues.

You can also scan the QR code to join our wechat group or follow our wechat official account.

Wechat GroupWechat Official Account
wechatwechat

Citation

@article{zhu2026omnivoice,
      title={OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models},
      author={Zhu, Han and Ye, Lingxuan and Kang, Wei and Yao, Zengwei and Guo, Liyong and Kuang, Fangjun and Han, Zhifeng and Zhuang, Weiji and Lin, Long and Povey, Daniel},
      journal={arXiv preprint arXiv:2604.00688},
      year={2026}
}

License

Our code is released under the Apache 2.0 License. The pre-trained model is licensed under the CC-BY-NC due to constraints from its training data (e.g., Emilia).

Disclaimer

Users are strictly prohibited from using this model for unauthorized voice cloning, voice impersonation, fraud, scams, or any other illegal or unethical activities. All users shall ensure full compliance with applicable local laws, regulations, and ethical standards. The developers assume no liability for any misuse of this model and advocate for responsible AI development and use, encouraging the community to uphold safety and ethical principles in AI research and applications.

Contributors

zhu-han

18 commits

MihaiPopa-1

1 commits

nielsr

1 commits

TahirC

1 commits