(tránh chép đè file conds.pt - đây là file config để voice mẫu hiện tại chạy chính xác) https://huggingface.co/dolly-vn/viterbox/tree/main
(model này nằm trong pipeline của Omnivoice, bắt buộc phải có) https://huggingface.co/eustlb/higgs-audio-v2-tokenizer/tree/main
(model này dùng để detect text cho âm thanh đầu vào - nếu không có text. lý do cần text cho audio mẫu là vì khi TTS bằng omniVoice sẽ chính xác hơn) (model này nằm trong pipeline của Omnivoice, bắt buộc phải có) https://huggingface.co/khanhld/chunkformer-ctc-large-vie/tree/main
https://huggingface.co/kjanh/KhanhTTS-OmniVoice
https://huggingface.co/k2-fsa/OmniVoice/tree/main
viterbox-TTS=GPU/
├── app.py # Gradio Web UI
├── inference.py # CLI inference script
└── general/ # Core library
├── general/requirements.txt # Dependencies (Windows/Linux)
├── general/requirements-mac.txt# Dependencies (macOS)
├── config_path.txt # lưu đường dẫn folder download audio
└── EQ_emotion_config/ # chứa các file config âm thanh bằng EQ
├── pyproject.toml # Package config
├── README.md
├── wavs/ # Thư mục chứa giọng mẫu
│ └── *.wav
├── OmniVoice/ # folder với model OmniVoice + file inference
│ ├── modelOmniLocal/ # Thư mục chứa model local OmniVoice
│ ├── omnivoice/ # Model components OmniVoice
│ └── omnivoice_inference/# Folder chứa phần suy luận của OmniVoice
│ └── ttsOmni.py # File suy luận cho Omnivoice
└── viterbox/ # Core library
├── modelViterboxLocal/ # Thư mục chứa model local Viterbox(base trên Chatterbox)
├── output-profile/ # Thư mục chứa file kết quả của Voice Profile
├── pretrained/ # Thư mục chứa audio + text cho Voice Profile
├── __init__.py
├── tts.py # Main Viterbox class
└── models/ # Model components
├── t3/ # T3 Text-to-Token model
├── s3gen/ # S3Gen vocoder
├── s3tokenizer/ # Speech tokenizer
├── voice_encoder/ # Speaker encoder
└── tokenizers/ # Text tokenizer
# Clone repo
git clone https://github.com/nowtranminh1-TTS/BetterBox-TTS.git
# vào thư mục viterbox
cd viterbox
# Tạo virtual environment (khuyến nghị) - tạo trong thư mục viterbox
python -m venv venv
# bật venv lên - bắt buộc để cài được lib
source venv/bin/activate # Linux/Mac
# hoặc: venv\Scripts\activate # Windows
# back ra ngoài
cd ..
# vào thư mục general - để cài các lib có trong file 'requirements.txt'
cd general
# Cài đặt dependencies
pip install -r requirements.txt
# sau khi đã cài venv + download model về local. sau này chỉ cần click file 'runApp.bat' - file tự động bật venv và chạy
pip install -e .
python app.py
Mở trình duyệt tại http://localhost:7860 hoặc http://127.0.0.1:7860
hoặc sau khi có venv, thì chạy file 'runApp.bat' - file tự động bật venv và chạy
| Tham số | Mô tả | Giá trị | Mặc định |
|---|---|---|---|
text | Văn bản cần đọc | string | (bắt buộc) |
language | Mã ngôn ngữ | "vi", "en" | "vi" |
audio_prompt | Audio mẫu cho voice cloning | path/tensor | None |
exaggeration | Mức độ biểu cảm | 0.0 - 2.0 | 0.5 |
cfg_weight | Độ bám sát giọng mẫu | 0.0 - 1.0 | 0.5 |
temperature | Độ ngẫu nhiên/sáng tạo | 0.1 - 1.0 | 0.8 |
top_p | Top-p sampling | 0.0 - 1.0 | 0.9 |
repetition_penalty | Phạt lặp từ | 1.0 - 2.0 | 1.2 |
sentence_pause_ms | Thời gian ngắt giữa câu | 0 - 2000 | 500 |
crossfade_ms | Thời gian crossfade | 0 - 100 | 50 |
Áp dụng cho dữ liệu trong folder viterbox/pretrained/:
clip1.mp3 + clip1.txt.speaker_emb và x-vector được tính từ toàn bộ audio (không cắt 80s).conds.pt trong viterbox/output-profile/.Copy -> modelViterboxLocal để app dùng ngay (cần restart app), file sẽ được copy vào viterbox/modelViterboxLocal/.CC BY-NC 4.0 (Creative Commons Attribution-NonCommercial 4.0)
Python
95.6%
Shell
3.3%
(tránh chép đè file conds.pt - đây là file config để voice mẫu hiện tại chạy chính xác) https://huggingface.co/dolly-vn/viterbox/tree/main
(model này nằm trong pipeline của Omnivoice, bắt buộc phải có) https://huggingface.co/eustlb/higgs-audio-v2-tokenizer/tree/main
(model này dùng để detect text cho âm thanh đầu vào - nếu không có text. lý do cần text cho audio mẫu là vì khi TTS bằng omniVoice sẽ chính xác hơn) (model này nằm trong pipeline của Omnivoice, bắt buộc phải có) https://huggingface.co/khanhld/chunkformer-ctc-large-vie/tree/main
https://huggingface.co/kjanh/KhanhTTS-OmniVoice
https://huggingface.co/k2-fsa/OmniVoice/tree/main
viterbox-TTS=GPU/
├── app.py # Gradio Web UI
├── inference.py # CLI inference script
└── general/ # Core library
├── general/requirements.txt # Dependencies (Windows/Linux)
├── general/requirements-mac.txt# Dependencies (macOS)
├── config_path.txt # lưu đường dẫn folder download audio
└── EQ_emotion_config/ # chứa các file config âm thanh bằng EQ
├── pyproject.toml # Package config
├── README.md
├── wavs/ # Thư mục chứa giọng mẫu
│ └── *.wav
├── OmniVoice/ # folder với model OmniVoice + file inference
│ ├── modelOmniLocal/ # Thư mục chứa model local OmniVoice
│ ├── omnivoice/ # Model components OmniVoice
│ └── omnivoice_inference/# Folder chứa phần suy luận của OmniVoice
│ └── ttsOmni.py # File suy luận cho Omnivoice
└── viterbox/ # Core library
├── modelViterboxLocal/ # Thư mục chứa model local Viterbox(base trên Chatterbox)
├── output-profile/ # Thư mục chứa file kết quả của Voice Profile
├── pretrained/ # Thư mục chứa audio + text cho Voice Profile
├── __init__.py
├── tts.py # Main Viterbox class
└── models/ # Model components
├── t3/ # T3 Text-to-Token model
├── s3gen/ # S3Gen vocoder
├── s3tokenizer/ # Speech tokenizer
├── voice_encoder/ # Speaker encoder
└── tokenizers/ # Text tokenizer
# Clone repo
git clone https://github.com/nowtranminh1-TTS/BetterBox-TTS.git
# vào thư mục viterbox
cd viterbox
# Tạo virtual environment (khuyến nghị) - tạo trong thư mục viterbox
python -m venv venv
# bật venv lên - bắt buộc để cài được lib
source venv/bin/activate # Linux/Mac
# hoặc: venv\Scripts\activate # Windows
# back ra ngoài
cd ..
# vào thư mục general - để cài các lib có trong file 'requirements.txt'
cd general
# Cài đặt dependencies
pip install -r requirements.txt
# sau khi đã cài venv + download model về local. sau này chỉ cần click file 'runApp.bat' - file tự động bật venv và chạy
pip install -e .
python app.py
Mở trình duyệt tại http://localhost:7860 hoặc http://127.0.0.1:7860
hoặc sau khi có venv, thì chạy file 'runApp.bat' - file tự động bật venv và chạy
| Tham số | Mô tả | Giá trị | Mặc định |
|---|---|---|---|
text | Văn bản cần đọc | string | (bắt buộc) |
language | Mã ngôn ngữ | "vi", "en" | "vi" |
audio_prompt | Audio mẫu cho voice cloning | path/tensor | None |
exaggeration | Mức độ biểu cảm | 0.0 - 2.0 | 0.5 |
cfg_weight | Độ bám sát giọng mẫu | 0.0 - 1.0 | 0.5 |
temperature | Độ ngẫu nhiên/sáng tạo | 0.1 - 1.0 | 0.8 |
top_p | Top-p sampling | 0.0 - 1.0 | 0.9 |
repetition_penalty | Phạt lặp từ | 1.0 - 2.0 | 1.2 |
sentence_pause_ms | Thời gian ngắt giữa câu | 0 - 2000 | 500 |
crossfade_ms | Thời gian crossfade | 0 - 100 | 50 |
Áp dụng cho dữ liệu trong folder viterbox/pretrained/:
clip1.mp3 + clip1.txt.speaker_emb và x-vector được tính từ toàn bộ audio (không cắt 80s).conds.pt trong viterbox/output-profile/.Copy -> modelViterboxLocal để app dùng ngay (cần restart app), file sẽ được copy vào viterbox/modelViterboxLocal/.CC BY-NC 4.0 (Creative Commons Attribution-NonCommercial 4.0)
Python
95.6%
Shell
3.3%