Speech Tokenization and Audio Encoding

4 repos

Tools and models for converting continuous speech and audio signals into discrete token sequences, enabling downstream processing by language models and other discrete-input systems. The cluster includes reference implementations like Kanade (at multiple frame rates) and SpeechTokenizer, along with supporting infrastructure for tokenization pipelines and ONNX model export. This represents a key preprocessing step in multimodal AI systems that need to bridge between continuous audio data and discrete token-based architectures.

Python · 1
onnx ·53
EMOVASpeechTokenizer ·2
en ·2
feature-extraction ·2
safetensors ·2
speech ·2
tokenization ·2
transformers ·2
custom_code ·2
zh ·2