15 repos
Quantized implementations of the Moshi real-time audio model across different precision formats and inference frameworks. The cluster focuses on model compression techniques—particularly bfloat16 and int8 quantization—applied to Moshi's audio-to-audio capabilities, with support for multiple backends including Candle and MLX. Repositories here enable efficient deployment of conversational audio models on resource-constrained hardware while maintaining quality.