Deep-unlearning/smol-audio

Practical, Colab-friendly notebooks for fine-tuning and running audio AI models

Jupyter Notebook

425

19 commits

updated Jul 22, 2026

See the code

README

Smol Audio

Smol Audio 🔊

Practical notebooks for shrinking, optimizing, and customizing audio AI models with the Hugging Face ecosystem.

Latest examples

  • Inference with Perception Encoder for Audio-Video (PE-AV)
  • Fine-tune Audio Flamingo 3
  • Granite Speech 4.0 1b ASR

[!NOTE] GitHub doesn't always render notebooks well. If you have trouble viewing them, try opening in Colab using the links below.

CategoryNotebookDescription
ASR Fine-tuningFine-tune WhisperFine-tune Whisper on a custom language/domain using transformers + datasets
ASR Fine-tuningFine-tune Granite Speech ItalianFine-tune IBM Granite Speech for Italian ASR with the YODAS-Granary dataset
Audio CaptioningFine-tune Audio Flamingo 3Fine-tune Audio Flamingo 3 for audio captioning (full + LoRA)
ASR Fine-tuningFine-tune ParakeetFine-tune NVIDIA Parakeet CTC for speech recognition (full + LoRA)
ASR Fine-tuningFine-tune Voxtral ASRFine-tune Voxtral for ASR with prompt masking (full + LoRA)
MultimodalInference with PE-AV-BaseZero-shot video classification and audio↔text retrieval (AudioCaps) with Meta's Perception Encoder for Audio-Video

Contributors

Deep-unlearning

19 commits

Deep-unlearning/smol-audio

Practical, Colab-friendly notebooks for fine-tuning and running audio AI models

Jupyter Notebook

425

19 commits

updated Jul 22, 2026

See the code

README

Smol Audio

Smol Audio 🔊

Practical notebooks for shrinking, optimizing, and customizing audio AI models with the Hugging Face ecosystem.

Latest examples

  • Inference with Perception Encoder for Audio-Video (PE-AV)
  • Fine-tune Audio Flamingo 3
  • Granite Speech 4.0 1b ASR

[!NOTE] GitHub doesn't always render notebooks well. If you have trouble viewing them, try opening in Colab using the links below.

CategoryNotebookDescription
ASR Fine-tuningFine-tune WhisperFine-tune Whisper on a custom language/domain using transformers + datasets
ASR Fine-tuningFine-tune Granite Speech ItalianFine-tune IBM Granite Speech for Italian ASR with the YODAS-Granary dataset
Audio CaptioningFine-tune Audio Flamingo 3Fine-tune Audio Flamingo 3 for audio captioning (full + LoRA)
ASR Fine-tuningFine-tune ParakeetFine-tune NVIDIA Parakeet CTC for speech recognition (full + LoRA)
ASR Fine-tuningFine-tune Voxtral ASRFine-tune Voxtral for ASR with prompt masking (full + LoRA)
MultimodalInference with PE-AV-BaseZero-shot video classification and audio↔text retrieval (AudioCaps) with Meta's Perception Encoder for Audio-Video

Contributors

Deep-unlearning

19 commits

Languages

Jupyter Notebook

90.7%

Python

9.3%