Python
0
33 commits
updated Jun 16, 2025
This project provides scripts to fine-tune an OpenAI Whisper model on a custom dataset using Modal for GPU-accelerated training and then run inference locally using the fine-tuned checkpoint, with optional comparison against OpenAI and Azure transcription APIs.
openai/whisper-small by default) on a custom Arabic dataset (e.g., Egyptian dialect).train_whisper_modal.py: Script to define and run the fine-tuning process on Modal.infer_whisper_local.py: Script to load a fine-tuned checkpoint and perform transcription locally or compare with APIs..env: (To be created by user) Stores API keys securely.README.md: This file.pip install modal-client
modal setup
git clone <your-repository-url>
cd <repository-directory>
infer_whisper_local.py to compare with APIs:
Modal uses Secrets to securely store credentials. Create the following secrets via the Modal Secrets page or CLI:
modal secret create huggingface-secret-write HF_TOKEN=<your-huggingface-write-token>
modal secret create wandb-secret WANDB_API_KEY=<your-wandb-api-key>
.env file)Create a file named .env in the root of the project directory to store API keys for local inference (if needed). Do not commit this file to Git.
# .env file
OPENAI_API_KEY="sk-..."
AZURE_SPEECH_KEY="..."
AZURE_SERVICE_REGION="YourAzureRegion" # e.g., westeurope, eastus
infer_whisper_local.py will automatically load these variables.
train_whisper_modal.py)This script defines a Modal function to perform the fine-tuning process on a GPU instance in the cloud.
Before running, review and modify the placeholders and hparams dictionary within train_whisper_modal.py:
Placeholder comments near the modal.App and modal.Volume.from_name definitions and replace them with your desired names.hparams["hf_dataset_id"]: Change "huggingfaceusername/datasetname" to the Hugging Face dataset identifier for your custom dataset (e.g., "MAdel121/arabic-egy-cleaned"). Ensure the dataset has 'audio' and 'text' columns. The script assumes standard splits like 'train', 'validation', 'test'.hparams["whisper_hub"]: Change "openai/whisper-small" if you want to fine-tune a different Whisper variant (e.g., openai/whisper-base, openai/whisper-medium). Ensure this matches the base model used in infer_whisper_local.py.hparams["language"]: Set the target language code (e.g., "ar", "en").hparams["task"]: Usually "transcribe".hparams["save_folder"], hparams["output_folder"]: Define paths within the Modal volume where checkpoints and logs will be saved. Usually, the defaults are fine.hparams["augment"]: True or False to enable/disable augmentation.hparams["augment_prob_master"]: Overall probability of applying augmentation per batch.hparams["use_*"] flags: Toggle specific augmentations (AddNoise, AddReverb, SpeedPerturb, etc.).hparams["epochs"], hparams["learning_rate"], hparams["loader_batch_size"], hparams["grad_accumulation_factor"], etc.: Adjust these based on your dataset size, GPU memory, and desired training regime.hparams["use_wandb"]: Set to True to enable logging.hparams["wandb_project"]: Change "you project's name on weights and biases " to your project name.hparams["wandb_entity"]: Optionally set your W&B username or team name.hparams["wandb_resume_id"]: Set to a specific run ID string (e.g., "ceeu3g6c") to resume a previous W&B run. Leave as None for a new run.gpu="A100-40GB" argument in the @app.function decorator if you need a different GPU type available on Modal (e.g., "T4", "A10G").Execute the script locally using the Modal CLI:
modal run train_whisper_modal.py
This command deploys the code to Modal and starts the train_whisper_on_modal function on a remote container with the specified GPU.
infer_whisper_local.py)This script loads a fine-tuned SpeechBrain checkpoint (saved to the Modal Volume during training) and performs transcription on a local audio file. It can also optionally transcribe the same audio using OpenAI and Azure APIs for comparison.
Before running inference, you need the fine-tuned checkpoint file (model.ckpt) saved during training. Download it from your Modal Volume:
/root/checkpoints/whisper_small_egy_save/CKPT+.../model.ckpt).modal volume get command knowledge - refer to Modal docs):
# Example - syntax might vary, check Modal docs for volume file operations
# modal volume get <your-volume-name> /root/checkpoints/whisper_small_egy_save/CKPT+.../model.ckpt ./downloaded_model.ckpt
Alternatively, add a simple Modal function to list files or download specific ones.Execute the script from your terminal:
python infer_whisper_local.py --ckpt_path /path/to/your/downloaded_model.ckpt --audio_path /path/to/your/audio.wav [OPTIONS]
Required Arguments:
--ckpt_path: Path to the downloaded .ckpt file from your fine-tuning run.--audio_path: Path to the audio file (e.g., .wav, .mp3) you want to transcribe.Optional Arguments:
--model_hub: Base Whisper model used for fine-tuning (e.g., "openai/whisper-small"). Crucially, this MUST match the whisper_hub used during training. Defaults to "openai/whisper-small".--language: Language code (e.g., "ar", "en"). Used for local model decoding and OpenAI API. Defaults to "ar".--task: Task for the local model ("transcribe" or "translate"). Defaults to "transcribe".--device: Device for inference ("cuda" or "cpu"). Defaults to CUDA if available.--num_beams: Number of beams for local model generation. Defaults to 5.--openai_api_key: OpenAI API key. If not provided, uses OPENAI_API_KEY from .env. Skips OpenAI if neither is found.--azure_speech_key: Azure Speech key. If not provided, uses AZURE_SPEECH_KEY from .env. Skips Azure if not found.--azure_service_region: Azure Speech region. If not provided, uses AZURE_SERVICE_REGION from .env. Skips Azure if not found.--azure_language_locale: Specific locale for Azure (e.g., "ar-EG", "en-US"). Defaults to "ar-EG".--output_dir: Directory to save the transcription results as text files. Defaults to the current directory (.).audio__local.txt, audio__openai.txt) will be saved in the specified --output_dir.language parameter in hparams (train_whisper_modal.py) and the --language / --azure_language_locale arguments (infer_whisper_local.py).hparams["hf_dataset_id"] to point to your Hugging Face dataset. Ensure it's compatible (audio, text columns).hparams["whisper_hub"] and the --model_hub argument consistently to use different Whisper model sizes. Note that larger models require more GPU memory.hparams to potentially improve model robustness.pip install --upgrade modal-client), are logged in (modal token set), and have correctly created the necessary secrets. Check Modal dashboard logs for detailed error messages from the container.infer_whisper_local.py):
--model_hub argument exactly matches the base model used for training (hparams["whisper_hub"]). Architecture mismatches are common errors.--ckpt_path points to the correct, fully downloaded model.ckpt file.infer_whisper_local.py):
.env file or command-line arguments.--device cuda for inference. Check GPU driver compatibility.hparams are valid.--audio_path, --ckpt_path).Python
100.0%
Python
0
33 commits
updated Jun 16, 2025
This project provides scripts to fine-tune an OpenAI Whisper model on a custom dataset using Modal for GPU-accelerated training and then run inference locally using the fine-tuned checkpoint, with optional comparison against OpenAI and Azure transcription APIs.
openai/whisper-small by default) on a custom Arabic dataset (e.g., Egyptian dialect).train_whisper_modal.py: Script to define and run the fine-tuning process on Modal.infer_whisper_local.py: Script to load a fine-tuned checkpoint and perform transcription locally or compare with APIs..env: (To be created by user) Stores API keys securely.README.md: This file.pip install modal-client
modal setup
git clone <your-repository-url>
cd <repository-directory>
infer_whisper_local.py to compare with APIs:
Modal uses Secrets to securely store credentials. Create the following secrets via the Modal Secrets page or CLI:
modal secret create huggingface-secret-write HF_TOKEN=<your-huggingface-write-token>
modal secret create wandb-secret WANDB_API_KEY=<your-wandb-api-key>
.env file)Create a file named .env in the root of the project directory to store API keys for local inference (if needed). Do not commit this file to Git.
# .env file
OPENAI_API_KEY="sk-..."
AZURE_SPEECH_KEY="..."
AZURE_SERVICE_REGION="YourAzureRegion" # e.g., westeurope, eastus
infer_whisper_local.py will automatically load these variables.
train_whisper_modal.py)This script defines a Modal function to perform the fine-tuning process on a GPU instance in the cloud.
Before running, review and modify the placeholders and hparams dictionary within train_whisper_modal.py:
Placeholder comments near the modal.App and modal.Volume.from_name definitions and replace them with your desired names.hparams["hf_dataset_id"]: Change "huggingfaceusername/datasetname" to the Hugging Face dataset identifier for your custom dataset (e.g., "MAdel121/arabic-egy-cleaned"). Ensure the dataset has 'audio' and 'text' columns. The script assumes standard splits like 'train', 'validation', 'test'.hparams["whisper_hub"]: Change "openai/whisper-small" if you want to fine-tune a different Whisper variant (e.g., openai/whisper-base, openai/whisper-medium). Ensure this matches the base model used in infer_whisper_local.py.hparams["language"]: Set the target language code (e.g., "ar", "en").hparams["task"]: Usually "transcribe".hparams["save_folder"], hparams["output_folder"]: Define paths within the Modal volume where checkpoints and logs will be saved. Usually, the defaults are fine.hparams["augment"]: True or False to enable/disable augmentation.hparams["augment_prob_master"]: Overall probability of applying augmentation per batch.hparams["use_*"] flags: Toggle specific augmentations (AddNoise, AddReverb, SpeedPerturb, etc.).hparams["epochs"], hparams["learning_rate"], hparams["loader_batch_size"], hparams["grad_accumulation_factor"], etc.: Adjust these based on your dataset size, GPU memory, and desired training regime.hparams["use_wandb"]: Set to True to enable logging.hparams["wandb_project"]: Change "you project's name on weights and biases " to your project name.hparams["wandb_entity"]: Optionally set your W&B username or team name.hparams["wandb_resume_id"]: Set to a specific run ID string (e.g., "ceeu3g6c") to resume a previous W&B run. Leave as None for a new run.gpu="A100-40GB" argument in the @app.function decorator if you need a different GPU type available on Modal (e.g., "T4", "A10G").Execute the script locally using the Modal CLI:
modal run train_whisper_modal.py
This command deploys the code to Modal and starts the train_whisper_on_modal function on a remote container with the specified GPU.
infer_whisper_local.py)This script loads a fine-tuned SpeechBrain checkpoint (saved to the Modal Volume during training) and performs transcription on a local audio file. It can also optionally transcribe the same audio using OpenAI and Azure APIs for comparison.
Before running inference, you need the fine-tuned checkpoint file (model.ckpt) saved during training. Download it from your Modal Volume:
/root/checkpoints/whisper_small_egy_save/CKPT+.../model.ckpt).modal volume get command knowledge - refer to Modal docs):
# Example - syntax might vary, check Modal docs for volume file operations
# modal volume get <your-volume-name> /root/checkpoints/whisper_small_egy_save/CKPT+.../model.ckpt ./downloaded_model.ckpt
Alternatively, add a simple Modal function to list files or download specific ones.Execute the script from your terminal:
python infer_whisper_local.py --ckpt_path /path/to/your/downloaded_model.ckpt --audio_path /path/to/your/audio.wav [OPTIONS]
Required Arguments:
--ckpt_path: Path to the downloaded .ckpt file from your fine-tuning run.--audio_path: Path to the audio file (e.g., .wav, .mp3) you want to transcribe.Optional Arguments:
--model_hub: Base Whisper model used for fine-tuning (e.g., "openai/whisper-small"). Crucially, this MUST match the whisper_hub used during training. Defaults to "openai/whisper-small".--language: Language code (e.g., "ar", "en"). Used for local model decoding and OpenAI API. Defaults to "ar".--task: Task for the local model ("transcribe" or "translate"). Defaults to "transcribe".--device: Device for inference ("cuda" or "cpu"). Defaults to CUDA if available.--num_beams: Number of beams for local model generation. Defaults to 5.--openai_api_key: OpenAI API key. If not provided, uses OPENAI_API_KEY from .env. Skips OpenAI if neither is found.--azure_speech_key: Azure Speech key. If not provided, uses AZURE_SPEECH_KEY from .env. Skips Azure if not found.--azure_service_region: Azure Speech region. If not provided, uses AZURE_SERVICE_REGION from .env. Skips Azure if not found.--azure_language_locale: Specific locale for Azure (e.g., "ar-EG", "en-US"). Defaults to "ar-EG".--output_dir: Directory to save the transcription results as text files. Defaults to the current directory (.).audio__local.txt, audio__openai.txt) will be saved in the specified --output_dir.language parameter in hparams (train_whisper_modal.py) and the --language / --azure_language_locale arguments (infer_whisper_local.py).hparams["hf_dataset_id"] to point to your Hugging Face dataset. Ensure it's compatible (audio, text columns).hparams["whisper_hub"] and the --model_hub argument consistently to use different Whisper model sizes. Note that larger models require more GPU memory.hparams to potentially improve model robustness.pip install --upgrade modal-client), are logged in (modal token set), and have correctly created the necessary secrets. Check Modal dashboard logs for detailed error messages from the container.infer_whisper_local.py):
--model_hub argument exactly matches the base model used for training (hparams["whisper_hub"]). Architecture mismatches are common errors.--ckpt_path points to the correct, fully downloaded model.ckpt file.infer_whisper_local.py):
.env file or command-line arguments.--device cuda for inference. Check GPU driver compatibility.hparams are valid.--audio_path, --ckpt_path).Python
100.0%