Whisper Medium Egyptian Arabic (whisper-medium-egy)
6
4 commits
2 linked in READMEs
updated May 21, 2025
This model is a fine-tuned version of openai/whisper-medium on a custom dataset of 72 hours of Egyptian Arabic speech. It's designed for Automatic Speech Recognition (ASR) for the Egyptian Arabic dialect.
openai/whisper-mediumMAdel121/arabic-egy-cleaned (approx. 72 hours)This model is intended for transcribing speech in Egyptian Arabic.
Intended Use:
Limitations:
You can use this model with the transformers library and the pipeline interface for ease of use.
from transformers import pipeline
import torch
device = "cuda:0" if torch.cuda.is_available() else "cpu"
pipe = pipeline(
"automatic-speech-recognition",
model="YOUR_HF_USERNAME/whisper-medium-egy", # Replace YOUR_HF_USERNAME with your Hugging Face username
device=device
)
# Example with a local audio file
# audio_file = "path/to/your/egyptian_arabic_audio.wav"
# transcription = pipe(audio_file, generate_kwargs={"language": "arabic"})["text"]
# print(transcription)
# Example with a Hugging Face dataset audio sample
# from datasets import load_dataset
# ds = load_dataset("MAdel121/arabic-egy-cleaned", "ar", split="validation") # Or your test split
# sample = ds[0]["audio"] # Make sure your dataset has an "audio" column
# result = pipe(sample.copy(), generate_kwargs={"language": "arabic"})
# print(result["text"])
Make sure to replace "YOUR_HF_USERNAME/whisper-medium-egy" with the actual model ID after uploading. The generate_kwargs={"language": "arabic"} is important for Whisper models to ensure correct tokenization and transcription for the target language.
The model was fine-tuned on the MAdel121/arabic-egy-cleaned dataset available on the Hugging Face Hub. This dataset contains approximately 72 hours of Egyptian Arabic audio paired with transcripts.
The model was trained using the transformers library. The fine-tuning process involved the following key hyperparameters:
openai/whisper-mediumuse_drop_freq: trueuse_drop_chunk: trueuse_drop_bit_resolution: trueuse_add_noise, use_speed_perturb, use_pitch_shift, use_add_reverb, use_codec_augment, use_gain were set to falseTraining was done on 1x A100 (80GB) on Modal Labs
The training was managed and tracked using Weights & Biases under the project whisper-medium-egyptian-arabic with resume ID r3sz4v27.
Can be found on Github here
Run can be found here : https://wandb.ai/m-adelomar1/whisper-medium-egyptian-arabic/
The model was evaluated on the validation split of the MAdel121/arabic-egy-cleaned dataset.
These metrics indicate the performance of the model on the validation set. Lower values are better.
@misc{madel_2025_whisper_medium_egy,
author = Madel
title = {Whisper Medium Fine-tuned for Egyptian Arabic},
year = {2025},
publisher = {Hugging Face},
journal = {Hugging Face Hub},
howpublished = {\\url{https://huggingface.co/MAdel121/whisper-medium-egy}} // Replace with actual URL
}
Whisper Medium Egyptian Arabic (whisper-medium-egy)
6
4 commits
2 linked in READMEs
updated May 21, 2025
This model is a fine-tuned version of openai/whisper-medium on a custom dataset of 72 hours of Egyptian Arabic speech. It's designed for Automatic Speech Recognition (ASR) for the Egyptian Arabic dialect.
openai/whisper-mediumMAdel121/arabic-egy-cleaned (approx. 72 hours)This model is intended for transcribing speech in Egyptian Arabic.
Intended Use:
Limitations:
You can use this model with the transformers library and the pipeline interface for ease of use.
from transformers import pipeline
import torch
device = "cuda:0" if torch.cuda.is_available() else "cpu"
pipe = pipeline(
"automatic-speech-recognition",
model="YOUR_HF_USERNAME/whisper-medium-egy", # Replace YOUR_HF_USERNAME with your Hugging Face username
device=device
)
# Example with a local audio file
# audio_file = "path/to/your/egyptian_arabic_audio.wav"
# transcription = pipe(audio_file, generate_kwargs={"language": "arabic"})["text"]
# print(transcription)
# Example with a Hugging Face dataset audio sample
# from datasets import load_dataset
# ds = load_dataset("MAdel121/arabic-egy-cleaned", "ar", split="validation") # Or your test split
# sample = ds[0]["audio"] # Make sure your dataset has an "audio" column
# result = pipe(sample.copy(), generate_kwargs={"language": "arabic"})
# print(result["text"])
Make sure to replace "YOUR_HF_USERNAME/whisper-medium-egy" with the actual model ID after uploading. The generate_kwargs={"language": "arabic"} is important for Whisper models to ensure correct tokenization and transcription for the target language.
The model was fine-tuned on the MAdel121/arabic-egy-cleaned dataset available on the Hugging Face Hub. This dataset contains approximately 72 hours of Egyptian Arabic audio paired with transcripts.
The model was trained using the transformers library. The fine-tuning process involved the following key hyperparameters:
openai/whisper-mediumuse_drop_freq: trueuse_drop_chunk: trueuse_drop_bit_resolution: trueuse_add_noise, use_speed_perturb, use_pitch_shift, use_add_reverb, use_codec_augment, use_gain were set to falseTraining was done on 1x A100 (80GB) on Modal Labs
The training was managed and tracked using Weights & Biases under the project whisper-medium-egyptian-arabic with resume ID r3sz4v27.
Can be found on Github here
Run can be found here : https://wandb.ai/m-adelomar1/whisper-medium-egyptian-arabic/
The model was evaluated on the validation split of the MAdel121/arabic-egy-cleaned dataset.
These metrics indicate the performance of the model on the validation set. Lower values are better.
@misc{madel_2025_whisper_medium_egy,
author = Madel
title = {Whisper Medium Fine-tuned for Egyptian Arabic},
year = {2025},
publisher = {Hugging Face},
journal = {Hugging Face Hub},
howpublished = {\\url{https://huggingface.co/MAdel121/whisper-medium-egy}} // Replace with actual URL
}