Compare the unmodified Whisper-large-v3-turbo with our dialect-balanced QLoRA fine-tune on the same audio clip. Trained on MSA + Egyptian + Levantine + Gulf (Maghrebi excluded); see the model card for held-out per-dialect WER and the paper draft for the full methodology.
This Space runs on free CPU, so transcription takes 30–60 seconds for a 10-second clip.
For production deployment use the model directly via transformers or faster-whisper —
the model card has copy-pasteable usage code.
2 commits
Compare the unmodified Whisper-large-v3-turbo with our dialect-balanced QLoRA fine-tune on the same audio clip. Trained on MSA + Egyptian + Levantine + Gulf (Maghrebi excluded); see the model card for held-out per-dialect WER and the paper draft for the full methodology.
This Space runs on free CPU, so transcription takes 30–60 seconds for a 10-second clip.
For production deployment use the model directly via transformers or faster-whisper —
the model card has copy-pasteable usage code.
2 commits