This model is a fine-tuned version of OpenAI's Whisper Medium EN model, specifically trained on Air Traffic Control (ATC) communication datasets. The fine-tuning process significantly improves transcription accuracy on domain-specific aviation communications, reducing the Word Error Rate (WER) by 84%, compared to the original pretrained model. The model is particularly effective at handling accent variations and ambiguous phrasing often encountered in ATC communications.
This model has been converted to an optimized .bin format, making it compatible with Faster-Whisper for faster and more efficient inference.
You can access the fine-tuned model on Hugging Face:
Whisper Medium EN fine-tuned for ATC is optimized to handle short, distinct transmissions between pilots and air traffic controllers. It is fine-tuned using data from the ATC Dataset, a combined and cleaned dataset sourced from the following:
The ATC Dataset merges these two original sources, filtering and refining the data to enhance transcription accuracy for domain-specific ATC communications. The model has been further optimized to a .bin format for compatibility with Faster-Whisper, ensuring faster and more efficient processing.
The fine-tuned Whisper model is designed for:
You can test the model online using the ATC Transcription Assistant, which lets you upload audio files and generate transcriptions.
While the fine-tuned model performs well in ATC-specific communications, it may not generalize as effectively to other domains of speech. Additionally, like most speech-to-text models, transcription accuracy can be affected by extremely poor-quality audio or heavily accented speech not encountered during training.
3 commits
This model is a fine-tuned version of OpenAI's Whisper Medium EN model, specifically trained on Air Traffic Control (ATC) communication datasets. The fine-tuning process significantly improves transcription accuracy on domain-specific aviation communications, reducing the Word Error Rate (WER) by 84%, compared to the original pretrained model. The model is particularly effective at handling accent variations and ambiguous phrasing often encountered in ATC communications.
This model has been converted to an optimized .bin format, making it compatible with Faster-Whisper for faster and more efficient inference.
You can access the fine-tuned model on Hugging Face:
Whisper Medium EN fine-tuned for ATC is optimized to handle short, distinct transmissions between pilots and air traffic controllers. It is fine-tuned using data from the ATC Dataset, a combined and cleaned dataset sourced from the following:
The ATC Dataset merges these two original sources, filtering and refining the data to enhance transcription accuracy for domain-specific ATC communications. The model has been further optimized to a .bin format for compatibility with Faster-Whisper, ensuring faster and more efficient processing.
The fine-tuned Whisper model is designed for:
You can test the model online using the ATC Transcription Assistant, which lets you upload audio files and generate transcriptions.
While the fine-tuned model performs well in ATC-specific communications, it may not generalize as effectively to other domains of speech. Additionally, like most speech-to-text models, transcription accuracy can be affected by extremely poor-quality audio or heavily accented speech not encountered during training.
3 commits