This repository contains the code for the COS760 Natural Language Research Project, which focuses on ASR models for African languages. The project is part of the COS-760 module provided by the University of Pretoria.
This project was developed by the following team members:
To get started, clone the repository:
# With HTTPS:
git clone https://github.com/GremBleen/COS-760-Project.git
# With SSH:
git clone git@github.com:GremBleen/COS-760-Project.git
├── analyse_results.sh
├── post-refinement-runs
├── pre-refinement-runs
├── presets.json
├── README.md
├── src
│ ├── common.py
│ ├── main.py
│ └── models
│ ├── facebook_mms.py
│ ├── lelapa.py
│ ├── SM4T.py
│ ├── wav2vec.py
│ └── whisper.py
└── stat_extract.py
analyse_results.sh is a shell script that can be used to analyze the results of the ASR models.
post-refinement-runs and pre-refinement-runs are directories that contain the results of the ASR models before and after refinement, respectively.
presets.json is a configuration file that contains the settings for running the models. It is necessary to alter this file to run the models with different configurations.
src contains all source code.
stat_extract.py is a deprecated script that was not used. Numpy was instead used.
After cloning the repository, you can set up the Python environment.
We use Python 3.11 for this project. It is recommended to use a virtual environment to manage dependencies and avoid conflicts with other projects. You can use venv, pipenv, conda, or any environment manager of your choice. Below are some options to create a virtual environment:
python3 -m venv venv
If you have pipenv installed:
pipenv --python 3.11
If you are using conda:
conda create -n venv python=3.11
After setting up the virtual environment, activate it and install the required dependencies.
Activate the virtual environment:
source venv/bin/activate # For Linux/MacOS
venv\Scripts\activate # For Windows
Note:
The project originally used conda for environment and dependency management. However, due to issues when creating environments on different operating systems, we now recommend installing dependencies manually. If you encounter errors when running the project, check if additional dependencies are required.
Below is a list of required dependencies:
datasets>=3.6.0
dotenv>=0.9.9
fuzzywuzzy>=0.18.0
huggingface-hub>=0.33.0
jiwer>=3.1.0
librosa>=0.11.0
matplotlib>=3.10.3
python-levenshtein>=0.27.1
torch>=2.7.1
torchaudio>=2.7.1
transformers>=4.52.4
After completing the setup, create a .env file in the root directory. This file stores environment variables and helps keep sensitive information secure.
The .env file should contain:
HUGGINGFACE_TOKEN=<INPUT TOKEN HERE>
To run the models, create a presets.json file in the main directory with the following structure:
{
"dataset_language": "<LANGUAGE CODE>",
"model": "<MODEL NAME>",
"batch_size": 20,
"refinement_method": true,
"debug": true
}
Replace <LANGUAGE CODE> with the code for your dataset's language (e.g., "afr" for Afrikaans, "xho" for Xhosa, "zul" for Zulu).
Replace <MODEL NAME> with the desired model (e.g., "facebook-mms", "lelapa", "sm4t", "wav2vec", "whisper-large", "whisper-medium").
You can adjust batch_size, but we recommend keeping it at 20 for a good balance between performance and memory usage.
Set refinement_method to true or false depending on whether you want to use the refinement method.
Set debug to true or false to control whether running information is printed to the console.
To run the project, use the following command from the root directory:
python3 src/main.py
This repository makes use of the NCHLT datasets through a Huggingface interface. Available at: NCHLT Datasets. It is not necessary to download the datasets as this is done implicitly by running the scripts.
We use the Git Flow workflow.
If Git Flow is installed, initialize it with:
git flow init
To start a new feature:
git flow feature start <feature-name>
To finish a feature:
git flow feature finish <feature-name>
Python
99.7%
This repository contains the code for the COS760 Natural Language Research Project, which focuses on ASR models for African languages. The project is part of the COS-760 module provided by the University of Pretoria.
This project was developed by the following team members:
To get started, clone the repository:
# With HTTPS:
git clone https://github.com/GremBleen/COS-760-Project.git
# With SSH:
git clone git@github.com:GremBleen/COS-760-Project.git
├── analyse_results.sh
├── post-refinement-runs
├── pre-refinement-runs
├── presets.json
├── README.md
├── src
│ ├── common.py
│ ├── main.py
│ └── models
│ ├── facebook_mms.py
│ ├── lelapa.py
│ ├── SM4T.py
│ ├── wav2vec.py
│ └── whisper.py
└── stat_extract.py
analyse_results.sh is a shell script that can be used to analyze the results of the ASR models.
post-refinement-runs and pre-refinement-runs are directories that contain the results of the ASR models before and after refinement, respectively.
presets.json is a configuration file that contains the settings for running the models. It is necessary to alter this file to run the models with different configurations.
src contains all source code.
stat_extract.py is a deprecated script that was not used. Numpy was instead used.
After cloning the repository, you can set up the Python environment.
We use Python 3.11 for this project. It is recommended to use a virtual environment to manage dependencies and avoid conflicts with other projects. You can use venv, pipenv, conda, or any environment manager of your choice. Below are some options to create a virtual environment:
python3 -m venv venv
If you have pipenv installed:
pipenv --python 3.11
If you are using conda:
conda create -n venv python=3.11
After setting up the virtual environment, activate it and install the required dependencies.
Activate the virtual environment:
source venv/bin/activate # For Linux/MacOS
venv\Scripts\activate # For Windows
Note:
The project originally used conda for environment and dependency management. However, due to issues when creating environments on different operating systems, we now recommend installing dependencies manually. If you encounter errors when running the project, check if additional dependencies are required.
Below is a list of required dependencies:
datasets>=3.6.0
dotenv>=0.9.9
fuzzywuzzy>=0.18.0
huggingface-hub>=0.33.0
jiwer>=3.1.0
librosa>=0.11.0
matplotlib>=3.10.3
python-levenshtein>=0.27.1
torch>=2.7.1
torchaudio>=2.7.1
transformers>=4.52.4
After completing the setup, create a .env file in the root directory. This file stores environment variables and helps keep sensitive information secure.
The .env file should contain:
HUGGINGFACE_TOKEN=<INPUT TOKEN HERE>
To run the models, create a presets.json file in the main directory with the following structure:
{
"dataset_language": "<LANGUAGE CODE>",
"model": "<MODEL NAME>",
"batch_size": 20,
"refinement_method": true,
"debug": true
}
Replace <LANGUAGE CODE> with the code for your dataset's language (e.g., "afr" for Afrikaans, "xho" for Xhosa, "zul" for Zulu).
Replace <MODEL NAME> with the desired model (e.g., "facebook-mms", "lelapa", "sm4t", "wav2vec", "whisper-large", "whisper-medium").
You can adjust batch_size, but we recommend keeping it at 20 for a good balance between performance and memory usage.
Set refinement_method to true or false depending on whether you want to use the refinement method.
Set debug to true or false to control whether running information is printed to the console.
To run the project, use the following command from the root directory:
python3 src/main.py
This repository makes use of the NCHLT datasets through a Huggingface interface. Available at: NCHLT Datasets. It is not necessary to download the datasets as this is done implicitly by running the scripts.
We use the Git Flow workflow.
If Git Flow is installed, initialize it with:
git flow init
To start a new feature:
git flow feature start <feature-name>
To finish a feature:
git flow feature finish <feature-name>
Python
99.7%