Nkululeko is a very common first name from the Zulu and Xhosa languages in South Africa and means "freedom". It is also the name of a software to detect speaker characteristics by machine learning experiments with a high-level interface. The idea is to have a framework (based on e.g. sklearn and torch) that can be used to rapidly and automatically analyse audio data and explore machine learning models based on that data.
Some abilities that Nkululeko provides: combines acoustic features and machine learning models (including feature selection and features concatenation); performs data exploration, selection and visualization the results; finetuning; ensemble learning models; soft labeling (predicting labels with pre-trained model); and inference the model on a test set.
Nkululeko orchestrates data loading, feature extraction, and model training, allowing you to specify your experiment in a configuration file. The framework handles the process from raw data to trained model and evaluation, making it easy to run machine learning experiments without directly coding in Python.
Nkululeko is for speech processing learners, researchers and ML practitioners focused on speaker characteristics, e.g., emotion, age, gender, or disorder detection.
Nkululeko requires Python 3.10 or higher with the following build status:
Create and activate a virtual Python environment and simply install Nkululeko:
# using python venv
python -m venv .env
source .env/bin/activate # specify OS versions, add a separate line for Windows users
pip install nkululeko
# using uv in development mode
uv venv --python 3.12
source .venv/bin/activate
uv pip install -e .
# or run directly using uv run after cloning
uv run python -m nkululeko.train --config examples/exp_polish_tree.ini
Nkululeko supports optional dependencies through extras:
# Install with PyTorch support
pip install nkululeko[torch]
# Install with CPU-only PyTorch
pip install nkululeko[torch-cpu]
# Install with TensorFlow support
pip install nkululeko[tensorflow]
# Install all optional dependencies
pip install nkululeko[all]
You can also install dependencies manually:
For CPU-only installation (recommended for most users):
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu
For GPU support (latest CUDA):
pip install torch torchvision torchaudio
If you encounter issues with CUDA dependencies being installed even with the torch-cpu extra (this can happen with some package managers), you can use this manual installation approach with uv:
# Create virtual environment
uv venv
source .venv/bin/activate
# Install CPU-only PyTorch
uv pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu --python .venv/bin/python
# Install nkululeko without dependencies
uv pip install -e . --no-deps --python .venv/bin/python
# Install core dependencies
uv pip install pandas numpy scipy matplotlib seaborn scikit-learn --python .venv/bin/python
# Install audio-related dependencies
uv pip install audb audeer audformat audplot audmodel auglib audinterface audiofile audmetric audonnx confidence-intervals --python .venv/bin/python
# Install remaining packages (excluding xgboost temporarily)
uv pip install opensmile praat-parselmouth sounddevice imageio datasets transformers umap-learn pylatex audiomentations googletrans sentencepiece shap numpy_audio_limiter splitutils --python .venv/bin/python
# Install xgboost (choose one option):
# Option 1: Use older version without CUDA dependencies
uv pip install "xgboost<2.0.0" --python .venv/bin/python
# Option 2: Use the CPU-only xgboost package (recommended for newer versions)
uv pip install xgboost-cpu --python .venv/bin/python
This approach ensures that only CPU-only versions of packages are installed, avoiding CUDA dependencies on systems without GPU support.
Some functionalities require extra packages to be installed, which we didn't include automatically:
pip install PyYAML # Install PyYAML first to avoid dependency issues
pip install nkululeko[spotlight]
Some examples for ini-files (which you use to control nkululeko) are in the examples folder.
The documentation, along with extensions of installation, usage, INI file format, and examples, can be found nkululeko.readthedocs.io.
Basically, you specify your experiment in an "ini" file (e.g. experiment.ini) and then call one of the Nkululeko interfaces to run the experiment like this:
python -m nkululeko.train --config experiment.ini
A basic configuration looks like this:
[EXP]
root = ./
name = exp_emodb
[DATA]
databases = ['emodb']
emodb = ./emodb/
emodb.split_strategy = speaker_split
target = emotion
labels = ['anger', 'boredom', 'disgust', 'fear']
[FEATS]
type = ['praat']
[MODEL]
type = svm
[EXPL]
model = tree
plot_tree = True
Read the Hello World example for initial usage with Emo-DB dataset.
Here is an overview of the interfaces/modules:
All of them take --config <my_config.ini> as an argument.
nkululeko.demo, nkululeko.feature_demo and nkululeko.testing modules.python, python should start with version >3 (NOT 2!). You can leave the Python Interpreter by typing exit()nkulu_worknkulu_work)nkulu_work folder:
python -m venv venvsource venv/bin/activatevenv\Scripts\activate.bat(venv) in front of your promptpip install nkululekonkulu_work)python -m nkululeko.train --config exp_emodb.iniexp_emodb/images/run_0/emodb_xgb_os_0_000_cnf.pngThe framework is targeted at the speech domain and supports experiments where different classifiers are combined with different feature extractors.
Here's a rough UML-like sketch of the framework (and here's the real one done with pyreverse).

Currently, the following linear classifiers are implemented (integrated from sklearn):
For visualization, besides confusion matrix, feature importance, feature distribution, t-SNE plot, data distribution (just names a few), Nkululeko can also be used for bias checking, uncertainty estimation, and epoch progression.
Here's an animation that shows the progress of classification done with nkululeko.
There's Felix blog with tutorials below:
Nkululeko can be used under the MIT license.
Contributions are welcome and encouraged. To learn more about how to contribute to nkululeko, please refer to the Contributing guidelines.
If you use Nkululeko, please cite the paper:
F. Burkhardt and B. Tris Atmaja, (2025). Nkululeko 1.0: A Python package to predict speaker characteristics with a high-level interface. Journal of Open Source Software, 10(115), 8049, https://doi.org/10.21105/joss.08049
@article{Burkhardt2025, doi = {10.21105/joss.08049}, url = {https://doi.org/10.21105/joss.08049}, year = {2025}, publisher = {The Open Journal}, volume = {10}, number = {115}, pages = {8049}, author = {Burkhardt, Felix and Atmaja, Bagus Tris}, title = {Nkululeko 1.0: A Python package to predict speaker characteristics with a high-level interface}, journal = {Journal of Open Source Software} }
Python
98.3%
Nkululeko is a very common first name from the Zulu and Xhosa languages in South Africa and means "freedom". It is also the name of a software to detect speaker characteristics by machine learning experiments with a high-level interface. The idea is to have a framework (based on e.g. sklearn and torch) that can be used to rapidly and automatically analyse audio data and explore machine learning models based on that data.
Some abilities that Nkululeko provides: combines acoustic features and machine learning models (including feature selection and features concatenation); performs data exploration, selection and visualization the results; finetuning; ensemble learning models; soft labeling (predicting labels with pre-trained model); and inference the model on a test set.
Nkululeko orchestrates data loading, feature extraction, and model training, allowing you to specify your experiment in a configuration file. The framework handles the process from raw data to trained model and evaluation, making it easy to run machine learning experiments without directly coding in Python.
Nkululeko is for speech processing learners, researchers and ML practitioners focused on speaker characteristics, e.g., emotion, age, gender, or disorder detection.
Nkululeko requires Python 3.10 or higher with the following build status:
Create and activate a virtual Python environment and simply install Nkululeko:
# using python venv
python -m venv .env
source .env/bin/activate # specify OS versions, add a separate line for Windows users
pip install nkululeko
# using uv in development mode
uv venv --python 3.12
source .venv/bin/activate
uv pip install -e .
# or run directly using uv run after cloning
uv run python -m nkululeko.train --config examples/exp_polish_tree.ini
Nkululeko supports optional dependencies through extras:
# Install with PyTorch support
pip install nkululeko[torch]
# Install with CPU-only PyTorch
pip install nkululeko[torch-cpu]
# Install with TensorFlow support
pip install nkululeko[tensorflow]
# Install all optional dependencies
pip install nkululeko[all]
You can also install dependencies manually:
For CPU-only installation (recommended for most users):
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu
For GPU support (latest CUDA):
pip install torch torchvision torchaudio
If you encounter issues with CUDA dependencies being installed even with the torch-cpu extra (this can happen with some package managers), you can use this manual installation approach with uv:
# Create virtual environment
uv venv
source .venv/bin/activate
# Install CPU-only PyTorch
uv pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu --python .venv/bin/python
# Install nkululeko without dependencies
uv pip install -e . --no-deps --python .venv/bin/python
# Install core dependencies
uv pip install pandas numpy scipy matplotlib seaborn scikit-learn --python .venv/bin/python
# Install audio-related dependencies
uv pip install audb audeer audformat audplot audmodel auglib audinterface audiofile audmetric audonnx confidence-intervals --python .venv/bin/python
# Install remaining packages (excluding xgboost temporarily)
uv pip install opensmile praat-parselmouth sounddevice imageio datasets transformers umap-learn pylatex audiomentations googletrans sentencepiece shap numpy_audio_limiter splitutils --python .venv/bin/python
# Install xgboost (choose one option):
# Option 1: Use older version without CUDA dependencies
uv pip install "xgboost<2.0.0" --python .venv/bin/python
# Option 2: Use the CPU-only xgboost package (recommended for newer versions)
uv pip install xgboost-cpu --python .venv/bin/python
This approach ensures that only CPU-only versions of packages are installed, avoiding CUDA dependencies on systems without GPU support.
Some functionalities require extra packages to be installed, which we didn't include automatically:
pip install PyYAML # Install PyYAML first to avoid dependency issues
pip install nkululeko[spotlight]
Some examples for ini-files (which you use to control nkululeko) are in the examples folder.
The documentation, along with extensions of installation, usage, INI file format, and examples, can be found nkululeko.readthedocs.io.
Basically, you specify your experiment in an "ini" file (e.g. experiment.ini) and then call one of the Nkululeko interfaces to run the experiment like this:
python -m nkululeko.train --config experiment.ini
A basic configuration looks like this:
[EXP]
root = ./
name = exp_emodb
[DATA]
databases = ['emodb']
emodb = ./emodb/
emodb.split_strategy = speaker_split
target = emotion
labels = ['anger', 'boredom', 'disgust', 'fear']
[FEATS]
type = ['praat']
[MODEL]
type = svm
[EXPL]
model = tree
plot_tree = True
Read the Hello World example for initial usage with Emo-DB dataset.
Here is an overview of the interfaces/modules:
All of them take --config <my_config.ini> as an argument.
nkululeko.demo, nkululeko.feature_demo and nkululeko.testing modules.python, python should start with version >3 (NOT 2!). You can leave the Python Interpreter by typing exit()nkulu_worknkulu_work)nkulu_work folder:
python -m venv venvsource venv/bin/activatevenv\Scripts\activate.bat(venv) in front of your promptpip install nkululekonkulu_work)python -m nkululeko.train --config exp_emodb.iniexp_emodb/images/run_0/emodb_xgb_os_0_000_cnf.pngThe framework is targeted at the speech domain and supports experiments where different classifiers are combined with different feature extractors.
Here's a rough UML-like sketch of the framework (and here's the real one done with pyreverse).

Currently, the following linear classifiers are implemented (integrated from sklearn):
For visualization, besides confusion matrix, feature importance, feature distribution, t-SNE plot, data distribution (just names a few), Nkululeko can also be used for bias checking, uncertainty estimation, and epoch progression.
Here's an animation that shows the progress of classification done with nkululeko.
There's Felix blog with tutorials below:
Nkululeko can be used under the MIT license.
Contributions are welcome and encouraged. To learn more about how to contribute to nkululeko, please refer to the Contributing guidelines.
If you use Nkululeko, please cite the paper:
F. Burkhardt and B. Tris Atmaja, (2025). Nkululeko 1.0: A Python package to predict speaker characteristics with a high-level interface. Journal of Open Source Software, 10(115), 8049, https://doi.org/10.21105/joss.08049
@article{Burkhardt2025, doi = {10.21105/joss.08049}, url = {https://doi.org/10.21105/joss.08049}, year = {2025}, publisher = {The Open Journal}, volume = {10}, number = {115}, pages = {8049}, author = {Burkhardt, Felix and Atmaja, Bagus Tris}, title = {Nkululeko 1.0: A Python package to predict speaker characteristics with a high-level interface}, journal = {Journal of Open Source Software} }
Python
98.3%