7 repos
Speech emotion recognition (SER) systems that detect and classify emotional states—valence, arousal, dominance, and categorical emotions—from audio signals. The cluster centers on WavLM-based models trained on the MSP-Podcast corpus, a primary benchmark dataset for emotion recognition research. Repositories here implement baseline models, fine-tuned variants for multiple languages, and evaluation frameworks for predicting emotional dimensions from spoken language.