Our contribution
(1) Dual-stream representation is designed. It fully trains a text processing network and fine-tunes a language processing network for encoding emotional cues in audio.
(2) A novel feature fusion module CAF is developed. It enables feature dimensionality alignment and generates a more informative representation for emotion recognition.
(3) The proposed SER framework is validated on three databases and achieves promising performance.
Citation
The work is under review. If it is helpful, please cite
Yu S, Meng J, Fan W, Chen Y, Zhu B, Yu H, Xie Y, Sun Q. Speech emotion recognition using dual-stream representation and cross-attention fusion. Electronics. 2024 Jun 4;13(11):2191.
5 commits
Python
100.0%
Our contribution
(1) Dual-stream representation is designed. It fully trains a text processing network and fine-tunes a language processing network for encoding emotional cues in audio.
(2) A novel feature fusion module CAF is developed. It enables feature dimensionality alignment and generates a more informative representation for emotion recognition.
(3) The proposed SER framework is validated on three databases and achieves promising performance.
Citation
The work is under review. If it is helpful, please cite
Yu S, Meng J, Fan W, Chen Y, Zhu B, Yu H, Xie Y, Sun Q. Speech emotion recognition using dual-stream representation and cross-attention fusion. Electronics. 2024 Jun 4;13(11):2191.
5 commits
Python
100.0%