Back to Search
Start Over
TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
- Publication Year :
- 2024
-
Abstract
- This paper presents a novel deep neural network-based architecture tailored for Speech Emotion Recognition (SER). The architecture capitalises on dense interconnections among multiple layers of bidirectional dilated convolutions. A linear kernel dynamically fuses the outputs of these layers to yield the final emotion class prediction. This innovative architecture is denoted as TBDM-Net: Temporally-Aware Bi-directional Dense Multi-Scale Network. We conduct a comprehensive performance evaluation of TBDM-Net, including an ablation study, across six widely-acknowledged SER datasets for unimodal speech emotion recognition. Additionally, we explore the influence of gender-informed emotion prediction by appending either golden or predicted gender labels to the architecture's inputs or predictions. The implementation of TBDM-Net is accessible at: https://github.com/adrianastan/tbdm-net<br />Comment: In Proceedings of 2024 IEEE International Workshop on Machine Learning for Signal Processing, London, UK
Details
- Database :
- arXiv
- Publication Type :
- Report
- Accession number :
- edsarx.2409.10056
- Document Type :
- Working Paper