Back to Search Start Over

TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition

Authors :
Striletchi, Vlad
Striletchi, Cosmin
Stan, Adriana
Publication Year :
2024

Abstract

This paper presents a novel deep neural network-based architecture tailored for Speech Emotion Recognition (SER). The architecture capitalises on dense interconnections among multiple layers of bidirectional dilated convolutions. A linear kernel dynamically fuses the outputs of these layers to yield the final emotion class prediction. This innovative architecture is denoted as TBDM-Net: Temporally-Aware Bi-directional Dense Multi-Scale Network. We conduct a comprehensive performance evaluation of TBDM-Net, including an ablation study, across six widely-acknowledged SER datasets for unimodal speech emotion recognition. Additionally, we explore the influence of gender-informed emotion prediction by appending either golden or predicted gender labels to the architecture's inputs or predictions. The implementation of TBDM-Net is accessible at: https://github.com/adrianastan/tbdm-net<br />Comment: In Proceedings of 2024 IEEE International Workshop on Machine Learning for Signal Processing, London, UK

Details

Database :
arXiv
Publication Type :
Report
Accession number :
edsarx.2409.10056
Document Type :
Working Paper