Back to Search Start Over

Speech emotion recognition based on hierarchical attributes using feature nets.

Authors :
Zhao, Huijuan
Ye, Ning
Wang, Ruchuan
Source :
International Journal of Parallel, Emergent & Distributed Systems. May2020, Vol. 35 Issue 3, p354-364. 11p.
Publication Year :
2020

Abstract

Speech emotion recognition is a challenging topic and has many important applications in our real life, especially in terms of human-computer interaction. Traditional methods are based on the pipeline of pre-processing, feature extraction, dimensionality reduction and emotion classification. Previous studies have focussed on emotion recognition based on two different models: discrete model and continuous model. Both the speaker's age and gender affect the speech emotion recognition in the two models. Moreover, investigation results shown that the dimensional attributes of emotion such as arousal, valence and dominance are related to each other. Based on these observations, we propose a new attributes recognition model using Feature Nets, aims to improve the emotion recognition performance and generalisation capabilities. The method utilises the corpus to train the age and gender classification model, which will be transferred to the main model: a hierarchical deep learning model, using age and gender as the high level attributes of the emotion. The public databases EMO-DB and IEMOCAP have been conducted to evaluate the performance both in the classification task and regression task. Experiment results show that the proposed approach based on attributes transferring can improve the recognition accuracy, no matter transferring age or gender. [ABSTRACT FROM AUTHOR]

Details

Language :
English
ISSN :
17445760
Volume :
35
Issue :
3
Database :
Academic Search Index
Journal :
International Journal of Parallel, Emergent & Distributed Systems
Publication Type :
Academic Journal
Accession number :
143635859
Full Text :
https://doi.org/10.1080/17445760.2019.1626854