Speaker Localization Based on Audio-Visual Bimodal Fusion.

Authors :: Zhu, Ying-Xin
Jin, Hao-Ran
Source :: Journal of Advanced Computational Intelligence & Intelligent Informatics. May2021, Vol. 25 Issue 3, p375-382. 8p.
Publication Year :: 2021
Abstract: The demand for fluency in human–computer interaction is on an increase globally; thus, the active localization of the speaker by the machine has become a problem worth exploring. Considering that the stability and accuracy of the single-mode localization method are low, while the multi-mode localization method can utilize the redundancy of information to improve accuracy and anti-interference, a speaker localization method based on voice and image multimodal fusion is proposed. First, the voice localization method based on time differences of arrival (TDOA) in a microphone array and the face detection method based on the AdaBoost algorithm are presented herein. Second, a multimodal fusion method based on spatiotemporal fusion of speech and image is proposed, and it uses a coordinate system converter and frame rate tracker. The proposed method was tested by positioning the speaker stand at 15 different points, and each point was tested 50 times. The experimental results demonstrate that there is a high accuracy when the speaker stands in front of the positioning system within a certain range. [ABSTRACT FROM AUTHOR]

Subjects :: *HUMAN-computer interaction
*SPATIOTEMPORAL processes
*ALGORITHMS
*MICROPHONE arrays
*MICROPOSITIONING systems

Language :: English
ISSN :: 13430130
Volume :: 25
Issue :: 3
Database :: Academic Search Index
Journal :: Journal of Advanced Computational Intelligence & Intelligent Informatics
Publication Type :: Academic Journal
Accession number :: 150389113
Full Text :: https://doi.org/10.20965/jaciii.2021.p0375