Back to Search Start Over

U2-KWS: Unified Two-pass Open-vocabulary Keyword Spotting with Keyword Bias

Authors :
Zhang, Ao
Zhou, Pan
Huang, Kaixun
Zou, Yong
Liu, Ming
Xie, Lei
Publication Year :
2023

Abstract

Open-vocabulary keyword spotting (KWS), which allows users to customize keywords, has attracted increasingly more interest. However, existing methods based on acoustic models and post-processing train the acoustic model with ASR training criteria to model all phonemes, making the acoustic model under-optimized for the KWS task. To solve this problem, we propose a novel unified two-pass open-vocabulary KWS (U2-KWS) framework inspired by the two-pass ASR model U2. Specifically, we employ the CTC branch as the first stage model to detect potential keyword candidates and the decoder branch as the second stage model to validate candidates. In order to enhance any customized keywords, we redesign the U2 training procedure for U2-KWS and add keyword information by audio and text cross-attention into both branches. We perform experiments on our internal dataset and Aishell-1. The results show that U2-KWS can achieve a significant relative wake-up rate improvement of 41% compared to the traditional customized KWS systems when the false alarm rate is fixed to 0.5 times per hour.<br />Comment: Accepted by ASRU2023

Details

Database :
arXiv
Publication Type :
Report
Accession number :
edsarx.2312.09760
Document Type :
Working Paper