LIS-Net: An end-to-end light interior search network for speech command recognition

作者:Nguyen Tuan Anh*; Hu, Yongjian; He, Qianhua; Tran Thi Ngoc Linh; Hoang Thi Kim Dung; Guang, Chen
来源:COMPUTER SPEECH AND LANGUAGE, 2021, 65: 101131.
DOI:10.1016/j.csl.2020.101131

摘要

With the rapid development of deep learning techniques, speech-based communication is getting more practically to be embedded into smart devices such as Alexa echo, TV, Fridge, etc. In this work, we have developed an efficient yet accurate Speech Command Recognition (SCR), that is particularly appropriate for low-resource devices. To this aim, a novel neural network, called Light Interior Search Network (LIS-Net), is presented that works with raw speech signal. LIS-Net is structurally composed of a sequence of parameterized LIS-Blocks, each of which is a stack of LIS-Cores, exploring the feature-map inheritance to learn highly distinctive and lightweight footprint of speech patterns. The proposed network is validated on Google Speech Commands benchmark speech datasets, demonstrating a significant improvement of accuracy and processing time in comparison with other state-of-the-art techniques.