A survey on end-to-end speech recognition systems
Abstract
Interest towards end to end SCR systems over the past few years have been growing because the system designs are more integrated and the methods have lower error rates than previous techniques. This survey also seeks to deliver a comprehensive understanding about these systems, analyse the improvement, status and potential developments in this field. Integrated systems avoid the lengthy process of converting sound to script by directly mapping the acoustic signal to the text string without the need for elaborate feature extraction. Some of the central sections that are discussed focuses towards the “Recurrent Neural Network (RNN)”, “Convolutional Neural Network (CNN)”, and transformer that have played crucial role in understanding optimization and advancement of accuracy for the recognition of speech. The survey also specifically discusses how these benchmarks datasets and evaluation methods are used to compare performance of these systems and offers information about these benchmarks. Some of the limitations that have been highlighted include; accent variation, noise and computational complexity with regard to the approaches used to address such concerns. Furthermore, the study delves into the implementation of end-to-end speech recognition systems in many fields, specifically focusing on the practical use of these systems in healthcare, telecommunications, and customer service industries. In order to explicate the path of “end-to-end systems of speech recognition” and their ability to transform the way man interacts with machines, this survey brings a review of the present literature and advancements.
How this paper connects to the literature. Drag to explore, click any node to open that paper.
