Speech Recognition and Synthesis
This cluster of papers focuses on the advances in speech recognition technology, covering topics such as acoustic modeling using deep neural networks, speaker verification, convolutional neural networks for speech recognition, end-to-end speech recognition systems, hidden Markov models, sequence-to-sequence models, automatic speech recognition, speaker diarization, and statistical language modeling.
Papers listed on taxonomy pages are the top few works per node from the OpenAlex snapshot. That list is not exhaustive and is not an endorsement. The topic map and the journal registry remain separate: there is still no authoritative topic-to-venue or topic-to-organization edge. Search is a lexical lookup, not a claim that a venue publishes a topic.
Most cited
- Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
- Framewise phoneme classification with bidirectional LSTM and other neural network architectures
- Kaldi Speech Recognition Toolkit
- An introduction to hidden Markov models
- Speaker Verification Using Adapted Gaussian Mixture Models
- Front-End Factor Analysis for Speaker Verification
Most recent
- Cross-task generative augmentation of multimodal speech features for task-incomplete MCI detection
- Learning Dense Multimodal Correspondence on Modest Hardware
- Utilising speech-derived biomarkers to detect Alzheimer's disease with BERT-based language models: a machine learning study
- Dynamic adaptive ensemble decoding under uncertainty for large language models
- Generative Fusion Decoding with MambaByte for Low-resource Regional Language Speech Recognition: A Case Study on Javanese and Sundanese
- Paralinguistic acoustic feature to natural language prompt conversion using hybrid transformer–LLaMA architecture