Fields · Physical Sciences · Computer Science · Artificial Intelligence

Speech Recognition and Synthesis

This cluster of papers focuses on the advances in speech recognition technology, covering topics such as acoustic modeling using deep neural networks, speaker verification, convolutional neural networks for speech recognition, end-to-end speech recognition systems, hidden Markov models, sequence-to-sequence models, automatic speech recognition, speaker diarization, and statistical language modeling.

97,620 works

Papers listed on taxonomy pages are the top few works per node from the OpenAlex snapshot. That list is not exhaustive and is not an endorsement. The topic map and the journal registry remain separate: there is still no authoritative topic-to-venue or topic-to-organization edge. Search is a lexical lookup, not a claim that a venue publishes a topic.

Most cited

  1. Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups

    Geoffrey E. Hinton, Li Deng, Dong Yu, George E. Dahl · 2012 · IEEE Signal Processing Magazine · 10,399 citations

  2. Framewise phoneme classification with bidirectional LSTM and other neural network architectures

    Alex Graves, Jürgen Schmidhuber · 2005 · Neural Networks · 5,619 citations

  3. Kaldi Speech Recognition Toolkit

    Daniel Povey · 2024 · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 4,895 citations

  4. An introduction to hidden Markov models

    L. R. Rabiner, Biing‐Hwang Juang · 1986 · IEEE ASSP Magazine · 4,812 citations

  5. Speaker Verification Using Adapted Gaussian Mixture Models

    Douglas A. Reynolds, Thomas F. Quatieri, R.B. Dunn · 2000 · Digital Signal Processing · 4,285 citations

  6. Front-End Factor Analysis for Speaker Verification

    Najim Dehak, Patrick Kenny, Réda Dehak, Pierre Dumouchel · 2010 · IEEE Transactions on Audio Speech and Language Processing · 3,605 citations

Most recent

  1. Cross-task generative augmentation of multimodal speech features for task-incomplete MCI detection

    Hao Li, Bo Wang, Qiang He, Jun Mou · 2026 · Pattern Recognition · 0 citations

  2. Learning Dense Multimodal Correspondence on Modest Hardware

    Noah Sinclair, Lobry Hsu, Ava Richardson, Liam Carter · 2026 · 0 citations

  3. Utilising speech-derived biomarkers to detect Alzheimer's disease with BERT-based language models: a machine learning study

    Zara Khanna, Dean Ho, Alexandria Remus, Marlena Raczkowska · 2026 · Frontiers in Artificial Intelligence · 0 citations

  4. Dynamic adaptive ensemble decoding under uncertainty for large language models

    Weitao Ma, Xiaocheng Feng, Xiaocheng Feng, Yichong Huang · 2026 · Knowledge-Based Systems · 0 citations

  5. Generative Fusion Decoding with MambaByte for Low-resource Regional Language Speech Recognition: A Case Study on Javanese and Sundanese

    Agung Santosa, Asril Jarin, Lyla Ruslana Aini, Eko Mulyanto Yuniarno · 2026 · International journal of intelligent engineering and systems · 0 citations

  6. Paralinguistic acoustic feature to natural language prompt conversion using hybrid transformer–LLaMA architecture

    Junyoung Kim, Ahyoung Choi · 2026 · Computer Speech & Language · 0 citations

Search papers

Keywords