Journal article
Robust Dysarthric Speech Recognition with GAN Enhancement and LLM Correction
Advanced Intelligent Systems, Vol.8(2), pp.1-16
2026
Appears in UniSC Supported Open Access Outputs
Abstract
Dysarthric speech recognition faces significant challenges of acoustic variability and data scarcity, and this study proposes a robust system by integrating generative adversarial network enhancement and large language model correction to address these issues effectively. The system employs three key components, including a multimodal recognition core that combines whisper‐medium encoder with LoRA‐fine‐tuned Llama‐3.1‐8B for end‐to‐end acoustic‐to‐semantic mapping, an improved CycleGAN module that generates synthetic dysarthric speech through Inception‐ResNet fusion blocks, and an intelligent error correction mechanism using N‐best hypothesis reranking with semantic constraints. Experiments on the UA‐Speech dataset show that the complete system achieves a 20.61% word error rate representing a 73.9% relative improvement over traditional end‐to‐end transformer automatic speech recognition. Under very low intelligibility conditions it maintains a 48.69% word error rate demonstrating robust recognition for severe pathological speech. Ablation studies validate each module's effectiveness, providing significant advances for dysarthric patient communication technologies.
Details
- Title
- Robust Dysarthric Speech Recognition with GAN Enhancement and LLM Correction
- Authors
- Yibo He - Xi’an Jiaotong-Liverpool UniversityKah Phooi Seng - University of the Sunshine Coast, Queensland, School of Science, Technology and EngineeringChee Shen Lim - Xi’an Jiaotong-Liverpool UniversityLi Minn Ang (Corresponding Author) - University of the Sunshine Coast, Queensland, School of Science, Technology and Engineering
- Publication details
- Advanced Intelligent Systems, Vol.8(2), pp.1-16
- Publisher
- Wiley-VCH Verlag GmbH & Co. KGaA
- Date published
- 2026
- DOI
- 10.1002/aisy.202500873
- ISSN
- 2640-4567
- Copyright note
- © 2025 The Author(s). Advanced Intelligent Systems published by Wiley-VCH GmbH. This is an open access article under the terms of the Creative Commons Attribution License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited.
- Data Availability
- The dataset can be accessed through a public link: https://speechtechnology.web.illinois.edu/.
- Organisation Unit
- School of Science, Technology and Engineering; Engage Research Lab
- Language
- English
- Record Identifier
- 991179072402621
- Output Type
- Journal article
Metrics
3 File views/ downloads
26 Record Views
InCites Highlights
These are selected metrics from InCites Benchmarking & Analytics tool, related to this output
- Collaboration types
- Domestic collaboration
- International collaboration
- Web Of Science research areas
- Automation & Control Systems
- Computer Science, Artificial Intelligence
- Robotics
UN Sustainable Development Goals (SDGs)
This output has contributed to the advancement of the following goals:
Source: SDGs from InCites