A Deployment-Oriented Evaluation of Hybrid Deep Learning and Traditional Signal Processing for Acoustic Echo Cancellation and Speech Enhancement
En cours de chargement...
Date
Authors
Nom de la revue
ISSN de la revue
Titre du volume
Éditeur
Université d'Ottawa / University of Ottawa
Résumé
Acoustic Echo Cancellation (AEC) is a key component of hands-free communication systems, enabling full-duplex operation by suppressing acoustic feedback between loudspeakers and microphones. Traditional adaptive filtering-based AEC systems remain widely used due to their efficiency and robustness but are limited by residual echo, nonlinear distortions, and background noise. Although hybrid deep learning-based enhancement methods have shown strong performance, they are often evaluated using idealized front-end configurations, making it difficult to isolate their contribution in real-time systems. This thesis presents a hybrid AEC framework that combines a production-oriented frequency-domain adaptive filter with a deep learning post-filter. The proposed model is derived from a convolutional recurrent architecture and redesigned for deployment using a Complex Ratio Mask (CRM)-based enhancement strategy, a simplified decoder, and Frequency-Temporal Long Short-Term Memory (F-T-LSTM) layers for temporal modelling. An identical linear echo cancellation front-end is used for both a traditional cascaded residual echo and noise suppression pipeline and the proposed neural system, enabling a controlled comparison. Evaluation is conducted using standardized ICASSP benchmark scenarios, including near-end single-talk, far-end single-talk, and double-talk conditions. Performance is assessed using perceptual metrics for echo suppression, speech distortion, and noise reduction, while Echo Return Loss Enhancement is used to evaluate suppression dynamics. Results show that the proposed system achieves comparable echo suppression to the traditional pipeline while improving speech preservation and overall perceptual quality. It also adapts more rapidly to changing acoustic conditions, whereas the baseline exhibits slower recovery. The proposed model achieved the highest perceptual evaluation scores, including an Echo Mean Opinion Score (MOS) of 4.53, a Degradation MOS of 3.55, and an overall perceptual MOS of 3.90. The system operates with 30 ms latency and a real-time factor of 0.68, demonstrating near real-time performance suitable for deployment.
Description
Mots-clés
Acoustic Echo Cancellation, DSP, Deep Learning, Speech Enhancement
