background

Articles

Multi-Input Deep Complex Convolution Recurrent Network for Joint Acoustic Echo Cancellation and Background Noise Suppression in Real- Time Speech Communication

Abstract

The proliferation of online communication platforms has intensified demand for high-quality speech enhancement in adverse acoustic environments. This paper presents a Multi-Input Deep Complex Convolution Recurrent Network (MIDCCRN) that extends conventional DCCRN architectures to handle multiple input streams through temporal concatenation, enabling unified processing of microphone signals, far-end references, and intermediate estimates. We propose a computationally efficient Complex LSTM implementation utilizing complex convolution operations, which reduces activation function calls by 50% while maintaining equivalent representational capacity. Building upon these innovations, we introduce a three-stage cascaded framework for joint acoustic echo cancellation (AEC) and background noise suppression (BNS) comprising echo estimation, acoustic echo cancellation, and background noise suppression modules, trained end-to-end with multi-task learning and a novel short-time weighted signal-to-distortion ratio (STwSDR) loss function. Comprehensive experiments on the ICASSP 2022 AEC Challenge dataset demonstrate state-of-the-art performance with a final score M = 0.85, comparable to the challenge-winning solution, while maintaining real-time capability (RTF = 0.302) and modest model size (4.28M parameters). Ablation studies validate the effectiveness of each proposed component.