Evaluation and analysis of pooling methods for audio applications
Loading...
Date
2025
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Pooling plays a critical role in spectrogram-based audio analysis by determining how time–frequency information is compressed and preserved within deep learning models. However, existing evaluations of pooling methods primarily emphasize downstream task performance, such as classification accuracy, rather than directly assessing their ability to preserve key spectrogram features. Furthermore, existing pooling methods often fail to account for the non-stationary and multi-scale nature of audio signals. This study addresses these gaps by developing unified frameworks for understanding, evaluating, and advancing pooling mechanisms in audio analysis applications. We first propose a standardized evaluation framework using 17 metrics across four domains to assess pooling performance comprehensively. We then examine how specific pooling strategies affect feature retention across 15 audio application types. By identifying each application’s key spectrogram characteristics and grouping them accordingly, we evaluate 12 pooling methods and analyze which features each method preserves. This provides task-level insight into selecting the most effective pooling technique, enabling more granular evaluation, improved explainability, and more efficient audio analysis pipelines. Building upon these insights, the study introduces a learnable adaptive wavelet pooling framework that combines multiscale feature decomposition with dynamic subband weighting, enabling the network to flexibly adjust the preservation of local and global features. To comprehensively evaluate performance and interpretability, we consider three major audio domains; speech, music, and environmental sounds and further analyze specific categories within each domain such as content, speaker/source, semantics, environmental factors etc. These categories allow systematic examination of how adaptive pooling responds to different information requirements, highlighting which spectral and temporal features are most critical for each type of application. Experimental results across representative datasets show that the proposed method consistently outperforms conventional and adaptive pooling baselines, achieving higher accuracy across all domains.
Description
Citation
Bandara, E.M.S.J. (2025). Evaluation and analysis of pooling methods for audio applications [Master’s theses, University of Moratuwa]. Institutional Repository University of Moratuwa. https://dl.lib.uom.lk/handle/123/25495
