Evaluation and analysis of pooling methods for audio applications

dc.contributor.advisorUthayasanker, T
dc.contributor.authorBandara, EMSJ
dc.date.accept2025
dc.date.accessioned2026-08-18T09:38:39Z
dc.date.issued2025
dc.description.abstractPooling plays a critical role in spectrogram-based audio analysis by determining how time–frequency information is compressed and preserved within deep learning models. However, existing evaluations of pooling methods primarily emphasize downstream task performance, such as classification accuracy, rather than directly assessing their ability to preserve key spectrogram features. Furthermore, existing pooling methods often fail to account for the non-stationary and multi-scale nature of audio signals. This study addresses these gaps by developing unified frameworks for understanding, evaluating, and advancing pooling mechanisms in audio analysis applications. We first propose a standardized evaluation framework using 17 metrics across four domains to assess pooling performance comprehensively. We then examine how specific pooling strategies affect feature retention across 15 audio application types. By identifying each application’s key spectrogram characteristics and grouping them accordingly, we evaluate 12 pooling methods and analyze which features each method preserves. This provides task-level insight into selecting the most effective pooling technique, enabling more granular evaluation, improved explainability, and more efficient audio analysis pipelines. Building upon these insights, the study introduces a learnable adaptive wavelet pooling framework that combines multiscale feature decomposition with dynamic subband weighting, enabling the network to flexibly adjust the preservation of local and global features. To comprehensively evaluate performance and interpretability, we consider three major audio domains; speech, music, and environmental sounds and further analyze specific categories within each domain such as content, speaker/source, semantics, environmental factors etc. These categories allow systematic examination of how adaptive pooling responds to different information requirements, highlighting which spectral and temporal features are most critical for each type of application. Experimental results across representative datasets show that the proposed method consistently outperforms conventional and adaptive pooling baselines, achieving higher accuracy across all domains.
dc.identifier.accnoTH6183
dc.identifier.citationBandara, E.M.S.J. (2025). Evaluation and analysis of pooling methods for audio applications [Master’s theses, University of Moratuwa]. Institutional Repository University of Moratuwa. https://dl.lib.uom.lk/handle/123/25495
dc.identifier.degreeMSc (Major Component Research)
dc.identifier.departmentDepartment of Computer Science
dc.identifier.facultyEngineering
dc.identifier.urihttps://dl.lib.uom.lk/handle/123/25495
dc.language.isoen
dc.subjectDEEP LEARNING-Audio Analysis-Pooling
dc.subjectACOUSTICS-Spectrograms
dc.subjectDEEP NEURAL NETWORKS-Wavelet Pooling
dc.subjectMSc (MAJOR COMPONENT RESEARCH)-Dissertations
dc.subjectCOMPUTER SCIENCE AND ENGINEERING-Dissertations
dc.subjectMSc (Major Component Research)
dc.titleEvaluation and analysis of pooling methods for audio applications
dc.typeThesis-Abstract

Files

Original bundle

Now showing 1 - 3 of 3
Loading...
Thumbnail Image
Name:
TH6183-1.pdf
Size:
77.49 KB
Format:
Adobe Portable Document Format
Description:
Pre-text
Loading...
Thumbnail Image
Name:
TH6183-2.pdf
Size:
54.5 KB
Format:
Adobe Portable Document Format
Description:
Post-text
Loading...
Thumbnail Image
Name:
TH6183.pdf
Size:
805.15 KB
Format:
Adobe Portable Document Format
Description:
Full-thesis

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: