基于改进EfficientNet的煤矸音频分类方法_中国煤炭行业知识服务平台

基于改进EfficientNet的煤矸音频分类方法

Title

Coal gangue audio classification method based on improved EfficientNet
作者

宋庆军焦守悦姜海燕宋庆辉郝文超
Author

SONG Qingjun;JIAO Shouyue;JIANG Haiyan;SONG Qinghui;HAO Wenchao
单位

山东科技大学智能装备学院
Organization

School of Intelligent Equipment, Shandong University of Science and Technology
摘要

针对煤矸音频特征提取过程中设备运行噪声干扰严重及单一提取方法易导致信息丢失的问题，提出了一种基于改进EfficientNet的煤矸音频分类方法。采用基于Mel频谱和Gammatone倒谱系数的特征提取方法，有效捕捉矸石声音中的低频信息和细节特征。选择EfficientNet−B0作为骨干网络，并对其进行以下改进：将原有的多尺度通道注意力模块换成卷积块注意力模块，得到卷积注意力特征融合（CAFF）模块，通过网络自学习为不同空间位置的特征分配不同的权重信息，生成新的有效特征；在原有的MBConv模块中并行嵌入频域通道注意力（FCA）模块，加强特征图的表达能力，从而提高整个网络的性能。实验结果表明：引入CAFF模块后，模型准确率提升了0.61%，F₁得分提升了0.52%，且模型收敛更快，说明CAFF模块有效提升了模型对频谱特征的捕捉能力；引入FCA模块后，准确率提升了0.45%，F₁得分提升了0.62%，说明模块的叠加可以进一步提高模型的泛化能力和处理复杂特征的能力；改进EfficientNe模型的准确率为91.90%，标准差为0.108，显著优于同类对比音频分类模型。
Abstract

To address the issues of severe interference of equipment operating noise and information loss caused by single extraction methods during coal gangue audio feature extraction, a coal gangue audio classification method based on improved EfficientNet is proposed. The method adopted a feature extraction approach combining Mel spectrogram and Gammatone frequency cepstral coefficients to effectively capture low-frequency information and detailed features in gangue audio. EfficientNet-B0 was selected as the backbone network, and the following improvements were made: the original multi-scale channel attention module was replaced with a convolutional block attention module, resulting in the Convolutional Attention Feature Fusion (CAFF) module. This module allowed the network to autonomously assign different weight information to features in different spatial positions, generating new effective features. Additionally, a Frequency-domain Channel Attention (FCA) module was embedded in parallel within the original MBConv module, strengthening the representation ability of feature maps and thereby improving overall network performance. The experimental results demonstrated that after introducing the CAFF module, the model's accuracy improved by 0.61%, the F₁ score increased by 0.52%, and convergence was faster, indicating that the CAFF module effectively enhanced the model's ability to capture spectral features. After integrating the FCA module, accuracy improved by 0.45%, and the F₁ score increased by 0.62%, showing that combining these modules further enhanced the model's generalization ability and its ability to process complex features. The improved EfficientNet model achieved an accuracy of 91.90%, with a standard deviation of 0.108, significantly outperforming other comparable audio classification models.
关键词

综放开采煤矸识别音频特征提取EfficientNetMel频谱特征Gammatone倒谱系数注意力机制
KeyWords

comprehensive mining;coal gangue recognition;audio feature extraction;EfficientNet;Mel spectrogram feature;Gammatone frequency cepstral coefficient;attention mechanism
基金项目(Foundation)

国家自然科学基金面上项目(52174145)；山东省科技型中小企业创新能力提升工程项目（2022TSGC1271，2023TSGC0620）。
DOI

10.13272/j.issn.1671-251x.2024090013
引用格式

宋庆军，焦守悦，姜海燕，等. 基于改进EfficientNet的煤矸音频分类方法[J]. 工矿自动化，2025，51（1）：138-144.
Citation

SONG Qingjun, JIAO Shouyue, JIANG Haiyan, et al. Coal gangue audio classification method based on improved EfficientNet[J]. Journal of Mine Automation，2025，51（1）：138-144.
图表
图(8) / 表(3)

煤问提

煤传媒

煤视界

科技创新50强

会员中心