基于离散小波变换和奇异值分解的沿层地震多属性U-Net断层识别模型

Research on layer-wise seismic multi-attribute fault recognition model based on discrete wavelet transform and singular value decomposition with U-Net

  • 摘要: 基于地震数据的机器学习断层识别中,复杂构造干扰与强噪声易导致微小断层漏判或将噪声误判为断层。针对该问题,以云南矿区为研究区,提出基于沿层地震属性引导的离散小波变换−奇异值分解−U-Net (Discrete Wavelet Transform−Singular Value Decomposition−U-Net,DSU)断层识别模型。该模型的核心特征如下:一是针对噪声干扰与微小断层响应微弱的问题,选取兼具正交性、紧支撑性与高消失矩的Daubechies 4 (Db4)、Symlet 4 (Sym4)、Coiflet 3 (Coif3) 3种母小波作为候选,通过快速傅里叶变换(Fast Fourier Transform,FFT)与离散余弦变换(Discrete Cosine Transform,DCT)分别从频域和空域计算断层信号与非断层信号的能量差异,构建频域能量差异指标 CFFT 和空域能量差异指标 CDCT 作为母小波优选的双重量化判据。经对比,Db4母小波在云南矿区的 CFFT 值达0.0358、CDCT值达2.4635,断点信噪比较原始数据提升35%,均优于Sym4和Coif3,故选定为最优母小波进行离散小波变换(Discrete Wavelet Transform,DWT)。在压制强噪声的同时,进一步沿目的层以5×5滑动窗提取地震属性,有效降低非目的层噪声对断层信号的干扰。二是针对训练数据集中20种沿层地震属性间存在显著相关性与信息冗余的问题,采用奇异值分解(Singular Value Decomposition,SVD)进行降维:基于硬阈值法确定最佳分解维度,依据累积能量贡献率不低于90%的准则,将属性维度从20维降至6维,优选出原始振幅、均方根振幅、混沌体、倾角偏差、方差和瞬时相位6种核心属性,在保留主要特征信息的前提下有效降低了高相关性属性引发的模型过拟合风险。在U-Net分类网络中,针对断层像素与非断层像素数量悬殊的类别不平衡问题,引入加权交叉熵损失函数,按2类样本数量的反比设置类别权重,使模型对断层样本给予更高的误分类惩罚。以研究区内断层落差较大的A区数据构建训练集,以中小型断层为主的B区数据构建测试集,设计U-Net、DWT−U-Net和DSU 3组对照试验。结果表明:DSU模型准确率达0.93131、查准率达0.95341、查全率达0.90695F1分数达0.9296,各项指标较仅使用U-Net均有显著提升。模型在B区对落差小于5 m的小断层仍能保持有效的识别能力,F67小断层经巷道揭露验证与预测结果一致。进一步将模型输出转换为后验概率值,实现断层识别结果的可靠性量化评价,为解释人员提供了直观的断层发育参考。

     

    Abstract: In machine learning-based fault identification using seismic data, complex structural interference and strong noise may easily lead to missed detection of small faults or misclassification of noise as faults. To address this problem, a fault identification model guided by along-layer seismic attributes, termed discrete wavelet transform−singular value decomposition−U-Net (DSU), was proposed with the Yunnan mining area as the study area. The key features of the model are as follows. First, to tackle the problem of noise interference and weak response of small faults, three mother wavelets with orthogonality, compact support, and high vanishing moments—Daubechies 4 (Db4), Symlet 4 (Sym4), and Coiflet 3 (Coif3)—were selected as candidates. Fast Fourier Transform (FFT) and Discrete Cosine Transform (DCT) were employed to compute the energy differences between fault and non-fault signals in the frequency and spatial domains, respectively, yielding two quantitative metrics for mother wavelet selection: the frequency-domain energy difference index CFFT and the spatial-domain energy difference index CDCT. The comparison showed that the Db4 mother wavelet achieved a CFFT value of 0.0358, a CDCT value of 2.4635, and a 35% improvement in the signal-to-noise ratio at fault locations over the original data in the Yunnan mining area, all of which were superior to those of Sym4 and Coif3. Therefore, Db4 was selected as the optimal mother wavelet for Discrete Wavelet Transform (DWT). While suppressing strong noise, seismic attributes were further extracted along the target layer using a 5 × 5 sliding window, effectively reducing the interference of non-target-layer noise on fault signals. Second, to address the significant correlation and information redundancy among the 20 along-layer seismic attributes in the training dataset, Singular Value Decomposition (SVD) was employed for dimensionality reduction. The optimal decomposition dimension was determined using the hard threshold method based on the criterion that the cumulative energy contribution rate should be no less than 90%, thereby reducing the attribute dimension from 20 to 6. Six core attributes were selected: original amplitude, root mean square amplitude, chaos volume, dip deviation, variance, and instantaneous phase. This approach effectively mitigated the risk of model overfitting caused by highly correlated attributes while retaining the majority of feature information. In U-Net classification network, to address the class imbalance problem arising from the substantial disparity between fault and non-fault pixel counts, a weighted cross-entropy loss function was introduced, in which the class weights were set inversely proportional to the sample sizes of the two classes, thereby imposing a higher misclassification penalty on fault samples. A training dataset was constructed from Area A, which is dominated by large-throw faults, and a testing dataset from Area B, which is dominated by small- and medium-throw faults. Three comparative experiments were designed using U-Net, DWT−U-Net, and DSU models. The results show that the DSU model achieves an accuracy of 0.93131, a precision of 0.95341, a recall of 0.90695, and an F1 score of 0.9296, all of which represent significant improvements over the U-Net-only baseline. The model maintains effective identification capability for small faults with throws of less than 5 m in Area B; the small fault F67 was confirmed by roadway exposure data to be consistent with the predicted result. Furthermore, the model output was converted into posterior probability values, enabling a quantitative reliability assessment of the fault identification results and providing interpreters with an intuitive reference map of fault development probability.

     

/

返回文章
返回