| 91 | 0 | 43 |
| 下载次数 | 被引频次 | 阅读次数 |
针对基因组数据高维度、高冗余及复杂非线性关系所导致的特征筛选困难和预测性能不足问题,本研究提出了一种基于沙普利加性解释(Shapley additive explanations, SHAP)引导的两阶段特征筛选与基因组预测框架。首先,对全体样本的基因型数据进行主成分分析(principal component analysis, PCA)降维,并保留95%的累计变异信息,以构建统一的低维特征表示;其次,在外层10折交叉验证框架下,于每一折训练集内部采用最小绝对收缩和选择算子(least absolute shrinkage and selection operator, LASSO)进行第一阶段全局稀疏筛选,以压缩候选特征空间;在此基础上,结合极端梯度提升(extreme gradient boosting, XGBoost)与SHAP对候选特征进行条件贡献评估,进一步筛选关键特征;最后,将筛选后的特征输入深度神经网络基因组预测(deep neural network genomic prediction, DNNGP)模型进行训练,并在对应测试集上完成独立预测评估。在wheat599和wheat2000数据集上的实验结果表明,所提出的方法在多个环境条件下优于基因组最佳线性无偏预测(genomic best linear unbiased prediction, GBLUP)、轻量级梯度提升机(light gradient boosting machine, Light GBM)、支持向量回归(support vector regression, SVR)和DNNGP等模型,在wheat599 Env2中,预测相关系数达到0.78,较DNNGP提高约23.8%。同时,在特征数量减少(最高约88%)的情况下,模型仍保持较高预测精度,展现了良好的稳定性与泛化能力。此外, SHAP分析显示,部分特征在不同环境中具有稳定的重要性,揭示了潜在的基因型与环境互作效应。综上,本文方法不仅提升了预测性能,还增强了模型可解释性,为高维基因组数据分析与作物分子育种提供了一种有效的技术路径。
Abstract:To address the problems of difficulty in feature selection and insufficient prediction performance caused by the high dimensionality, high redundancy, and complex nonlinear relationships of genomic data, this paper proposes a two-stage feature selection and genomic prediction framework based on Shapley additive explanations( SHAP). First, principal component analysis( PCA) was performed to reduce the dimensionality of the genotypic data of all samples while retaining 95% of the cumulative variation, thereby constructing a unified low-dimensional feature representation. Then, the least absolute shrinkage and selection operator( LASSO) was applied for the first-stage global sparse selection within the training set of each fold under an outer 10-fold cross-validation framework, reducing the candidate feature space. Extreme gradient boosting( XGBoost) combined with SHAP was used to evaluate the conditional contributions of candidate features and further select key features. The selected features were input into the deep neural network genomic prediction( DNNGP) model for training, and independent prediction evaluation was conducted on the corresponding test set. Experimental results on the wheat599 and wheat2000 datasets showed that the proposed method outperformed models including genomic best linear unbiased prediction( GBLUP), light gradient boosting machine( Light GBM), support vector regression( SVR), and DNNGP under multiple environmental conditions. For instance, in Env2 of the wheat599 dataset, the prediction correlation coefficient reached 0. 78, representing an approximately 23. 8% improvement over DNNGP. The model maintained high accuracy even with up to an approximately 88% reduction in the number of features, demonstrating good stability and generalization. Additionally, SHAP analysis revealed that certain features showed stable importance across different environments, highlighting potential genotype-by-environment interaction effects. In conclusion, the proposed method not only improved prediction performance but also enhanced model interpretability, offering a valuable approach for high-dimensional genomic data analysis and crop molecular breeding.
Abdi H,Alipour H,Bernousi I,et al.,2023.Identification of novel putative alleles related to important agronomic traits of wheat using robust strategies in GWAS[J].Sci.Rep.,13(1):9927.
Abdollahi-Arpanahi R,Gianola D,Peñagaricano F,2020.Deep learning versus parametric and ensemble methods for genomic prediction of complex phenotypes[J].Genet.Sel.Evol.,52(1):12.
Al-Sabri E H A,Shah A A,Iqbal K,et al.,2025.Enhancing machine learning performance through statistical feature selection in high-dimensional genomic and financial data[J].Int.J.Appl.Math.,38(12s):698-715.
Bauer A M,Reetz T C,Léon J,2006.Estimation of breeding values of inbred lines using best linear unbiased prediction (BLUP) and genetic similarities[J].Crop Sci.,46(6):2685-2691.
Bonilla-Huerta E,Hernández-Montiel A,Caporal R M,et al.,2016.Hybrid framework using multiple-filters and an embedded approach for an efficient selection and classification of microarray data[J].IEEE/ACM Trans.Comput.Biol.Bioinform.,13(1):12-26.
Chen C S,Powell O,Dinglasan E,et al.,2023.Genomic prediction with machine learning in sugarcane,a complex highly polyploid clonally propagated crop with substantial non-additive variation for key traits[J].Plant Genome,16(4):e20390.
Chiang L H,Pell R J,2004.Genetic algorithms combined with discriminant analysis for key variable identification[J].J.Process Control,14(2):143-155.
Clark S A,van der Werf J,2013.Genomic best linear unbiased prediction (g BLUP) for the estimation of genomic breeding values[M]//Gondro C,van der Werf J,Hayes B,et al.,Genome-wide association studies and genomic prediction.Totowa,NJ:Humana Press:321-330.
Crossa J,Campos G D E L,Pérez P,et al.,2010.Prediction of genetic values of quantitative traits in plant breeding using pedigree and molecular markers[J].Genetics,186 (2):713-724.
Crossa J,Jarquín D,Franco J,et al.,2016.Genomic prediction of gene bank wheat landraces[J].G3 (Bethesda),6(7):1819-1834.
Debelee T G,Kebede S R,Waldamichael F G,et al.,2023.Wheat yield prediction using machine learning:a survey[C]//Debelee T G,Ibenthal A,Schwenker F,et al.,PanAfrican Conference on Artificial Intelligence.Cham:Springer:114-132.
Dhar T,Dey N,Borra S,et al.,2023.Challenges of deep learning in medical image analysis:improving explainability and trust[J].IEEE Trans.Technol.Soc.,4(1):68-75.
Duan M M,Li N,Li S Y,2014.A modified LQG benchmark for economic performance assessment of model predictive control[C]//Proceedings of the 33rd Chinese Control Conference.IEEE:7811-7816.
Elsharawy H,Refat M,2023.CRISPR/Cas9 genome editing in wheat:enhancing quality and productivity for global food security-a review[J].Funct.Integr.Genom.,23 (3):265.
Hira Z M,Gillies D F,2015.A review of feature selection and feature extraction methods applied on microarray data[J].Adv.Bioinform.,2015:198363.
Jung Y,2018.Multiple predicting K-fold cross-validation for model selection[J].J.Nonparam.Stat.,30(1):197-215.
Krassowski M,Das V,Sahu S K,et al.,2020.State of the field in multi-omics research:from computational needs to data mining and sharing[J].Front.Genet.,11:610798.
Kuzudisli C,Bakir-Gungor B,Bulut N,et al.,2023.Review of feature selection approaches based on grouping of features[J].Peer J,11:e15666.
Ma W L,Qiu Z X,Song J,et al.,2018.A deep convolutional neural network approach for predicting phenotypes from genotypes[J].Planta,248(5):1307-1318.
Mc Laren C G,Bruskiewich R M,Portugal A M,et al.,2005.The international rice information system.a platform for meta-analysis of rice crop data[J].Plant Physiol.,139(2):637-642.
Mercado Rueda A P,2023.Analysis of variance:ANOVA[M]//Bakal J A,De Froda S,Owens B D,et al.,Translational Sports Medicine.London:Academic Press:157-160.
Meuwissen T H,Hayes B J,Goddard M E,2001.Prediction of total genetic value using genome-wide dense marker maps[J].Genetics,157(4):1819-1829.
Mulugeta B,Tesfaye K,Ortiz R,et al.,2023.Marker-trait association analyses revealed major novel QTLs for grain yield and related traits in durum wheat[J].Front.Plant Sci.,13:1009244.
Prechelt L,2012.Early stopping:but when?[M]//Montavon G,Orr G B,Müller K R,et al.,Neural Networks:Tricks of the Trade:Second Edition.Berlin,Heidelberg:Springer:53-67.
Ras G,Xie N,Van Gerven M,et al.,2022.Explainable deep learning:a field guide for the uninitiated[J].J.Artif.Intell.Res.,73:329-397.
Shaheenuzzamn M,Liu T X,Shi S D,et al.,2020.Development of sequencing technology and role of next generation sequencing technologies in wheat research:a review[J].Pak.J.Bot.,52(5):1867-1878.
Tibshirani R,1996.Regression shrinkage and selection via the Lasso[J].J.R.Stat.Soc.Ser.B Stat.Methodol.,58(1):267-288.
Wang J B,Zhang J,Hao W J,et al.,2025.An interpretable integrated machine learning framework for genomic selection[J].Smart Agric.Technol.,12:101138.
Wang K L,Abid M A,Rasheed A,et al.,2023.DNNGP,a deep neural network-based method for genomic prediction using multi-omics data in plants[J].Mol.Plant,16(1):279-293.
Yang H,Song J W,Qiao C B,et al.,2023.Genome-wide association studies of salt-tolerance-related traits in rice at the seedling stage using In Del markers developed by the genome re-sequencing of Japonica rice accessions[J].Agriculture,13(8):1573.
Yousef M,Kumar A,Bakir-Gungor B,2020.Application of biological domain knowledge based feature selection on gene expression data[J].Entropy,23(1):2.
Zhu T T,Wang L,Rimbert H,et al.,2021.Optical maps refine the bread wheat Triticum aestivum cv.Chinese Spring genome assembly[J].Plant J.,107(1):303-314.
基本信息:
DOI:10.13417/j.gab.045.000856
中图分类号:TP18;Q811.4
引用信息:
[1]卢开心,刘建平.基于SHAP引导的两阶段特征筛选与基因组预测框架研究[J].基因组学与应用生物学,2026,45(04):856-872.DOI:10.13417/j.gab.045.000856.
基金信息:
国家自然科学基金(32460444); 宁夏全职引进高层次人才科研启动项目(2024BEH04130); 宁夏自然科学基金(2025AAC050001); 北方民族大学研究生创新项目(CYX25216)共同资助
2026-06-17
2026-06-17
2026-06-17