Bioinformatics
Mahboubeh Ayoubi; Babak Teimourpour; Mostafa Akhavan-Safar
Abstract
Background and Objectives: Identifying and classifying cancer-driving genes by analyzing their complex relationships within gene regulatory networks (GRN) can significantly aid in the development, progression, and discovery of targeted cancer therapies. The cancer-driving genes are responsible for tumorigenesis ...
Read More
Background and Objectives: Identifying and classifying cancer-driving genes by analyzing their complex relationships within gene regulatory networks (GRN) can significantly aid in the development, progression, and discovery of targeted cancer therapies. The cancer-driving genes are responsible for tumorigenesis and disease progression. However, the current methods frequently concentrate on network rebuilding, which restricts their capacity to identify regulatory linkages. By utilizing the structural and functional characteristics of GRNs, this work seeks to create a strong graph-based framework for precise cancer driver gene classification.Methods: Network and graph-based methodologies are employed to analyze these complex gene networks. Using graph neural networks (GNN), complex intergenic patterns can be identified in genetic and cellular data. In this study, a GNN-based framework is proposed to classify genes in gene regulatory networks in order to improve the detection of cancer-driving genes. The proposed graph-based framework effectively integrates multi-omics data and mitigates class imbalance through an artificial oversampling strategy. The proposed GNN-based framework facilitates the modeling of both topological structure and feature information within gene interaction networks. Additionally, it addresses the challenges of class imbalance between driver and non-driver genes through the implementation of the GraphSMOTE technique.Results: To construct the gene regulatory graph, the Regetworks regulatory dataset was combined with three gene expression datasets related to breast, lung, and colon cancers. The results demonstrate that the proposed model consistently attains robust classification performance, with AUC-ROC scores exceeding 0.77 in all cases and F1 scores above 0.70, outperforming previous network-based methods.Conclusion: The evaluation criteria show that the proposed model has a high ability to generalize tumor types with differences in network topology and class imbalance.
Bioinformatics
M. Akhavan-Safar; B. Teimourpour; M. Ayyoubi
Abstract
Background and Objectives: One of the important topics in oncology treatment and prevention is the identification of genes that initiate cancer in cells. These genes are known as cancer driver genes (CDGs). Identification of the CDGs is important both for a basic understanding of cancer and to help find ...
Read More
Background and Objectives: One of the important topics in oncology treatment and prevention is the identification of genes that initiate cancer in cells. These genes are known as cancer driver genes (CDGs). Identification of the CDGs is important both for a basic understanding of cancer and to help find new therapeutic or biomarker goals. Several computational methods to find the genes responsible for cancer have been developed based on genome data. However, many of these methods find key mutations in genomic data to predict which genes are responsible for cancer. These methods depend on the mutation and genome data and often show a high rate of false positives in the results. In this study, we proposed an influence maximization-based approach, CinfuMax, which can detect the genes responsible for cancer without needing information on mutations.Methods: In this method, the concept of influence maximization and the independent cascade model are employed. Firstly, the gene regulatory network for breast, lung and colon cancers was built using regulatory interactions and gene expression data. Next, we implemented an independent cascade diffusion algorithm on the networks to compute each gene's coverage. Finally, the genes with the highest coverage were classified as driver.Results: The results of the proposed method were compared to 19 other computational and network-based methods based on the F-measure and the number of detected driver genes. The results demonstrated that the proposed method produces better results than other methods. Also, CinfuMax is able to detect 18, 19 and 22 individual driver genes in three breast, lung and colon cancers, respectively, which have not been identified in any of the previous methods.Conclusion: The results show that independent cascading methods to identify driver genes perform better than linear threshold methods. Driver genes are also classified in terms of influence speed and have identified the genes with the highest diffusion rate in each type of cancer. Identification of these genes can be useful for molecular therapies and drug purposes.