1. Improve hot region prediction by analyzing different machine learning algorithms
- Author
-
Longwei Zhou, Nansheng Chen, Jing Hu, Bo Li, and Xiaolong Zhang
- Subjects
DBSCAN ,Support Vector Machine ,Computer science ,QH301-705.5 ,SVM ,Computer applications to medicine. Medical informatics ,R858-859.7 ,Hot spot (veterinary medicine) ,Machine learning ,computer.software_genre ,Biochemistry ,Machine Learning ,Naive Bayes classifier ,Structural Biology ,Cluster Analysis ,Hot spot ,Biology (General) ,Cluster analysis ,Molecular Biology ,Artificial neural network ,business.industry ,Applied Mathematics ,Research ,Proteins ,Bayes Theorem ,Computer Science Applications ,Random forest ,Support vector machine ,Protein–protein interaction ,Hot region ,Artificial intelligence ,Precision and recall ,business ,computer ,Algorithm ,Algorithms ,Gaussian Naïve Bayes - Abstract
Background In the process of designing drugs and proteins, it is crucial to recognize hot regions in protein–protein interactions. Each hot region of protein–protein interaction is composed of at least three hot spots, which play an important role in binding. However, it takes time and labor force to identify hot spots through biological experiments. If predictive models based on machine learning methods can be trained, the drug design process can be effectively accelerated. Results The results show that different machine learning algorithms perform similarly, as evaluating using the F-measure. The main differences between these methods are recall and precision. Since the key attribute of hot regions is that they are packed tightly, we used the cluster algorithm to predict hot regions. By combining Gaussian Naïve Bayes and DBSCAN, the F-measure of hot region prediction can reach 0.809. Conclusions In this paper, different machine learning models such as Gaussian Naïve Bayes, SVM, Xgboost, Random Forest, and Artificial Neural Network are used to predict hot spots. The experiment results show that the combination of hot spot classification algorithm with higher recall rate and clustering algorithm with higher precision can effectively improve the accuracy of hot region prediction.
- Published
- 2021