Prediction of Hotspots in Proteins Using Machine Learning Techniques

نویسندگان

1 Electronics and Communication Engineering Department, Koneru Lakshmiah Education Foundation, Hyderabad, India

2 Electronics and Communication Engineering Department, Koneru Lakshmiah Education Foundation, Hyderabad, India

doi
10.5829/ije.2026.39.07a.06
چکیده

In order to create medications that effectively modulate Protein-Protein Interaction (PPI) and aid in the treatment of illnesses, protein hotspots are essential. One computational technique that aids in the prediction of protein hotspots is machine learning. The use of machine learning methods to predict protein hotspots is investigated in this work. Feature selection methods were used to analyze a protein dataset. k-Nearest Neighbor (kNN), Support Vector Machine (SVM), Random Forest, XGBoost, and Logistic Regression were among the machine learning algorithms used. With models such as Logistic Regression and SVM, Physicochemical features (PHY) and Solvent Accessible surface Area (ASA) demonstrated strong predictive performance among the features examined, attaining accuracy scores of up to 0.730, AUC values up to 0.885, and F1-scores up to 0.642. These metrics indicate the effectiveness with which they can predict hotspots which supports computational biology applications.