An integrated approach combining autoencoders with fuzzy C-means and Chebyshev distance for enhanced genomic clustering in cancer research

نویسندگان

1 Department of Mathematics, Thapar Institute of Engineering & Technology, Patiala 147004, Punjab, India.

2 Faculty of Informatics and Computing, Universiti Sultan Zainal Abidin, Besut Campus, 22200 Besut, Terengganu, Malaysia.

3 Faculty of Informatics and Computing, Universiti Sultan Zainal Abidin, Besut Campus, 22200 Besut, Terengganu, Malaysia.

4 Faculty of Informatics and Computing, Universiti Sultan Zainal Abidin, Besut Campus, 22200 Besut, Terengganu, Malaysia.

5 Faculty of Medicine, Universiti Sultan Zainal Abidin, Medical Campus, 20400 Kuala Terengganu, Terengganu, Malaysia.

doi
10.22105/jfea.2025.526247.1930
چکیده

Genomic clustering is one of the vital tools in cancer research for uncovering hidden patterns within complex genetic datasets. Despite advancements in various clustering techniques, conventional methods often struggle with high-dimensional genomic data due to noise sensitivity and difficulty in capturing overlapping or subtle patterns. This study proposes an integrated framework that combines Autoencoders, Fuzzy C-Means (FCM) clustering, and the Chebyshev distance metric to enhance clustering accuracy. In the proposed approach, Autoencoders perform dimensionality reduction by preserving essential features while mitigating the curse of dimensionality and improving computational efficiency. FCM allows overlapping clusters, effectively representing the inherent biological variability in genomic data, while the Chebyshev distance refines cluster boundaries by considering variations across multiple dimensions. The proposed framework was evaluated using high-dimensional genomic datasets and compared with Principal Component Analysis (PCA)-based approaches. The results demonstrate that the Autoencoder-FCM-Chebyshev framework outperforms PCA-based models, achieving a Silhouette Score of 0.9718, a Davies-Bouldin Index of 0.0015, and lower Within-Cluster Sum of Squares. These findings highlight the robustness and effectiveness of the proposed framework in capturing complex patterns within highdimensional genomic data.