A novel three-term conjugate gradient approach for deep neural network training

نویسندگان

1 Department of Mathematics, Faculty of Mathematics, Statistics and Computer Science, Semnan University, P.O. Box 35195-363, Semnan, Iran

2 Department of Mathematics, Faculty of Mathematics, Statistics and Computer Science, Semnan University, P.O. Box 35195-363, Semnan, Iran

doi
10.22067/ijnao.2025.94956.1709
چکیده

This paper presents a new three-term conjugate gradient (CG) method for large-scale unconstrained optimization, with application to training deep neural networks. The approach employs enhanced CG parameter strategies together with the strong Wolfe line search to guarantee sufficient descent and global convergence under standard smoothness assumptions. Extensive numerical experiments are conducted on both the standard CUTEr benchmark test problems and the MNIST digit classi cation task. The proposed MPRP{IFR hybrid algorithm demonstrates superior efficiency compared to traditional CG variants, achieving faster convergence and reduced computational cost on CUTEr problems as well as high classi cation accuracy and stable training dynamics for deep learning models. These findings highlight the robustness and e Effectiveness of the proposed CG scheme for large-scale optimization.