A novel three-term conjugate gradient approach for deep neural network training
نویسندگان
1 Department of Mathematics, Faculty of Mathematics, Statistics and Computer Science, Semnan University, P.O. Box 35195-363, Semnan, Iran
2 Department of Mathematics, Faculty of Mathematics, Statistics and Computer Science, Semnan University, P.O. Box 35195-363, Semnan, Iran
doi
10.22067/ijnao.2025.94956.1709چکیده
This paper presents a new three-term conjugate gradient (CG) method for large-scale unconstrained optimization, with application to training deep neural networks. The approach employs enhanced CG parameter strategies together with the strong Wolfe line search to guarantee sufficient descent and global convergence under standard smoothness assumptions. Extensive numerical experiments are conducted on both the standard CUTEr benchmark test problems and the MNIST digit classication task. The proposed MPRP{IFR hybrid algorithm demonstrates superior efficiency compared to traditional CG variants, achieving faster convergence and reduced computational cost on CUTEr problems as well as high classication accuracy and stable training dynamics for deep learning models. These findings highlight the robustness and eEffectiveness of the proposed CG scheme for large-scale optimization.