>
>
Convergence Rate Bounds for Stochastic Gradient Descent With Adaptive Step Sizes on Non-Convex Loss Surfaces
Convergence Rate Bounds for Stochastic Gradient Descent With Adaptive Step Sizes on Non-Convex Loss Surfaces
Publisher : PJPCR
Author(s)
Valentina K. Petrakis; Olumide T. Adesanya; William J. Harrington
Abstract
This study investigates convergence rate bounds for stochastic gradient descent with adaptive step-size schedules on non-convex loss surfaces within the context of mathematical optimization theory and machine learning theory, an area of growing scientific importance given its implications for training efficiency of large-scale deep learning models and hyperparameter-free optimization algorithm design. Using theoretical convergence analysis with empirical verification on non-convex benchmark functions and neural network training tasks, we examine adaptive pre-conditioning of gradient noise enabling escape from saddle points and convergence to first-order stationary points in 12 adaptive SGD variants across 6 benchmark optimization landscapes and 3 neural network architectures drawn from controlled numerical experiment environments with fixed random seeds for reproducibility. Results indicate that the proposed step-size schedule achieves an O(log(T)/sqrt(T)) convergence rate, outperforming constant step-size baselines by 31.4% in iterations to target gradient norm (p < 0.001), with 31.4% improvement in convergence speed as the primary quantitative benchmark. Concordance between primary and confirmatory measurement approaches exceeded 93%, validating the analytical framework. These findings contribute empirically to mathematical optimization theory and machine learning theory and carry actionable implications for the design of programs and policies targeting training efficiency of large-scale deep learning models and hyperparameter-free optimization algorithm design.
