Personal tools

ML Optimization and Gradient Descent

Lapland, Finland - sennarelax]
[Lapland, Finland - sennarelax]


- Overview

In machine learning (ML), optimization is the process of adjusting a model's internal parameters to make its predictions as accurate as possible. The core mechanism behind this is gradient descent, an iterative optimization algorithm used to minimize a model's error, which is mathematically represented by a loss or cost function. 

Please refer to the following for more information:

 

1. The Core Concepts: 

To understand how gradient descent works, it helps to break down its three fundamental components:

  • The Loss Function (L(θ)): This functions as a barometer for accuracy. A loss function calculates the error for a single training example, while a cost function averages that error across the entire dataset. The ultimate goal of optimization is to tweak the model's parameters (θ, such as weights and biases) to make this value as close to zero as possible.
  • The Gradient (∇ L(θ)): This is a vector of partial derivatives that points in the direction of the steepest increase of the loss function. Think of it like standing on a foggy mountainside; the gradient tells you which way goes straight up.
  • The Learning Rate (α or η): A crucial hyperparameter that dictates the size of the step taken down the slope during each iteration.

 

2. The Mathematical Update Rule:

Because the gradient points uphill, the algorithm takes steps in the opposite direction to minimize the error. The parameters are updated iteratively using the following formula:

 

𝜃𝑛⁢𝑒⁢𝑤=𝜃𝑜⁢𝑙⁢𝑑–𝑎⋅∇𝐽⁡(𝜃)

 

  • If the gradient is positive, the step moves backward.
  • If the gradient is negative, the step moves forward.
  • This cycle repeats until the slope becomes nearly flat, a state known as convergence.

 

3. The Three Primary Variants: 

Depending on how much data is processed before updating the weights, gradient descent is split into three main types: 

  • Batch Gradient Descent: Computes the gradients using the entire dataset at once. It provides stable updates but is computationally expensive and slow for massive data pools. 
  • Stochastic Gradient Descent (SGD): Updates parameters using only one random data point per iteration. It is fast but introduces a noisy, erratic path toward the minimum.
  • Mini-batch Gradient Descent: Divides the data into small batch sizes (e.g., 32 or 64) for updates. According to IBM, this approach strikes the ideal balance between computational efficiency and speed.

 

4. Common Challenges in Optimization: 

Standard gradient descent faces several structural roadblocks, especially in deep or non-convex neural network architectures:

  • Local Minima & Saddle Points: For complex functions, the algorithm can get trapped in a "local minimum" (a valley that isn't the absolute lowest point) or a "saddle point" (where the slope is zero but it's a maximum on one side and a minimum on the other). 
  • Overshooting vs. Slow Convergence: If the learning rate is set too high, the algorithm can overshoot the minimum and diverge. If it is too low, the model takes tiny steps and takes an impractical amount of time to train.
  • Vanishing & Exploding Gradients: In deep networks, as gradients are passed backward through layers, they can shrink to zero (causing the model to stop learning) or grow exponentially large (causing numerical instability and NaN errors).

 

5. Advanced Evolution: Adaptive Optimizers:  

To overcome these limits, modern machine learning frameworks rely on advanced optimizers that adjust the learning rate dynamically:

  • Momentum: Accelerates optimization by adding a fraction of the previous step’s direction, acting like a ball gaining speed rolling down a hill.
  • RMSprop: Restricts vertical oscillations by keeping a moving average of recent squared gradients, allowing faster horizontal progress.
  • Adam (Adaptive Moment Estimation): Combines the principles of Momentum and RMSprop. It maintains unique, adaptive learning rates for every single parameter, making it the industry standard for training complex architectures.


 

[More to come ...]


Document Actions