/ THE IDEA
A loss function gives the model a score for being wrong. The gradient tells us how that score would change if each parameter moved slightly. Move opposite the gradient and, locally, loss falls fastest. The learning rate controls step size. Too small and training crawls. Too large and the model can overshoot a useful region or bounce without settling.
THE FORMAL IDEA
θ_new = θ_old − η ∇L(θ)
| θ (theta) = all adjustable model parameters | | L = loss; ∇L = direction of steepest increase | | η (eta) = learning rate controlling step size |
|
RUN THE TINY EXAMPLE
Walk toward the number 3
Loss L(x) = (x − 3)²; start x = 0 → loss 9 Slope dL/dx = 2(x − 3); at x = 0, slope = 2(0 − 3) = −6 Choose η = 0.1 New x = 0 − 0.1(−6) = 0.6 → loss 5.76
|
One local step reduced the error. Repeating recalculates the slope from the new position and continues toward the minimum at x = 3.
/ SO WHAT?
This explains why training is iterative rather than a model instantly absorbing data. It also explains loss curves: early steps can improve quickly, then flatten as gradients shrink near a solution.
ONE CAVEAT |
| With millions of adjustable numbers, the error surface is not one smooth bowl. Gradients can be noisy, data can encode bias, and lower training loss does not guarantee accurate or useful behaviour on new data. |
KEEP THIS
Gradient descent learns by measuring local error slope and taking many controlled steps downhill.
|
NEXT: How to find a whisper inside static
|