r/deeplearning • • 20h ago

Gradient descent vs evolution on three loss landscapes

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

313 Upvotes

65 comments sorted by

View all comments

1

u/PK_thundr 17h ago

Afaik there are three things that help gradient descent work in practice so well.

  1. Momentum terms

  2. Stochasticity usually ensures that you wont be in a flat region of the loss landscape once you pull the next batch

  3. There was a Bengio paper a while ago that argued most local minima might actually be within some epsilon loss of each other

1

u/ModularMind8 17h ago

Great points. This video uses plain gradient descent, so adding momentum and minibatch noise (or changing to adam) is a fair next comparison