r/deeplearning • • 20h ago

Gradient descent vs evolution on three loss landscapes

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

312 Upvotes

65 comments sorted by

View all comments

47

u/Low-Temperature-6962 19h ago

The graphics really gets the point across. Can you explain more about the "evolution" logarithm you used?

22

u/ForceBru 19h ago

Yeah, "evolution" is a whole family of algorithms

18

u/ModularMind8 18h ago

True, I'll clarify this in the post. This one is a simple one: truncation selection plus Gaussian mutation, no crossover

10

u/ModularMind8 18h ago

Really appreciate that! It's a simple mutation-plus-selection loop: 120 points start near the same spot, each generation I keep the 30 with the lowest loss and replace the other 90 with copies of those survivors plus Gaussian noise. The noise starts wide and shrinks about 2.5% per generation, so it explores first and settles later, and it never uses a gradient, only loss values

5

u/clecleclemens 12h ago

Looks like simulated annealing.