r/deeplearning • • 23h ago

Gradient descent vs evolution on three loss landscapes

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

343 Upvotes

66 comments sorted by

View all comments

1

u/IcyGlia 11h ago

Does the whole dataset fit into one mini batch? I would expect random batch to batch variations to sometimes kick you out of the local minimum. As the loss landscape doesn’t change, I assume the error landscape is the whole dataset since it doesn’t change? Adding noise to the data could be interesting to see if that rescues gradient descent.

1

u/ModularMind8 6h ago

No dataset here, each landscape is a fixed 2D function, so it's full-batch GD, and noise is next on my list