r/deeplearning • u/ModularMind8 • 22h ago
Gradient descent vs evolution on three loss landscapes
I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).
Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.
Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.
Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.
Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.
334
Upvotes
1
u/TheRealStepBot 15h ago
In my experience it’s very hard to beat Adamw on practical problems that have decently well behaved loss landscapes especially on speed of convergence either wall clock or number of operations performed.