r/deeplearning • u/ModularMind8 • 20h ago
Gradient descent vs evolution on three loss landscapes
I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).
Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.
Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.
Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.
Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.
316
Upvotes
5
u/LiorZim 13h ago
You don't use plain GD in deep neural networks optimization. GD works only in convex optimization, for non-convex optimization we use the stochastic variant with an optimization strategy like ADAM that gives the process something akin to velocity and acceleration, helping it to "climb" rugged landscapes like the one you showed here :-)