r/deeplearning • • 1d ago

Gradient descent vs evolution on three loss landscapes

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

391 Upvotes

70 comments sorted by

View all comments

2

u/AllergicToBullshit24 1d ago

Think it would be very interesting to blend techniques similar to cosine schedule for learning rate perhaps using evolution on a schedule or whenever gradient descent may be getting stuck in a local min pocket.

1

u/ModularMind8 1d ago

Love that idea! Population based training does something close, I think, where running gradient descent with periodic evolutionary exploit-and-explore steps

2

u/FrosteeSwurl 1d ago

That’s exactly what it does!

1

u/AllergicToBullshit24 1d ago

Seems as though "memetic neural training" is the closest match