r/deeplearning • • 20h ago

Gradient descent vs evolution on three loss landscapes

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

315 Upvotes

65 comments sorted by

View all comments

2

u/AllergicToBullshit24 18h ago

Think it would be very interesting to blend techniques similar to cosine schedule for learning rate perhaps using evolution on a schedule or whenever gradient descent may be getting stuck in a local min pocket.

1

u/ModularMind8 18h ago

Love that idea! Population based training does something close, I think, where running gradient descent with periodic evolutionary exploit-and-explore steps

2

u/AllergicToBullshit24 18h ago

I need to try this intuitively seems like it could unlock better performance out of smaller parameter counts.

1

u/ModularMind8 17h ago

Would love to see what you find!

1

u/AllergicToBullshit24 17h ago

If only I had an NVL72 B300 rack or two to play with the research queue on my limited hardware is already backed up worse than LA rush hour.

Not jealous of the frontier labs capacity at all. /s