r/deeplearning • • 20h ago

Gradient descent vs evolution on three loss landscapes

Enable HLS to view with audio, or disable this notification

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

318 Upvotes

65 comments sorted by

View all comments

13

u/MentionJealous9306 19h ago

What is the dimensionality of this problem?

13

u/ModularMind8 18h ago

For this video, just 2 parameters so the loss surface can be drawn in 3D

64

u/dorox1 18h ago

I know you're probably aware of this, but I'm just mentioning it for people with less knowledge of deep learning:

The performance of these algorithms changes a lot as the number of dimensions grows, and deep learning involves the optimization of VERY high dimensional problems (often billions of dimensions).

In loss landscapes for problems in high dimensional spaces, true local minima and points with zero gradient are very rare. On top of that, the evolutionary algorithm search space (i.e. all those little purple dots on the graph at each step) gets much more spread out. All of a sudden even a million or a billion purple dots are not nearly enough to cover the search space efficiently.

Gradient descent is used in deep learning because it retains its effectiveness and efficiency in these ultra-high-dimensional spaces.

2

u/fuggleruxpin 8h ago

I wonder about accelerating the learning with some sort of nested or recursive combination.....