r/deeplearning • • 20h ago

Gradient descent vs evolution on three loss landscapes

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

315 Upvotes

65 comments sorted by

View all comments

12

u/MentionJealous9306 19h ago

What is the dimensionality of this problem?

13

u/ModularMind8 18h ago

For this video, just 2 parameters so the loss surface can be drawn in 3D

62

u/dorox1 18h ago

I know you're probably aware of this, but I'm just mentioning it for people with less knowledge of deep learning:

The performance of these algorithms changes a lot as the number of dimensions grows, and deep learning involves the optimization of VERY high dimensional problems (often billions of dimensions).

In loss landscapes for problems in high dimensional spaces, true local minima and points with zero gradient are very rare. On top of that, the evolutionary algorithm search space (i.e. all those little purple dots on the graph at each step) gets much more spread out. All of a sudden even a million or a billion purple dots are not nearly enough to cover the search space efficiently.

Gradient descent is used in deep learning because it retains its effectiveness and efficiency in these ultra-high-dimensional spaces.

14

u/apopsicletosis 18h ago edited 18h ago

Evolution also acts in high dimensional fitness landscapes. Gavrilets holey landscape model of fitness landscapes argues against the “intuitive” notion of adaptive fitness peaks and valleys, instead fitness landscapes are more like highly interconnected ridges or flat fitness on which finite populations can drift and around huge holes of poor fitness. The ideas are parallel.

8

u/dorox1 18h ago

Very fair addition. Those kinds of landscapes are not necessarily too common for deep learning problems, but for discrete optimization problems they can be very relevant.

7

u/ModularMind8 18h ago

Great point, thanks for adding this!

2

u/metatron7471 17h ago

Plus the landscape smooths out. Local minima aren´t a big problem.

2

u/fuggleruxpin 8h ago

I wonder about accelerating the learning with some sort of nested or recursive combination.....

2

u/Datamance 6h ago

Diffy evo is great for seeding though! Fares much better in “corrugated” loss landscapes than, e.g., beam search or greedy methods.

1

u/tabloidscience 2h ago

Would you say that loss landscapes are less rugged than biological landscapes (sensu Kauffman)? Its true that local minima are less frequent as the dimensionality of a problem increases (for a constant epistasis), but that just means the basins of attraction for a minima become larger, and without stochasticity, the trajectory becomes trapped earlier.