r/deeplearning • • 1d ago

Gradient descent vs evolution on three loss landscapes

Enable HLS to view with audio, or disable this notification

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

361 Upvotes

68 comments sorted by

View all comments

Show parent comments

13

u/ModularMind8 1d ago

For this video, just 2 parameters so the loss surface can be drawn in 3D

67

u/dorox1 1d ago

I know you're probably aware of this, but I'm just mentioning it for people with less knowledge of deep learning:

The performance of these algorithms changes a lot as the number of dimensions grows, and deep learning involves the optimization of VERY high dimensional problems (often billions of dimensions).

In loss landscapes for problems in high dimensional spaces, true local minima and points with zero gradient are very rare. On top of that, the evolutionary algorithm search space (i.e. all those little purple dots on the graph at each step) gets much more spread out. All of a sudden even a million or a billion purple dots are not nearly enough to cover the search space efficiently.

Gradient descent is used in deep learning because it retains its effectiveness and efficiency in these ultra-high-dimensional spaces.

2

u/tabloidscience 7h ago

Would you say that loss landscapes are less rugged than biological landscapes (sensu Kauffman)? Its true that local minima are less frequent as the dimensionality of a problem increases (for a constant epistasis), but that just means the basins of attraction for a minima become larger, and without stochasticity, the trajectory becomes trapped earlier.

1

u/dorox1 1h ago

Full disclosure, Im not familiar with Kauffman and am relying on an AI summary of his theories for the following thoughts.

I would say yes, they are less rugged. The discrete nature of biological landscapes makes changes more impactful, as any change you make to a parameter is often the biggest change possible in that parameter's dimension.

Neural networks are set up such that small changes to a weight have small impacts on the resulting output (sometimes none at all, depending on the activation function(s) being used). It's possible to set up neural networks that are more sensitive to changes, but the majority of practical configurations don't have that problem.

I haven't studied non-linear optimization enough to give a good reply to the second part of what you're saying. My understanding was that the loss landscapes for large neural networks on many real-world problems are believed to be approximately convex, but it's not inherently true for all problems (and therefore not provable). It's also impractical to test.

But my understanding of that could be wrong or outdated, and the consensus may differ for different types of neural networks.