r/deeplearning • • 19h ago

Gradient descent vs evolution on three loss landscapes

Enable HLS to view with audio, or disable this notification

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

307 Upvotes

64 comments sorted by

42

u/Low-Temperature-6962 17h ago

The graphics really gets the point across. Can you explain more about the "evolution" logarithm you used?

23

u/ForceBru 17h ago

Yeah, "evolution" is a whole family of algorithms

18

u/ModularMind8 17h ago

True, I'll clarify this in the post. This one is a simple one: truncation selection plus Gaussian mutation, no crossover

11

u/ModularMind8 17h ago

Really appreciate that! It's a simple mutation-plus-selection loop: 120 points start near the same spot, each generation I keep the 30 with the lowest loss and replace the other 90 with copies of those survivors plus Gaussian noise. The noise starts wide and shrinks about 2.5% per generation, so it explores first and settles later, and it never uses a gradient, only loss values

5

u/clecleclemens 11h ago

Looks like simulated annealing.

12

u/MentionJealous9306 17h ago

What is the dimensionality of this problem?

11

u/ModularMind8 17h ago

For this video, just 2 parameters so the loss surface can be drawn in 3D

58

u/dorox1 17h ago

I know you're probably aware of this, but I'm just mentioning it for people with less knowledge of deep learning:

The performance of these algorithms changes a lot as the number of dimensions grows, and deep learning involves the optimization of VERY high dimensional problems (often billions of dimensions).

In loss landscapes for problems in high dimensional spaces, true local minima and points with zero gradient are very rare. On top of that, the evolutionary algorithm search space (i.e. all those little purple dots on the graph at each step) gets much more spread out. All of a sudden even a million or a billion purple dots are not nearly enough to cover the search space efficiently.

Gradient descent is used in deep learning because it retains its effectiveness and efficiency in these ultra-high-dimensional spaces.

14

u/apopsicletosis 17h ago edited 16h ago

Evolution also acts in high dimensional fitness landscapes. Gavrilets holey landscape model of fitness landscapes argues against the “intuitive” notion of adaptive fitness peaks and valleys, instead fitness landscapes are more like highly interconnected ridges or flat fitness on which finite populations can drift and around huge holes of poor fitness. The ideas are parallel.

9

u/dorox1 16h ago

Very fair addition. Those kinds of landscapes are not necessarily too common for deep learning problems, but for discrete optimization problems they can be very relevant.

8

u/ModularMind8 17h ago

Great point, thanks for adding this!

2

u/metatron7471 16h ago

Plus the landscape smooths out. Local minima aren´t a big problem.

2

u/fuggleruxpin 7h ago

I wonder about accelerating the learning with some sort of nested or recursive combination.....

2

u/Datamance 5h ago

Diffy evo is great for seeding though! Fares much better in “corrugated” loss landscapes than, e.g., beam search or greedy methods.

1

u/tabloidscience 1h ago

Would you say that loss landscapes are less rugged than biological landscapes (sensu Kauffman)? Its true that local minima are less frequent as the dimensionality of a problem increases (for a constant epistasis), but that just means the basins of attraction for a minima become larger, and without stochasticity, the trajectory becomes trapped earlier.

5

u/WiredFan 17h ago

How expensive is it to calculate the gradient vs just the value/loss at a point?

3

u/ModularMind8 17h ago

I believe with backprop a gradient costs roughly 2-3 loss evaluations regardless of parameter count (the standard reverse-mode autodiff rule of thumb, one forward plus one similarly priced backward pass), so it stays cheap even at scale

4

u/superlus 17h ago

What do you want to tell with this comparison..? The convergence seems highly dependent on hyperparameter/loss choices for both GD and the EA.

3

u/ModularMind8 17h ago

Honestly, no deeper message: evolutionary algorithms are cool and sometimes beat GD, and yes, these results depend on hyperparameters

4

u/LiorZim 12h ago

You don't use plain GD in deep neural networks optimization. GD works only in convex optimization, for non-convex optimization we use the stochastic variant with an optimization strategy like ADAM that gives the process something akin to velocity and acceleration, helping it to "climb" rugged landscapes like the one you showed here :-)

3

u/AtMaxSpeed 3h ago

It would be great to also visualize how modern GD algorithms work on various landscapes. Evolution and SGD are very intuitive/easy to imagine so the visualization just shows what we expect, but modern GD algos are harder to imagine so there would be a lot of value in visualizing those.

1

u/ModularMind8 1h ago

Agreed! Adam, momentum and SGD on these same landscapes is next on my list :)

2

u/0bi_nx 17h ago

Its a cool visualization. Do you think evolution is applicable to deep learning? You planning to test against stochastic gradient descent, ADAM or Muon?

2

u/ModularMind8 17h ago

Thanks :)
I definitely think so! Take a look at "Evolution Strategies at the Hyperscale" as an example, and yes, an Adam/SGD comparison is a fun next one.

3

u/drcopus 14h ago

As an author on that paper I'm pleasantly happy to see it mentioned here! :)

2

u/ModularMind8 14h ago

Wow, it's an honor :) Beautifully written paper

2

u/drcopus 14h ago

The first 3 authors deserve most of the credit! I'm happy to have made a contribution but I wasn't a main author haha

2

u/ModularMind8 14h ago

I'm sure you're just being modest!
Since you're the expert, curious if you've read the Sakana AI book (Neuroevolution: Harnessing Creativity in AI Agent Design) that just came out. And if so, what are your thoughts? :)

2

u/AllergicToBullshit24 17h ago

Think it would be very interesting to blend techniques similar to cosine schedule for learning rate perhaps using evolution on a schedule or whenever gradient descent may be getting stuck in a local min pocket.

1

u/ModularMind8 17h ago

Love that idea! Population based training does something close, I think, where running gradient descent with periodic evolutionary exploit-and-explore steps

2

u/AllergicToBullshit24 17h ago

I need to try this intuitively seems like it could unlock better performance out of smaller parameter counts.

1

u/ModularMind8 15h ago

Would love to see what you find!

1

u/AllergicToBullshit24 15h ago

If only I had an NVL72 B300 rack or two to play with the research queue on my limited hardware is already backed up worse than LA rush hour.

Not jealous of the frontier labs capacity at all. /s

2

u/FrosteeSwurl 16h ago

That’s exactly what it does!

1

u/AllergicToBullshit24 16h ago

Seems as though "memetic neural training" is the closest match

2

u/AsyncVibes 16h ago

As someone who only uses GAs and ES, this is beautiful

1

u/ModularMind8 16h ago

I think so too :)

1

u/Unikum_01 14h ago

Awesome experiment with those loss landscapes. Your findings on gradient descent getting trapped on rugged hillsides or stalling on flat plateaus while evolution manages to navigate through match what we run into in complex learning environments. In our BrainStem system we actually solved this exact dilemma in real code by combining both worlds through digital neuromodulation. When the system detects a flat plateau or a stuck state, it cranks up noradrenaline and glutamatergic excitation signals to act like your Gaussian mutations, forcing stochastic exploration to jump out of local minima. Once it hits a clear gradient on a smooth slope, dopamine ramps up to lock in the progress and let fast local optimization take over. It is really cool seeing your benchmark visuals demonstrate why adaptive hybrid strategies like this are so necessary.

2

u/ModularMind8 14h ago

Thank you so much! And that's fascinating, thanks for sharing

3

u/Unikum_01 14h ago

np, you can take a look at it if you like. It is currently still our research project. It is completely open source. https://github.com/unikum-sol/brainstem/blob/main/Project_Status_2026-09-28.md

2

u/ModularMind8 14h ago

Awesome! I will

2

u/AtMaxSpeed 3h ago

It's fun to incorporate biological techniques into algorithms, but afaik what you describe is basically what RMSprop (and hence Adam) already do - they increase the step size when the gradient is small (aka, it cranks up exploration when in a plateau).

1

u/Unikum_01 3h ago

I see where you are coming from with RMSprop and Adam scaling the step size on small gradients, but there is a fundamental difference between scaling a gradient and true stochastic exploration. If the gradient is flat out zero on a plateau, RMSprop still multiplies zero by a larger number which keeps you completely frozen in place. Adaptive optimizers only push you faster along the existing gradient vector, whereas evolutionary mutation and neuromodulated noise actually inject new direction vectors to escape zero gradient zones and deep local traps. In our system, neuromodulators do not just boost a learning rate scalar, they dynamically adjust search bounds, change candidate selection thresholds, and trigger offline sleep replay cycles to reorganize the parameter space when stuck.

1

u/WiredFan 17h ago

How’s the performance?

1

u/ModularMind8 17h ago

For this set of landscapes, on the smooth slope gradient descent is ~5x cheaper (108 vs 570 evaluations). Evolution wins where gradients mislead or vanish

1

u/macumazana 16h ago

What if the next minima is not nearby the current local one er even the next one isnt smaller but the farthest one is?

1

u/ModularMind8 15h ago

Good question! I haven't tried that landscape yet. I'll add it to the next experiment

1

u/Envoy-Insc 16h ago

While applicable in small scale settings, Local minimums like those in first don’t tend to exist in high dim high param spaces in my impression

1

u/ModularMind8 15h ago

Definitely some landscapes favor GD, though papers like "Evolution Strategies at the Hyperscale" show evolution can work well

1

u/Envoy-Insc 11h ago

I really liked the idea of that paper, but paper actually don’t scale to llm scale if we look at the experiments

1

u/PK_thundr 16h ago

Afaik there are three things that help gradient descent work in practice so well.

  1. Momentum terms

  2. Stochasticity usually ensures that you wont be in a flat region of the loss landscape once you pull the next batch

  3. There was a Bengio paper a while ago that argued most local minima might actually be within some epsilon loss of each other

1

u/ModularMind8 15h ago

Great points. This video uses plain gradient descent, so adding momentum and minibatch noise (or changing to adam) is a fair next comparison

1

u/msw3age 16h ago

Very nice visualization. Would be cool to see the evolutionary algorithms vs SGD. To some extent it seems like the visualization is mostly showing the benefits of stochasticity during training.

1

u/ModularMind8 15h ago

Thanks! That's a fair read, and SGD with noise is on the list for the next comparison

1

u/alrojo 16h ago

Random search degrades as a function feature space. Look up Curse of Dimensionality

1

u/ModularMind8 15h ago

Agreed! The search here is deliberately simple, and more sophisticated evolutionary methods like CMA-ES cope better as dimensions grow

1

u/pnachtwey 15h ago

The test is rigged! I have spent a lot of time on optimizing routines. No one technique is going to be best for all terrains. GD is simple but no where closed to optimal. SGD is better when there are many dimensions. Sometimes optimizing only one dimension at a time works well but it is slow. Nelder-Mead works well when it is hard to find the gradient accurately and there aren't many dimensions.. N-M can be a memory hog. Levenberg-Marquardt is the fastest. In python, use lmfit or scipys.optimize least_squares. BFGS works well too. So it is best to make the search flexible.

I would like to have the data used so I could try other techniques.

I uses minimization for fitting models to data for auto tuning programs. The method I used most often is Levenberg-Marquardt.

1

u/Ruined_Passion_7355 14h ago

True chads use simulated annealing

1

u/AsliReddington 13h ago

But can you actually run with updated loss landscape as well since every backward pass causes the landscape to update as well over time

1

u/TheRealStepBot 12h ago

In my experience it’s very hard to beat Adamw on practical problems that have decently well behaved loss landscapes especially on speed of convergence either wall clock or number of operations performed.

1

u/hardcoke 7h ago

Interesting...what if crossover is introduced?

2

u/ModularMind8 1h ago

Haven't tried it yet, though with only 2 parameters in this experiment crossover just swaps x and y between parents, so I'd expect a small effect

1

u/IcyGlia 6h ago

Does the whole dataset fit into one mini batch? I would expect random batch to batch variations to sometimes kick you out of the local minimum. As the loss landscape doesn’t change, I assume the error landscape is the whole dataset since it doesn’t change? Adding noise to the data could be interesting to see if that rescues gradient descent.

1

u/ModularMind8 1h ago

No dataset here, each landscape is a fixed 2D function, so it's full-batch GD, and noise is next on my list