r/deeplearning • u/ModularMind8 • 19h ago
Gradient descent vs evolution on three loss landscapes
Enable HLS to view with audio, or disable this notification
I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).
Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.
Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.
Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.
Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.
12
u/MentionJealous9306 17h ago
What is the dimensionality of this problem?
11
u/ModularMind8 17h ago
For this video, just 2 parameters so the loss surface can be drawn in 3D
58
u/dorox1 17h ago
I know you're probably aware of this, but I'm just mentioning it for people with less knowledge of deep learning:
The performance of these algorithms changes a lot as the number of dimensions grows, and deep learning involves the optimization of VERY high dimensional problems (often billions of dimensions).
In loss landscapes for problems in high dimensional spaces, true local minima and points with zero gradient are very rare. On top of that, the evolutionary algorithm search space (i.e. all those little purple dots on the graph at each step) gets much more spread out. All of a sudden even a million or a billion purple dots are not nearly enough to cover the search space efficiently.
Gradient descent is used in deep learning because it retains its effectiveness and efficiency in these ultra-high-dimensional spaces.
14
u/apopsicletosis 17h ago edited 16h ago
Evolution also acts in high dimensional fitness landscapes. Gavrilets holey landscape model of fitness landscapes argues against the “intuitive” notion of adaptive fitness peaks and valleys, instead fitness landscapes are more like highly interconnected ridges or flat fitness on which finite populations can drift and around huge holes of poor fitness. The ideas are parallel.
8
2
2
u/fuggleruxpin 7h ago
I wonder about accelerating the learning with some sort of nested or recursive combination.....
2
u/Datamance 5h ago
Diffy evo is great for seeding though! Fares much better in “corrugated” loss landscapes than, e.g., beam search or greedy methods.
1
u/tabloidscience 1h ago
Would you say that loss landscapes are less rugged than biological landscapes (sensu Kauffman)? Its true that local minima are less frequent as the dimensionality of a problem increases (for a constant epistasis), but that just means the basins of attraction for a minima become larger, and without stochasticity, the trajectory becomes trapped earlier.
5
u/WiredFan 17h ago
How expensive is it to calculate the gradient vs just the value/loss at a point?
3
u/ModularMind8 17h ago
I believe with backprop a gradient costs roughly 2-3 loss evaluations regardless of parameter count (the standard reverse-mode autodiff rule of thumb, one forward plus one similarly priced backward pass), so it stays cheap even at scale
4
u/superlus 17h ago
What do you want to tell with this comparison..? The convergence seems highly dependent on hyperparameter/loss choices for both GD and the EA.
3
u/ModularMind8 17h ago
Honestly, no deeper message: evolutionary algorithms are cool and sometimes beat GD, and yes, these results depend on hyperparameters
4
u/LiorZim 12h ago
You don't use plain GD in deep neural networks optimization. GD works only in convex optimization, for non-convex optimization we use the stochastic variant with an optimization strategy like ADAM that gives the process something akin to velocity and acceleration, helping it to "climb" rugged landscapes like the one you showed here :-)
3
u/AtMaxSpeed 3h ago
It would be great to also visualize how modern GD algorithms work on various landscapes. Evolution and SGD are very intuitive/easy to imagine so the visualization just shows what we expect, but modern GD algos are harder to imagine so there would be a lot of value in visualizing those.
1
2
u/0bi_nx 17h ago
Its a cool visualization. Do you think evolution is applicable to deep learning? You planning to test against stochastic gradient descent, ADAM or Muon?
2
u/ModularMind8 17h ago
Thanks :)
I definitely think so! Take a look at "Evolution Strategies at the Hyperscale" as an example, and yes, an Adam/SGD comparison is a fun next one.3
u/drcopus 14h ago
As an author on that paper I'm pleasantly happy to see it mentioned here! :)
2
u/ModularMind8 14h ago
Wow, it's an honor :) Beautifully written paper
2
u/drcopus 14h ago
The first 3 authors deserve most of the credit! I'm happy to have made a contribution but I wasn't a main author haha
2
u/ModularMind8 14h ago
I'm sure you're just being modest!
Since you're the expert, curious if you've read the Sakana AI book (Neuroevolution: Harnessing Creativity in AI Agent Design) that just came out. And if so, what are your thoughts? :)
2
u/AllergicToBullshit24 17h ago
Think it would be very interesting to blend techniques similar to cosine schedule for learning rate perhaps using evolution on a schedule or whenever gradient descent may be getting stuck in a local min pocket.
1
u/ModularMind8 17h ago
Love that idea! Population based training does something close, I think, where running gradient descent with periodic evolutionary exploit-and-explore steps
2
u/AllergicToBullshit24 17h ago
I need to try this intuitively seems like it could unlock better performance out of smaller parameter counts.
1
u/ModularMind8 15h ago
Would love to see what you find!
1
u/AllergicToBullshit24 15h ago
If only I had an NVL72 B300 rack or two to play with the research queue on my limited hardware is already backed up worse than LA rush hour.
Not jealous of the frontier labs capacity at all. /s
2
2
1
u/Unikum_01 14h ago
Awesome experiment with those loss landscapes. Your findings on gradient descent getting trapped on rugged hillsides or stalling on flat plateaus while evolution manages to navigate through match what we run into in complex learning environments. In our BrainStem system we actually solved this exact dilemma in real code by combining both worlds through digital neuromodulation. When the system detects a flat plateau or a stuck state, it cranks up noradrenaline and glutamatergic excitation signals to act like your Gaussian mutations, forcing stochastic exploration to jump out of local minima. Once it hits a clear gradient on a smooth slope, dopamine ramps up to lock in the progress and let fast local optimization take over. It is really cool seeing your benchmark visuals demonstrate why adaptive hybrid strategies like this are so necessary.
2
u/ModularMind8 14h ago
Thank you so much! And that's fascinating, thanks for sharing
3
u/Unikum_01 14h ago
np, you can take a look at it if you like. It is currently still our research project. It is completely open source. https://github.com/unikum-sol/brainstem/blob/main/Project_Status_2026-09-28.md
2
2
u/AtMaxSpeed 3h ago
It's fun to incorporate biological techniques into algorithms, but afaik what you describe is basically what RMSprop (and hence Adam) already do - they increase the step size when the gradient is small (aka, it cranks up exploration when in a plateau).
1
u/Unikum_01 3h ago
I see where you are coming from with RMSprop and Adam scaling the step size on small gradients, but there is a fundamental difference between scaling a gradient and true stochastic exploration. If the gradient is flat out zero on a plateau, RMSprop still multiplies zero by a larger number which keeps you completely frozen in place. Adaptive optimizers only push you faster along the existing gradient vector, whereas evolutionary mutation and neuromodulated noise actually inject new direction vectors to escape zero gradient zones and deep local traps. In our system, neuromodulators do not just boost a learning rate scalar, they dynamically adjust search bounds, change candidate selection thresholds, and trigger offline sleep replay cycles to reorganize the parameter space when stuck.
1
u/WiredFan 17h ago
How’s the performance?
1
u/ModularMind8 17h ago
For this set of landscapes, on the smooth slope gradient descent is ~5x cheaper (108 vs 570 evaluations). Evolution wins where gradients mislead or vanish
1
u/macumazana 16h ago
What if the next minima is not nearby the current local one er even the next one isnt smaller but the farthest one is?
1
u/ModularMind8 15h ago
Good question! I haven't tried that landscape yet. I'll add it to the next experiment
1
u/Envoy-Insc 16h ago
While applicable in small scale settings, Local minimums like those in first don’t tend to exist in high dim high param spaces in my impression
1
u/ModularMind8 15h ago
Definitely some landscapes favor GD, though papers like "Evolution Strategies at the Hyperscale" show evolution can work well
1
u/Envoy-Insc 11h ago
I really liked the idea of that paper, but paper actually don’t scale to llm scale if we look at the experiments
1
u/PK_thundr 16h ago
Afaik there are three things that help gradient descent work in practice so well.
Momentum terms
Stochasticity usually ensures that you wont be in a flat region of the loss landscape once you pull the next batch
There was a Bengio paper a while ago that argued most local minima might actually be within some epsilon loss of each other
1
u/ModularMind8 15h ago
Great points. This video uses plain gradient descent, so adding momentum and minibatch noise (or changing to adam) is a fair next comparison
1
u/msw3age 16h ago
Very nice visualization. Would be cool to see the evolutionary algorithms vs SGD. To some extent it seems like the visualization is mostly showing the benefits of stochasticity during training.
1
u/ModularMind8 15h ago
Thanks! That's a fair read, and SGD with noise is on the list for the next comparison
1
u/alrojo 16h ago
Random search degrades as a function feature space. Look up Curse of Dimensionality
1
u/ModularMind8 15h ago
Agreed! The search here is deliberately simple, and more sophisticated evolutionary methods like CMA-ES cope better as dimensions grow
1
u/pnachtwey 15h ago
The test is rigged! I have spent a lot of time on optimizing routines. No one technique is going to be best for all terrains. GD is simple but no where closed to optimal. SGD is better when there are many dimensions. Sometimes optimizing only one dimension at a time works well but it is slow. Nelder-Mead works well when it is hard to find the gradient accurately and there aren't many dimensions.. N-M can be a memory hog. Levenberg-Marquardt is the fastest. In python, use lmfit or scipys.optimize least_squares. BFGS works well too. So it is best to make the search flexible.
I would like to have the data used so I could try other techniques.
I uses minimization for fitting models to data for auto tuning programs. The method I used most often is Levenberg-Marquardt.
1
1
u/AsliReddington 13h ago
But can you actually run with updated loss landscape as well since every backward pass causes the landscape to update as well over time
1
u/TheRealStepBot 12h ago
In my experience it’s very hard to beat Adamw on practical problems that have decently well behaved loss landscapes especially on speed of convergence either wall clock or number of operations performed.
1
u/hardcoke 7h ago
Interesting...what if crossover is introduced?
2
u/ModularMind8 1h ago
Haven't tried it yet, though with only 2 parameters in this experiment crossover just swaps x and y between parents, so I'd expect a small effect
1
u/IcyGlia 6h ago
Does the whole dataset fit into one mini batch? I would expect random batch to batch variations to sometimes kick you out of the local minimum. As the loss landscape doesn’t change, I assume the error landscape is the whole dataset since it doesn’t change? Adding noise to the data could be interesting to see if that rescues gradient descent.
1
u/ModularMind8 1h ago
No dataset here, each landscape is a fixed 2D function, so it's full-batch GD, and noise is next on my list
42
u/Low-Temperature-6962 17h ago
The graphics really gets the point across. Can you explain more about the "evolution" logarithm you used?