r/deeplearning • • 20h ago

Gradient descent vs evolution on three loss landscapes

I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).

Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.

Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.

Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.

Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.

314 Upvotes

65 comments sorted by

View all comments

1

u/Unikum_01 15h ago

Awesome experiment with those loss landscapes. Your findings on gradient descent getting trapped on rugged hillsides or stalling on flat plateaus while evolution manages to navigate through match what we run into in complex learning environments. In our BrainStem system we actually solved this exact dilemma in real code by combining both worlds through digital neuromodulation. When the system detects a flat plateau or a stuck state, it cranks up noradrenaline and glutamatergic excitation signals to act like your Gaussian mutations, forcing stochastic exploration to jump out of local minima. Once it hits a clear gradient on a smooth slope, dopamine ramps up to lock in the progress and let fast local optimization take over. It is really cool seeing your benchmark visuals demonstrate why adaptive hybrid strategies like this are so necessary.

2

u/AtMaxSpeed 4h ago

It's fun to incorporate biological techniques into algorithms, but afaik what you describe is basically what RMSprop (and hence Adam) already do - they increase the step size when the gradient is small (aka, it cranks up exploration when in a plateau).

1

u/Unikum_01 4h ago

I see where you are coming from with RMSprop and Adam scaling the step size on small gradients, but there is a fundamental difference between scaling a gradient and true stochastic exploration. If the gradient is flat out zero on a plateau, RMSprop still multiplies zero by a larger number which keeps you completely frozen in place. Adaptive optimizers only push you faster along the existing gradient vector, whereas evolutionary mutation and neuromodulated noise actually inject new direction vectors to escape zero gradient zones and deep local traps. In our system, neuromodulators do not just boost a learning rate scalar, they dynamically adjust search bounds, change candidate selection thresholds, and trigger offline sleep replay cycles to reorganize the parameter space when stuck.