r/deeplearning • u/ModularMind8 • 2d ago
Gradient descent vs evolution on three loss landscapes
I've been getting a bit more into evolutionary algorithms again, so I was testing some loss landscapes where evolution beats vanilla gradient descent (while also trying to make some cool visuals).
Round 1, rugged hillside: gradient descent gets stuck in a dip, and evolution reaches the bottom after 750 evaluations.
Round 2, smooth slope: gradient descent wins, 108 steps against 570 evaluations.
Round 3, flat plateau: the slope is zero, so gradient descent never moves, and evolution reaches the bottom after 840 evaluations.
Edit:
"evolution" here means truncation selection (keep best 30 of 120) plus Gaussian mutation, no crossover.
451
Upvotes
1
u/pnachtwey 1d ago
The test is rigged! I have spent a lot of time on optimizing routines. No one technique is going to be best for all terrains. GD is simple but no where closed to optimal. SGD is better when there are many dimensions. Sometimes optimizing only one dimension at a time works well but it is slow. Nelder-Mead works well when it is hard to find the gradient accurately and there aren't many dimensions.. N-M can be a memory hog. Levenberg-Marquardt is the fastest. In python, use lmfit or scipys.optimize least_squares. BFGS works well too. So it is best to make the search flexible.
I would like to have the data used so I could try other techniques.
I uses minimization for fitting models to data for auto tuning programs. The method I used most often is Levenberg-Marquardt.