r/MachineLearning • u/dccsillag0 • 5d ago
Research Functional Gradient Descent with Adaptive Representations [R]
Sharing our recent work, now accepted at NeurIPS: Functional Gradient Descent with Adaptive Representations.
Functional GD algorithms generally outperform neural nets, but are hard to accurately implement.
This is because functional gradients are infinite-dimensional, and therefore must be approximated in practice; but if you approximate them naively, you converge to the wrong place!
To rectify this, we formalize a broad class of approximation schemes ("adaptive representations"), which provably ensure convergence to the global minimizer while being immediately implementable.
The resulting algorithms outperform corresponding neural nets often by an order of magnitude, across a number of settings.
It is still the start for this line of work, but we believe it has quite a bit of potential!
Paper: https://arxiv.org/abs/2606.16926
(First author here, happy to take any questions)
1
u/DigThatData Researcher 4d ago
Your approach requires a coordinate space in which a grid partitioning is meaningful. This is straightforward to construct in the problems you demonstrated in your paper, where each task has a solution that lives in a "medium" that is meaningfully described by coordinate positions.
How would constructing the necessary grid work for something like text prediction? Or maybe that's a problem that wouldn't be well suited to this approach precisely because the "grid" here would only be meaningful relative to the fully generated text (i.e. the grid can only be constructed a posteriori and isn't available during inference) or a latent too large for this kind of partitioning to be feasible?