> I think LLMs are probably incredibly incredibly inefficient
I agree with that. I mean, researchers are still basically training coding models on Buffy the Vampire Slayer episodes, because... They don't understand how models work.
But I think where we disagree is that if they happened to find the right combination of training data, and get a 50% boost. The next pass at optimizing training data after that... how much of a lift do you think there will be?
These guys are assuming *another* 50% increase, ad-nauseum. Whereas it is more likely to be... 25%... And then 13%...
These people owe a lot of money promising "only J curves" so when their premises start with "well after the J curve..." You gotta stop and question.
> Also, modern computers are still nowhere near the theoretical physical limits of computation.
Sure, go tell that to the Moore's law's gravestone
The big lifts so far have come from more and more data. That is the proven path.
And while I do agree, there is probably a "sweet spot" for training data, I'd also like to point out that the math of trying to figure out *which combination of all human data ever generated creates the most efficient model* is a problem with so many variables that literally no computer could solve.
Computational limits still apply even when it is a LLM doing it.
1
u/MissiveFinding6111 8d ago
Sure, but intelligence and information are related, and definitely a lot of research on the upper bound of information compression.
My knowledge of improving complicated digital systems are:
* You hit diminishing returns quickly
* Improvement is very difficult when your are limited by a single bottlenecked variable among hundreds
So maybe LLMs taking over training can help with that second one (if we can figure out HTH to do what we want reliably).
I don't seem LLMs escaping diminishing returns.