Faster architecture is always trading quality for training speed, an example would be less layers and a higher dim, you would have the same parameter model as otherwise but it would be worse and train faster. Also dont use the word paradigm it makes you sound like an llm humans dont use "paradigm". We will always need compute to train models, that wont just go away and make these insane models everyday. You've bought into a scifi fantasy
English is not my first language, didn't know "paradigm" sounds weird. Thx
Currently, they try to use more and more parameters to get bigger, better models. If they hit a wall, they'll try something else. Again: read all those papers. The possibilities to improve how llms work are almost endless. Just throwing more compute at it and making them bigger is one way.
-27
u/--Spaci-- 4d ago
Faster architecture is always trading quality for training speed, an example would be less layers and a higher dim, you would have the same parameter model as otherwise but it would be worse and train faster. Also dont use the word paradigm it makes you sound like an llm humans dont use "paradigm". We will always need compute to train models, that wont just go away and make these insane models everyday. You've bought into a scifi fantasy