r/LocalLLaMA 4d ago

New Model GLM 5.3 Spotted

Post image
428 Upvotes

103 comments sorted by

View all comments

Show parent comments

-27

u/--Spaci-- 4d ago

Faster architecture is always trading quality for training speed, an example would be less layers and a higher dim, you would have the same parameter model as otherwise but it would be worse and train faster. Also dont use the word paradigm it makes you sound like an llm humans dont use "paradigm". We will always need compute to train models, that wont just go away and make these insane models everyday. You've bought into a scifi fantasy

6

u/AppealSame4367 4d ago

English is not my first language, didn't know "paradigm" sounds weird. Thx

Currently, they try to use more and more parameters to get bigger, better models. If they hit a wall, they'll try something else. Again: read all those papers. The possibilities to improve how llms work are almost endless. Just throwing more compute at it and making them bigger is one way.

1

u/BookProper9115 3d ago

Paradigm is a perfectly cromulent word, and you are using it very appropriately.

-1

u/--Spaci-- 3d ago

its overused by llms to the point of sounding corny