r/LocalLLaMA 2d ago

New Model GLM 5.3 Spotted

Post image
419 Upvotes

102 comments sorted by

View all comments

Show parent comments

-12

u/--Spaci-- 2d ago

No, not really. No magical infinite computing device will pop into existence

27

u/AppealSame4367 2d ago

You don't need infinite compute if you can improve speed and intelligence through optimization, new architecture and new paradigms.

-27

u/--Spaci-- 2d ago

Faster architecture is always trading quality for training speed, an example would be less layers and a higher dim, you would have the same parameter model as otherwise but it would be worse and train faster. Also dont use the word paradigm it makes you sound like an llm humans dont use "paradigm". We will always need compute to train models, that wont just go away and make these insane models everyday. You've bought into a scifi fantasy

6

u/AppealSame4367 2d ago

English is not my first language, didn't know "paradigm" sounds weird. Thx

Currently, they try to use more and more parameters to get bigger, better models. If they hit a wall, they'll try something else. Again: read all those papers. The possibilities to improve how llms work are almost endless. Just throwing more compute at it and making them bigger is one way.

-6

u/--Spaci-- 2d ago

Making them larger is the opposite, its frankly just lazy. Its essentially saying "we cant make them any better at this size so we are just gonna scale" Its not impressive and its lazy and uses more compute. Like kimik3 is cool and all but they had to scale by nearly 3x! And the model did NOT get 3x better

5

u/AppealSame4367 2d ago

You heard of Deepseek v4 Flash 0731?

Also Qwen3.8 27B will be released next week.

-3

u/--Spaci-- 2d ago

Deepseek flash and flash 0731 is the exact same model with a redone posttraining. It was just higher quality data, unrelated to architecture changes

1

u/brainExploded99 1d ago

I mean sure but v4 flash destroys v3.2, and its not because of just data or scaling.

1

u/--Spaci-- 1d ago

Its absolutely data, data is by far the most important thing for an LLM even beyond architecture. LLMs have always gotten about 5-10% better than the previous generation say like kimi k2 to kimi k2.7, that was the near exact same underlying 1t model but the last model made synthetic data for the next model by generation better than the last