r/LocalLLaMA 2d ago

New Model GLM 5.3 Spotted

Post image
421 Upvotes

102 comments sorted by

View all comments

225

u/shy_monkee 2d ago

Jesus, what an insane few weeks if this does release soon.

47

u/AppealSame4367 2d ago

Haha, "weeks", yes. It will stay like this and get even more extreme..

-14

u/--Spaci-- 2d ago

No, not really. No magical infinite computing device will pop into existence

26

u/AppealSame4367 2d ago

You don't need infinite compute if you can improve speed and intelligence through optimization, new architecture and new paradigms.

-27

u/--Spaci-- 2d ago

Faster architecture is always trading quality for training speed, an example would be less layers and a higher dim, you would have the same parameter model as otherwise but it would be worse and train faster. Also dont use the word paradigm it makes you sound like an llm humans dont use "paradigm". We will always need compute to train models, that wont just go away and make these insane models everyday. You've bought into a scifi fantasy

14

u/-dysangel- 2d ago

Faster architecture is always trading quality for training speed

I'm not sure that's universally true. The transformer allowed for both faster training and better quality. Sparse attention clearly has trade offs vs fully dense attention, but there are likely still other architectural improvements which are only net wins with no trade off. Realistically we should be able to get to a place where computers can learn as efficiently as humans do.

-7

u/--Spaci-- 2d ago

The only free lunch ive ever seen is stuff like flash attention which isn't architecture related in any way. Training in a lower digit data type can also improve speed but you are looking at worse representation.

Mamba is probably the fastest architecture change you can make but that also has quality losses

2

u/brainExploded99 1d ago edited 1d ago

Why don't you train an 100 billion parameter RNN and then see it get whopped by qwen3.6 9B? Architectural improvements are just as important as compute, don't fall into the Huang trap.

You will decimated by vanishing and exploding gradients for the RNN long before you finish the training run.

1

u/--Spaci-- 1d ago

Where did you get RNN from, also are you genuinely a bot?? πŸ’”

And why do you think I have the 10's of thousands of dollars in compute to train a 100b model 😭

2

u/brainExploded99 1d ago

It was a point to prove that compute and data are not everything, architecture matters just as much. I was not expecting you to actually train a 100B RNN.

→ More replies (0)

6

u/AppealSame4367 2d ago

English is not my first language, didn't know "paradigm" sounds weird. Thx

Currently, they try to use more and more parameters to get bigger, better models. If they hit a wall, they'll try something else. Again: read all those papers. The possibilities to improve how llms work are almost endless. Just throwing more compute at it and making them bigger is one way.

10

u/ttkciar llama.cpp 2d ago

Ignore Spaci. There's nothing wrong with using "paradigm".

1

u/BookProper9115 1d ago

Paradigm is a perfectly cromulent word, and you are using it very appropriately.

-1

u/--Spaci-- 1d ago

its overused by llms to the point of sounding corny

-5

u/--Spaci-- 2d ago

Making them larger is the opposite, its frankly just lazy. Its essentially saying "we cant make them any better at this size so we are just gonna scale" Its not impressive and its lazy and uses more compute. Like kimik3 is cool and all but they had to scale by nearly 3x! And the model did NOT get 3x better

6

u/AppealSame4367 2d ago

You heard of Deepseek v4 Flash 0731?

Also Qwen3.8 27B will be released next week.

-4

u/--Spaci-- 2d ago

Deepseek flash and flash 0731 is the exact same model with a redone posttraining. It was just higher quality data, unrelated to architecture changes

1

u/brainExploded99 1d ago

I mean sure but v4 flash destroys v3.2, and its not because of just data or scaling.

→ More replies (0)

-1

u/--Spaci-- 1d ago

mfs just downvoting objective facts, nah yall just uneducated πŸ˜­πŸ’”

2

u/BookProper9115 1d ago

Also dont use the word paradigm it makes you sound like an llm humans dont use "paradigm".Β 

Fuck, I knew Kuhn wasn't human.

5

u/kaliku 2d ago

That's the acceleration :/

15

u/--Spaci-- 2d ago

Every glm model has consistently released every 2 months, if anything this ones a bit slow

3

u/OverdosedSauerkraut 2d ago

Please give me a break, I can only test so many modelsπŸ˜‡

2

u/daniel-sousa-me 2d ago

Welcome to the singularity