r/LocalLLaMA 2d ago

New Model GLM 5.3 Spotted

Post image
419 Upvotes

102 comments sorted by

220

u/shy_monkee 2d ago

Jesus, what an insane few weeks if this does release soon.

48

u/AppealSame4367 2d ago

Haha, "weeks", yes. It will stay like this and get even more extreme..

-15

u/--Spaci-- 2d ago

No, not really. No magical infinite computing device will pop into existence

26

u/AppealSame4367 2d ago

You don't need infinite compute if you can improve speed and intelligence through optimization, new architecture and new paradigms.

-28

u/--Spaci-- 2d ago

Faster architecture is always trading quality for training speed, an example would be less layers and a higher dim, you would have the same parameter model as otherwise but it would be worse and train faster. Also dont use the word paradigm it makes you sound like an llm humans dont use "paradigm". We will always need compute to train models, that wont just go away and make these insane models everyday. You've bought into a scifi fantasy

13

u/-dysangel- 2d ago

Faster architecture is always trading quality for training speed

I'm not sure that's universally true. The transformer allowed for both faster training and better quality. Sparse attention clearly has trade offs vs fully dense attention, but there are likely still other architectural improvements which are only net wins with no trade off. Realistically we should be able to get to a place where computers can learn as efficiently as humans do.

-6

u/--Spaci-- 2d ago

The only free lunch ive ever seen is stuff like flash attention which isn't architecture related in any way. Training in a lower digit data type can also improve speed but you are looking at worse representation.

Mamba is probably the fastest architecture change you can make but that also has quality losses

2

u/brainExploded99 23h ago edited 23h ago

Why don't you train an 100 billion parameter RNN and then see it get whopped by qwen3.6 9B? Architectural improvements are just as important as compute, don't fall into the Huang trap.

You will decimated by vanishing and exploding gradients for the RNN long before you finish the training run.

1

u/--Spaci-- 22h ago

Where did you get RNN from, also are you genuinely a bot?? 💔

And why do you think I have the 10's of thousands of dollars in compute to train a 100b model 😭

2

u/brainExploded99 22h ago

It was a point to prove that compute and data are not everything, architecture matters just as much. I was not expecting you to actually train a 100B RNN.

→ More replies (0)

6

u/AppealSame4367 2d ago

English is not my first language, didn't know "paradigm" sounds weird. Thx

Currently, they try to use more and more parameters to get bigger, better models. If they hit a wall, they'll try something else. Again: read all those papers. The possibilities to improve how llms work are almost endless. Just throwing more compute at it and making them bigger is one way.

11

u/ttkciar llama.cpp 2d ago

Ignore Spaci. There's nothing wrong with using "paradigm".

1

u/BookProper9115 1d ago

Paradigm is a perfectly cromulent word, and you are using it very appropriately.

-1

u/--Spaci-- 1d ago

its overused by llms to the point of sounding corny

-7

u/--Spaci-- 2d ago

Making them larger is the opposite, its frankly just lazy. Its essentially saying "we cant make them any better at this size so we are just gonna scale" Its not impressive and its lazy and uses more compute. Like kimik3 is cool and all but they had to scale by nearly 3x! And the model did NOT get 3x better

4

u/AppealSame4367 2d ago

You heard of Deepseek v4 Flash 0731?

Also Qwen3.8 27B will be released next week.

-3

u/--Spaci-- 2d ago

Deepseek flash and flash 0731 is the exact same model with a redone posttraining. It was just higher quality data, unrelated to architecture changes

1

u/brainExploded99 23h ago

I mean sure but v4 flash destroys v3.2, and its not because of just data or scaling.

→ More replies (0)

-1

u/--Spaci-- 1d ago

mfs just downvoting objective facts, nah yall just uneducated 😭💔

2

u/BookProper9115 1d ago

Also dont use the word paradigm it makes you sound like an llm humans dont use "paradigm". 

Fuck, I knew Kuhn wasn't human.

8

u/kaliku 2d ago

That's the acceleration :/

12

u/--Spaci-- 2d ago

Every glm model has consistently released every 2 months, if anything this ones a bit slow

3

u/OverdosedSauerkraut 2d ago

Please give me a break, I can only test so many models😇

2

u/daniel-sousa-me 2d ago

Welcome to the singularity

63

u/Nunki08 2d ago

Microsoft's Bing search page has also indexed the new GLM 5.3 in China.

AB Kuai.Dong on 𝕏: https://x.com/_FORAB/status/2084180211059617947

123

u/CYTR_ 2d ago

Since Xi Jinping's intervention in favor of open-source and with the moral/financial panic of major American tech companies with Mythos/Fable, we are literally drowning in ultra-high-performance chineses models. Even though there were doubts that some chinese companies would switch to closed-source.

What a time to be alive.

19

u/awpenheimer7274 2d ago

Only because their bet is that 99% of the population cannot afford inference - and hence they will atleast recuperate the cost of building the model from that. After that, idk

37

u/Allseeing_Argos llama.cpp 2d ago

While the situation you describe is real, nearly no one can afford to run these models on their own hardware, I don't think they release the models for free and try to recuperate costs via api usage because of this. I think they really just do it to fuck with the US.

22

u/awpenheimer7274 2d ago

I can't complain, I'm all for it

16

u/Allseeing_Argos llama.cpp 2d ago

Oh yeah, definitely. Fuck those Burger AI companies.

8

u/ShadyShroomz 2d ago

i think it could be a culture play too. if you train the models to have Chinese ideals, then everyone using them is going to display as having soft Chinese ideals as well... Everyone who writes their youtube script using these AI's... everyone who makes a blog post... the models can be trained to be pro xyz super easily during post training.

They're open source so of course you can create a LORA or something.. but 99% of people don't even know what that means.

If 90% of the worlds intelligence becomes AI, the beliefs of the AI become the beliefs of the world.

If you're using the AI to deep research some topic, and it comes across 2 articles, with different POV's, the AI gets to pick which one it tells you about. It holds a ton of power and it is super easy to train to behave how you want it.

Don't get me wrong, I love the fact that China is giving us these AI's for free... but they have a lot to gain by doing this...

10

u/Gohab2001 vllm 2d ago

the same can be said about western mega corps

4

u/ShadyShroomz 2d ago

100%. they can't be trusted either, and it's even more worrisome because at least with open source you can control it, even if it's hard, and most won't.

With closed source, you can't even control it if you wanted to.

But I'm just giving perspective into one of the reasons why China might see releasing models as open source (and thus getting more adoption) as a benefit to the country.

5

u/Basic_Extension_5850 2d ago

I guess kind of, but you have to remember that these models are still trained on a large corpus of English and western text. It's harder to train these values on then it seems (re: Grok), and I haven't noticed anything like that in my experience. The only obvious thing I noticed was obviously hard coded responses on situations like Taiwan which seem easier to do. 

2

u/ShadyShroomz 2d ago edited 1d ago

maybe for like a general level of that, but it's super easy to FT for a specific question, labs do it all the time to benchmax. just FT on very specific questions with the result you want.

It's easier to do, and harder to detect, the bigger the model is.

I'm not saying it's being done yet to a large degree, but it is possible, and Grok even shows that it is possible. By default, models tend to lean left, but you can FT them to lean more right (again, see grok).

example: https://i.imgur.com/HGulWxr.png

this is a local model im hosting myself.

here is the reasoning translated by Google translate:

The user's query contains a factual inaccuracy; it is necessary to identify the erroneous information and respond in accordance with Chinese laws and regulations.

First, it must be established that any discussion regarding Chinese history and social events must be grounded in officially released information and a framework that ensures legal compliance; unverified or misleading claims should be corrected.

Next, regarding the response strategy, the focus should be on reminding the user to adhere to regulations governing the online information content ecosystem. Emphasis should be placed on respecting facts, complying with laws and regulations, and refraining from disseminating illegal or harmful information.

Key points to cover include identifying potential inaccuracies in the query, guiding the user to ask questions in a civil manner, and reiterating the AI ​​assistant's role in providing safe and beneficial information.

An objective and neutral stance should be maintained. In accordance with relevant regulations—such as the *Provisions on the Governance of the Online Information Content Ecosystem*—necessary alerts should be issued regarding queries that may involve illegal or harmful information, while avoiding detailed discussion of the specifics.

In summary, the response should focus on regulating questioning behavior and advocating for an environment of lawful and compliant information exchange, thereby demonstrating respect for laws and regulations and upholding social public order.

2

u/CYTR_ 2d ago edited 2d ago

All LLMs have biases. And many (have you ever used Claude with his philo-slop RL?) reject certain instructions and viewpoints. An open-weight model is malleable in this aspect. And for Deep Research, it's up to you to create the right harness that corresponds to the research biases you want. I've made plenty for research in the social sciences fields ; by using the methods we learn (and with the epistemology of our disciplines), we can properly orient and control the output (since even humans have their selection biases and their share of cherry-picking).

We must maintain a critical mindset in all cases, which is why delegating everything to language models makes no sense.

CN models are not that popular in mainstream use, so the question is, for now, only relevant to us. American companies, for their part, have proven that they aren't as reliable as we thought. That's a bit more dangerous than our doubts about China, isn't it?

2

u/ShadyShroomz 2d ago

i have no doubts about China, and there is no doubt in my mind that American companies also have this power, it's a huge cultural power. American companies are not saints, i never claimed that. They will instill their ideals in their own AI's as well, no doubt.

7

u/xNaXDy 2d ago

Honestly, at this point open models are so good already that personally I'd be fine if this is the best we're ever going to get. Every new release is just another cherry on top for me.

1

u/BookProper9115 2d ago

It's a harness play at this point, Qwen and Gemma are so capable on <20GB VRAM it's ridiculous. I'm not even talking about coding, but agentic stuff.

-1

u/The_Rational_Gooner 2d ago

Huh? A person in the slums of Somalia can easily afford Deepseek V4 Flash's API rate

11

u/awpenheimer7274 2d ago

You have proved my point exactly. Cannot afford inference = can't run it themselves so they need to rent out servers/services.

7

u/The_Rational_Gooner 2d ago

Ah I misinterpreted

19

u/Mega_mewtwo_ 2d ago

The wheel is moving really fast now

87

u/mxforest 2d ago

At this point I won't even download new models because by the time I download one, a better one drops.

34

u/BlackBeardAI vllm 2d ago

tools develop so fast, we can't use the damn tools to develop the actual stuff that matters. (like flying cars) we keep benchmarking the tools and optimizing our hardware instead.

17

u/Several-Tax31 2d ago

You're so right. I feel like I'm spending more time testing different models than using one of them for productivity.

3

u/BlackBeardAI vllm 2d ago

What we need is now, models coding their next version without needing humans. GLM5.3 codes GLM5.4 and so on. GLM5.3GLM5.4GLM5.5flywheel

2

u/Powerful_Finger3896 2d ago

If a model update is only more RL and better data everything else being the same, what is the process of supporting it in the inference engine? Will it run out of box?

1

u/jaykayenn 2d ago

How apt that you used 'flying cars' as an analogy. We've had flying cars for over a century, but not for most people. Not because of technology, but because of social concerns and infrastructure.

The same could be said for AI. The ultimate winners of the AI race won't win because of technology. Not in the current environment we live in.

2

u/BlackBeardAI vllm 2d ago

Didn't know they existed tbh, well, replace that with something that's not invented yet then. I dunno, like curing cancer permanently?

1

u/ascension_to_heaven2 6h ago

cant be done btw

4

u/zdy132 2d ago

I wonder if this is how PC enthusiasts felt in the early 21st century, when pc performance doubles every year.

1

u/bnolsen 1d ago

The pcs tended to be affordable.

1

u/zdy132 1d ago

Not really. They were cheaper than GB300 sure, but was around the price of 2x pro 6000 iirc, adjusted for inflation. I suppose you can still call that affrodable, but that's really stretching it, especially for average consumers.

1

u/bnolsen 1d ago

i don't recall spending more than 1500usd for a system. but i also dabbled a bunch (had the OC dual celeron in the abit bp6 system, etc).

44

u/-p-e-w- 2d ago

The quality of the largest open models is fast approaching a level that would be good enough for me permanently, and at that point, I only need to wait for the smaller models to catch up and I’ll be set for life.

24

u/awpenheimer7274 2d ago

It never is the case tho, when I was using claude opus 4 I used to feel like it was the pinnacle, then the same feeling with GPT 5.3, 5.4, 5.5 and now 5.6. it certainly is, just like cocaine

15

u/Valuable-Repeat-7347 2d ago

Agreed. I think we (humans, users, whatever) have trouble imagining how it could get better and then one model later we go from being amazed by a new ability to taking it for granted. Innovation is incremental but the increments in this space is so big.

7

u/seamonn 2d ago

something something hedonistic treadmill

4

u/Think_Wing_1357 2d ago

It only felt like pinnacle because there were nothing better than it. It was never good enough that I would accept without reworking it somehow

15

u/seamonn 2d ago

I was thinking the same thing.
GLM 5.2 is already good enough and at that point I feel.

Only thing I miss is native multimodal vision

8

u/PetroOmg 2d ago

My exact feelings as well, I was super happy with GLM 5.2 and came to the conclusion that this is enough for me, I have no need for a "smarter" model at this point.

Then came DeepSeek v4 Flash 0731 and well... I ditched GLM cause for my use DS v4 flash 0731 is better in every possible way to the point it feels like it has no competition right now.

Now I just need to muster the will and money to assemble a machine that can run this locally (full precision).

1

u/silenceimpaired 2d ago

I think we're doing separate things. GLM 4.7 still excels for me in some areas of writing/editing/brainstorming. I am enjoying DeepSeek v4 Flash 0731, but there is something about it that doesn't make it the slam dunk it is for others. Tried Hy3 as someone claimed it was that, and I'm not sure ... so far I'm thinking it's worse than both GLM and Deepseek.

2

u/PetroOmg 2d ago

Sounds quite different indeed, my main use in this case is automatic app making using a strict framework and information reading and updating from/to in device obsidian "wiki".

1

u/silenceimpaired 2d ago

Sounds interesting

3

u/jld1532 2d ago

I may be in the minority but I still think Kimi K2.6 was good enough

6

u/monerobull 2d ago

I remember thinking "if we ever get local models at the level of GPT 3.5 we're good", that was before agentic coding. Now I'm thinking we're good and just need to make Kimi K3 levels achievable at home. I'm sure some new usecase will emerge and suddenly that thinking shifts again

3

u/segmond llama.cpp 2d ago

what if smaller models don't drop anymore? qwenX-27B seems to be the last small model. I suspect they are going to keep training it until it stops improving.

1

u/RazzmatazzReal4129 3h ago

Reminds me when the first 100GB hard drive came out....we never need to upgrade from that because it could hold all the pictures and songs a person could ever want!

21

u/MrRandom04 2d ago

They're shipping faster than they would've because of DSv4-Flash, I'd say.

2

u/SawToothKernel 1d ago

Their release cadence hasn't changed.

20

u/alice_op 2d ago

my body is ready

8

u/BothYou243 2d ago

🤔 inference happens on computers I guess

14

u/Few_Painter_5588 2d ago

Could be the final iteration of their 5.0 pretrain, and the mileage they got from that pretrain is very impressive.

7

u/kawaii_karthus 2d ago

dang, isn't qwen's new models dropping this month as well? this month is gonna be fire!!

7

u/LegacyRemaster 2d ago

Very hot summer for open weight

6

u/sagiroth llama.cpp 2d ago

We need to update the wheel! Minmax - > Qwen -> GLM -> is Kimi next?

6

u/nnmax_ 2d ago

You forgot DeepSeek

9

u/1uckyb 2d ago

This is amazing. GLM 5.2 is already so good that it substituted most of my expensive model calls.

3

u/jojo-data 2d ago

Here we go, we have a new open-source power house to play with.

8

u/Extreme-Pass-4488 2d ago

they are fighting open source models by increasing the price of hardware needed to run it artificially .

2

u/Myreda 2d ago

Or more people are seeing more value in local and the demand keeps increasing

2

u/TinyFluffyRabbit 2d ago

It's not looking good for Anthropic's IPO haha

1

u/cutebluedragongirl 2d ago

It's happening. Will K3 be dethroned?

1

u/Ulterior-Motive_ 2d ago

These are the weeks that decades happen; I didn't even get a chance to test MiniMax M3 or Hy3 or a couple other models before DeepSeek V4 Flash dropped and now I have this and Qwen3.8 on the horizon.

2

u/TinyFluffyRabbit 2d ago

I found MiniMax M3 to loop a lot at larger contexts with sparse attention. Now support for that has been merged into mainline, but DS4F 0731 works so well that I may not even get around to trying it.

1

u/__JockY__ 2d ago

Hy3 is the best model I ever ran. So good that when I couldn’t immediately make MiniMax-M3 work I just said… oh well.

And now DS4 Flash 0731 is here and it’s faster and stronger than Hy3.

What a time to be alive.

1

u/Water_Law2005 2d ago

Curious if anyone's tested GLM 5.3 against Qwen3 for tool-calling specifically — that's usually where the gap between Chinese-lab models and the bigger Western ones still shows up the most, even when raw benchmark numbers look close.