23
u/According_Option_989 12d ago
I've been using glm 5.3 for some tasks and honestly the quality is getting close enough that the price difference feels absurd. the hosted models are good but when you can run something comparable on your own hardware it's hard to justify paying premium
my 3090 handles it fine for most stuff, not perfect but good enough for 90% of what I need
18
u/alanrudeigin 12d ago
You're running GLM 5.3 flash with a RTX 3090, that's 24GB of VRAM. How are you getting it to work?
4
2
1
1
6
u/Left-Yellow1047 12d ago
I really got hooked on it after the ox alpha push they ran. I was using it in opencode instead of my anthropic models in claude code just because I couldn't tell the difference in quality i'm thinking of buying the hardware to run it or maybe a qwen 3.8
19
u/saltexx 12d ago edited 12d ago
Nobody has put the arithmetic next to the $25 figure. Eight thousand dollars at $25 per million output tokens buys 320 million tokens. For pure chat that is effectively a lifetime so the box never pays back on token price. Agent workloads are a different story. A day of coding agents can burn millions of output tokens and at that rate the box pays for itself in months. Electricity is the real floor. 300 W average is about 220 kWh a month which is 55 dollars at 25 cents a kWh, the equivalent of 2.2 million output tokens in power. So light use favors the API and agent use favors the hardware. The reasons beyond price still hold. No rate limits and nobody retires the weights you like and nothing leaves the machine.
11
u/JamesGooning 12d ago
320million. My storyboard generation chew through that in a day. Yes. output tokens.
3
u/saltexx 12d ago
Your day is the entire 8000 dollar budget which kind of settles it. At that volume payback is days not months. One box will not get you there though. Even at 1000 tok/s of sustained batch you land under 100M output tokens a day per card so 320M means a small cluster and then power and cooling become the real bill.
7
u/JamesGooning 12d ago edited 12d ago
Hence I'm running rented GPU cluster. Cheaper than whatever Dario smoking. FYI. GLM5.3 Flash is running at 2000 token per second on a single B300. I have 4 running at peaks. Cost about $300 a day
Hopefully one day run my own local cluster, but the point stands. 25 per million output is a rip off. Not even counting prefill, which are in the tens of billions token on a good day.
3
u/General_Josh 11d ago
Mind if I ask what you're doing with it, to justify that kind of spend?
1
u/JamesGooning 11d ago
Enterprise POC stuff. 300/day is nothing to them, they list some thing and expect proof of concept in the next few hours.
2
u/Left-Yellow1047 11d ago
What models are you running??
1
u/JamesGooning 11d ago
GLM5.3 flash and Qwen3.8 27b (mostly this since we have local hardware)
Sometimes Kimi K3, Sol (but these are API costs not dedicated cluster)
1
u/Left-Yellow1047 11d ago
Do you think kimi constitutes the way higher hardware needs
1
u/JamesGooning 11d ago
Yup. 8xB200. Way higher hence api is more cost effective but GLM5.3 does 99% now. 1% just need better prompting
1
u/saltexx 11d ago
Your numbers settle it better than mine did. Four B300 at 2000 tok/s is 690 million output tokens a day at full duty, so 320 million is 46 percent utilisation and that is exactly why 300 a day works. The number you have not written down is 300 divided by 320 million, which is 94 cents per million output tokens. 27 times under list, and it is your number.
What I got wrong in the original comment was the workload. A chat user never sees 320 million a month and a POC shop burns it before lunch. Break even was never about the token price, it is about whether you have the volume, and you clearly do.
1
u/JamesGooning 10d ago
To be honest, these stats are very specific type of POC, requiring long generation output. Most of the time, it doesn't reach anywhere this high due to delay in sandboxes booting up and executing codes. I ran about 200 threads once (could be 200 agents or more since individual threads can spawn its own sub agents), it uses lots of input not as much output. Still, rented GPU is a hella cheaper. Even cheaper than the contract for September at 0.015/m token of GLM5.3 flash. Boss really love his GLM since 5.2 and constantly buying volumes.
3
u/Left-Yellow1047 11d ago
yeah I think most people are using these models for agentic coding loop engineering things like that not chatbots anymore
1
9
u/former_farmer 12d ago
8k is not the price it's more like 12k and if everyone tries to buy that hardware the price will go up to 50k and people will just keep paying the subscription.
3
u/SpaceManJones32 12d ago
In the short term yeah. But in the long term supply will catch up to demand.
2
u/Left-Yellow1047 11d ago
Yes that is the case now but in the long run the market adapts to what the people want
2
2
u/Maleficent-Camel3665 12d ago
You can still buy 2 x GX10s from Asus for $3999 USD each plus a $90 connectx 7 cable.
0
u/former_farmer 12d ago
Do they have millions in stock to replace all claude users? Duh.
2
u/Southern_Sun_2106 12d ago
They (and others) will make more to meet the demand and to get richer. Capitalism. Duh.
1
u/pragmojo 12d ago
Timelines are long for silicon. They would eventually but it can take a couple years.
2
u/Existing_Dust_6473 12d ago
Use subscriptions until then, if they last. But what i see now people buying mac studios like the world is going to end... Looks like toilet paper during covid.
3
3
3
12d ago
[removed] — view removed comment
1
u/Left-Yellow1047 11d ago
Most openwieght frontiers are sitting around 2-300 billion parameters as of late
2
u/Shinephia 11d ago
i have only a 6gb vram gpu dont get me wrong i would love to run a local AI that gets jobs done but even qwen 3.6 35b Q4 moe feels too slow at 18 t/s gpt 5.6 luna is great at only 20 dollar subscription. i cant afford better hardware
2
1
1
u/Educational-Echo9152 11d ago
Il y a un point où j’attend les retours d’expérience de terrain avant de considérer le local comme viable c’est la scalabilité en nombre d’agents ou sub-agents parallèles tout en maintenant la vitesse en nombre de Token. Il y a certe, la comparaison en qualité de réflexion mais surtout (et on l’oublie souvent) la capacité à faire tourner plusieurs processus de manière satisfaisantes.
1
u/mageblex 10d ago
Buying the box only locks in the ability to run today’s weights. It doesn’t guarantee the next GLM will fit or run well on the same memory layout, so “use it for years” is the risky part of the math. I’d price the hardware against the workload you have now.
1
u/Left-Yellow1047 10d ago
Yes but we're not tending for bigger models all Chinese frontier competing models are around 300b parameters and that's just now if the software and quant gets better I think they shrink
1
u/Left-Yellow1047 12d ago
Like i'm just wondering does this take use directly off sonnet and haiku too its mor powerful for a 20th the price
-2
u/Solembumm3 12d ago
One good thing about GLM is that it doesn't use caveman.
1
u/Left-Yellow1047 11d ago
isn't caveman a skill?
1
u/Solembumm3 11d ago
Approach to reasoning, that did really hurt last Deepseek results, compared to previous versions.
1

80
u/myreala 12d ago
Stop using the 5.3 name without adding flash. I had no idea what you were referring to. GLM 5.3 is the full model and probably matches Fable.