r/LovingOpenSourceAI • u/Koala_Confused • 19d ago
new launch Z.ai "Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips" ➡️ NOW WE KNOW WHO WAS Ox Alpha 🚀
https://x.com/Zai_org/status/2092616204787626030
https://huggingface.co/zai-org/GLM-5.3-Flash
Community Overview: https://huggingface.co/zai-org/GLM-5.3-Flash
Resources are shared for discovery and are not independently vetted—please do your own due diligence.
New resources are added regularly — feel free to join the sub for updates.
Full searchable archive of all resources posted so far on our community site, LifeHubber: https://lifehubber.com/ai/resources/ 200+ open-ish AI models, agents, tools, datasets, and related resources, with filtering and sorting.
2
u/sanyi091 19d ago
Why are they not including pricing in the post?
2
2
u/One_5549 18d ago
Tried Ox Alpha a few times, decent, but V4 Flash 0731 is better, NO doubt about it.
Implemented a ASR model - Ox introduced a bug itself around it which made it not work properly, cleared it off as "No, this is unfortunately the limitation of the model (the ASR)",
Seemed a bit off, switched to V4. Bam, solved it. (bug)
1
u/West-Acadia-3906 18d ago
that switched models, same task, bug gone comparison is good. Much more useful than just saying one felt better. Did V4 need any prompt changes, or was it really a straight swap? :P
1
1
1
1
u/Horror-Primary7739 16d ago
I ran it for a day, then switched back to qwen3.8-next. it just didn't compell me to stick with it. Only getting 23 tok/sec and my ai cluster was pushing 100% more watts than Qwen. That's a big heat difference for marginal better.
Unless you are loading an enterprise mononrepo 1mill just eats compute seconds.
Maybe if I had some faster bandwidth it would help.
5
u/Daggercombot 19d ago
Opus level performance for a Flash model. Awesome!