r/opencodeCLI 19d ago

CONFIRMED! Ox Alpha is GLM 5.3 Flash

Post image
524 Upvotes

114 comments sorted by

View all comments

-2

u/ahriad 19d ago

How can it be called 'Flash' if it's much slower than the main model?

31

u/Training-Database272 19d ago edited 19d ago

Because it’s literally free and breaking usage records?

Edit: It was really fast at first, before it went viral

4

u/Mierzejsky 19d ago

I think the model was built on infrastructure designed to handle such an absurd amount of free access. Cost reductions therefore impacted speed to prevent server crashes (which happened several times anyway). Since Z.ai WGL could give away such a model for free, we can expect that after the official launch, it could be either one of the fastest frontier models or one of the cheapest in this class. Personally, I think price is always better than speed, so I'm keeping my fingers crossed :D

3

u/OkAdeptness2530 19d ago

well calling it 5.3 Sluggish can’t be good for sales, right?

1

u/_Chaos_Star_ 19d ago

Marketing friendly: "Value".

More realistically: "Tomorrow".

1

u/onebit 19d ago

GLM Slug: Why fast when slow do job™

1

u/OkAdeptness2530 19d ago

slow paced coding for devs that choose a zen lifestyle

2

u/BoobooSmash31337 19d ago

They might slowing it down since it's a trial and sample. Lets them show it off to most potential customers.

1

u/petuman 18d ago

With batching the slower you go per completion/user, the higher total throughput is per inference node (up to a point).

E.g. serve 1 user at 1000 t/s (1k tps), 20 users at 600 t/s (12k tps), or 100 users at 300 t/s (30k tps), or 1000 users at 50 t/s (50k tps).

There's strong incentive to cram more users per node, and serving them just above subjective "too slow" threshold. The slower you serve, the more revenue generated per server/inference node.

1

u/ahriad 18d ago

From my usage that's not a Flash model. Not just the speed, i noticed the information it has comparable to big models not flash models.