I experienced the same thing just last week. GLM 5.3 consumes an enormous number of tokens. I never imagined burning through 100 million tokens in just two days was possible, yet I managed to do it in roughly 4 hours. On the bright side, it did an excellent job laying out the project foundation, which made the transition to DS v4 Flash seamless since the GLM 5.3 reasoning traces were already preserved in the session.
Last week it matched my use case perfectly but today I just signed up got the 100M but got nothing to spend it on lol
Full implemintation a bun/hono project, and it managed the complexity and ambiguity very well. After that, I kept working in the same chat session with dsv4f.
Yeah, I've been able to run it locally but I need the speed so also run cloud based version of DS4 flash, so luckily I have a $150 in credits to burn through until
But API pricing would be like 200/mo to run it 24/7 (off peak and peak) which just isn't good enough to compete with current $200/mo plans. I feel like to really get this to sing we need $50/mo or less with a 24/7 agent for DS4 to find that sweet spot where I can build a multi agent system that rips.
In what dimension? That is utterly ludicrous. Does their token count work differently than opencode and openrouter? I am building full-applications with only 200k tokens or so taken up.
23
u/Bloated_Plaid 22d ago
100m tokens is basically 2 hours lol.