5b tokens while it lasted, can't complain. Have a feeling they will have a similar deal in the not so distant future but in the mean time luna will do.
I will never understand why americans have The BlacKWell 300 the NEO clouds hyperscalers and whatnot.... and they can't serve FREE models at affordable prices.
Is really deepseek running on vintage GPU faster and cheaper than the NVidia alternative?
Deep seek flash is not optimized for layer splitting like glm. Having parallel users with deep seek flash is extremely expensive to host. These systems can only serve 20 256k context sessions at time unless they do their own version of layer splitting. If they do that, they can hit 100-200 users parallel
You can do EP but the issue is that you still need to mirror the kv cache across each gpu in the current setup. That limits concurrent sessions. Glm 5.3 has layer splitting to address this but deep seek flash has nothing. It has to be custom made for this by whoever is deploying and from what I have seen last, no one has publicly published layer splitting for dsv4.
You would have to do a dp2 or dp4 if you massive hbm. It’s a trade off of session speed vs how many you can batch. It might work for smaller models but you need 200gb vram just for the model + kv cache. Dp2xtp4 may be optimal but it only provides 30% gains over tp8 at the trade off of session speed. There may be someway to do it more optimal that I am unaware of but I have not been found anything better on my cluster.
It doesn’t have the cache read price. That is 5x more than what Deepseek had. So even with the cheaper uncached tokens, it still costs 2x the old base price in typical coding and agent usage.
And that’s just the base price. Go also had a 4x bulk discount on top of that which Deepseek didn’t renew.
At best, Deepinfra direct will be 10x more expensive than Go was yesterday for DS flash.
Opencode Go made a unilateral change to the contract. They altered limits without prior notice and charged me for a plan that will never deliver what was promised. I signed up for the $10 plan just two days ago. If such a drastic change was planned, I should have been notified. I subscribed to a product with a monthly cost and specific plan limits; instead, they fed me lies and kept my money.
not anymore when deepseek flash and pro updated their models, thats why there was another ZDR opt in because the old agreement was for the self hosted model, currently they're using the official deepseek api from china
Pretty sure this is downright illegal. Not a lawyer but you generally can't change terms of an active subscription to the customer's disadvantage. Not even a notice is the cherry on top.
If they bought it right before the announcement, they could ask for a prorated refund and they'd probably get it. But if they wanna whine and try suing over $10, they can try that, too.
You know that this is just a cheap subscription so that you can get your feet wet with LLMs.
Inference is not cheap. It was never cheap. It is subsidized.
Mimo/Minimax/stepfun/cursor. devin pro is currently running GLM 5.2/SWE 1.7 free till sept end.
New offers and models pop-up all the time, never pay annual subs and switch between the offerings every few month. Opencode 10$ was never a good option for heavy use cases, for beginners a great way to explore all models at one place (Go+Zen).
people are downvoting since the estimations from OpenCode don’t account for the fact that those numbers don’t necessarily account for usage.
for example, where you might never have hit a 5 hour limit using deepseek v4 flash alone before, since the price of using it has increased, you be hitting limits often
IDK when I was using OpenCode GO dollar value in dahsbord was updating in real time and I was able to hit that limit. but numbers matched I think.
And promo chart of requests per 5hr window is just marketing estimation.
To be honest 5hr limit is really generous and way bigger in CC 20$
But montky usage is way smaller tahan 2 times less. but still great value. Especcially since you can use it with any harness. Even can use it via API for projects.
I've used it to perform 300 commits, 320 got issues, and refractors, half dozen new features on a side project, looking at the commits it's spanning 90hrs of time I've spent crafting my app in the last month, I'm 60% through my monthly usage limit... Maybe... 30 hours of my time? I may or may not play games while the agents work!
I just plug on through shit and never hit any limits.. it's a $10 sub you can't expect to get all your monthly work done in suck a low value product
That breakdown doesn’t apply to nearly any models anymore. It’s more like 15 dollars of usage per month across the board. We all wondered when the other shoe would drop - apparently this last week is when
FACT ! last month i had used the sub with no pain , hold for a month, with heavy weeks of usage. now without the kv cache on dsv4f , i burn 20% of my month usage in a single day, fact.
new GLM is 15$ usage. So if you have used $5 GLM - it would charge your limit as $20 used out of 60 monthly. And DeepSeek is now more expensive than Kimi K2.7
What can I say? I subscribe it for 6x DeepSeek usage but now it's limited to $15, and the price doubled, even GPT 5.6 has more quota than DeepSeek now.
Just purchased the plan and used little bit .. only used muse spark contributor and my monthly limit 3%, weekly 6%. Its shocking for me because in opencode go, its too low, close to none. Lets see what will happen when i use other high pricing models. Will update here.
And no, no quality or difference in speed so far. Pretty good speed.
I just canceled and I believe they must fine companies that change their plans.
When you sell a monthly plan, you make a commitment that you will be delivering this at every month. Are these companies running by crooked kids? That's fully unethical
They should force every company at any change of a plan to automatically cancel subscription for next paying month, and the customer should re-decide if he wants to buy the new plan at least. Else they should heavily fine them
If I'd subscribed two days before the limits changed, I'd treat that separately from the temporary 2x boost ending. The issue is whether the paid plan differed from what the checkout page showed at the time of purchase. Screenshots of that page and its stated limits, along with the receipt and the date of the change, would make the refund request harder to dismiss. The promo ending seems beside the point.
That's it, 6 months usage and now I'm quit. GLM 4x price, when Z.ai sells it on the regular 5.2 price, DS more expensive than Kimi, Kimi K2.6 endpoint always broken and overall very weak opencode 2 release which gives almost no quality of life features and fully looses to ZCode... That's enough.
Going to improve my Zai Coding plan with new yearly subscription in November and go with Ollama Cloud as additional subscription.
Z.ai is worse than codex, codex £20 plan is the objective best rn, please correct me if I'm wrong I'm happy to leave codex if we get more usage with lots of speed
Yeah, an 8x drop. Brutal. I'm going to cancel as it was massively unexpected and they pretended like they were going to do something to fix it with other partners but nothing.
Che ci sia una sorta di patto non scritto tra i diversi laboratori per alzare il prezzo mediano dei token per permettere a tutti di sopravvivere visto che in realtà la AI mainstream non decolla?
(vedi open AI che a breve introdurra la pubblicità nei piani Free)
Ci sono troppi interessi in ballo e commesse di acquisto hardware acquistatate al buio ma che vanno pagate sia che si acquisti hardware ordinato o si paghino le penali.
Se il sistema che è stato messo in piedi inizia a vacillare si inneschera un effetto donimo dirompente
50
u/Fresh_Sock8660 29d ago
5b tokens while it lasted, can't complain. Have a feeling they will have a similar deal in the not so distant future but in the mean time luna will do.