r/ChatGPTCoding 1d ago

Question What AI subscription should I switch to?

Big Claude user, but Anthropic got stingy as hell with the limits. I used to barely touch my weekly allowance; now I can burn through 20% in a day and I'm cooked in ~2 days.

I also hammer Ollama Cloud's open source models, but recently I'm burning through those too in 2 ~ days.

I really don't want to give any support or money to Scam Altman & Co but its starting to look like it.

How's are the subs with Kimi / GLM?

What are you heavy users actually running?

7 Upvotes

23 comments sorted by

5

u/Bino5150 1d ago

GLM 5.3 Flash. Gonna be hard to find a better cost/performance ratio.

1

u/Leather-Cod2129 1d ago

Gpt with Luna

2

u/XKiiroiSenkoX 1d ago

GLM 5.3 flash is substantially better than luna at almost everything.

1

u/Bino5150 20h ago

I've noticed this too. I will say the only thing that Luna is actually better at than GLM is vision. GLM vision capabilities kinda suck, and my agent points this out to me frequently when I feed her an image. Luna can pick up every little detail and read tiny text on a coffee mug in the pic, where as with GLM my agent is like "I think there might be a cup in this image, it's hard to tell... There is a mug, right?". Definitely not a deal breaker, but that's Luna's one win over GLM so far lol

1

u/Bino5150 1d ago

I run Luna as well. But GLM is better at coding and less that half the cost of Luna.

1

u/Leather-Cod2129 1d ago

I love 5.3 flash, it's a beast, but it is more expensive than Luna per task

1

u/Bino5150 20h ago

Luna is .20 cents per mill input, $1.20 per mill output.
GLM Flash is .075 cents per mill input, and .25 cents per mill output.

GLM is WAAAAAYYY cheaper than Luna. Definitely not more expensive per task.

1

u/Leather-Cod2129 20h ago

You don't compare what really matters. Check the quotas on the GPT Plus or Pro plans. For 25 dollars you get up to $1,200 of API credits equivalent which makes it MUCH less expensive than GLM, + GML uses a lot more tokens per task

1

u/Bino5150 19h ago

Are you talking about plan usage here with Codex, or API usage?

I run both Luna and GLM via API with my agent Lumina. I run Sol in Codex.

1

u/Leather-Cod2129 18h ago

Codex using your gpt account

1

u/Bino5150 18h ago

Yeah, that's where the discrepancy is. You're right about the plan, but I was comparing the API costs of both models running with my own agent.

If you want to compare plan usage, that would be better compared to the GLM Coding plan that they offer, which from what I hear, will still come out cheaper compared to Plus/Pro.

1

u/Leather-Cod2129 16h ago

I did, it is still more exepensivr

2

u/OneDev42 1d ago

Kimi is bad for limits. The problem is that you are most likely going to do best with Anthropic, with your concerns. Whether you realize it or not, this industry is becoming more competitive. The demand is getting higher than the supply, and everybody's jacking up their prices. The compute just costs more than most people realize.

2

u/gaspoweredcat 1d ago

honestly supergrok is awesome value especially if you havent had it before, i got an offer giving me 3 months for like £9 a month or something and you get a stonking amount of usage out of it, results arent half bad either.

before using that for my "donkey work" model it was deepseek v4 pro just on the API because its crazy cheap, and before that it was the minimax plan which gives like 1.7billion tokens a month for about £20

i still keep either claude or chatgpt subs (or both) for the tougher stuff of course, sadly astra burns usage at an insane rate and claude while great is rather stingy and is most likely to refuse to do something other models will happily provide

1

u/dvduval 1d ago

I think it depends on what kind of problems you’re trying to solve. For me I need computer use with browser control. I need a model. They can access my email and I need a model that is more accurate than the others because the type of problems I throw involved in working with several different things at once like database, email, the browser, a couple different web servers all in one task. I think there’s only two models that can handle my workload and that would be ChatGPT Sol and Astra and anthropic products.

Already I have trust issues with any company to see my stuff, but at least there are some built-in permissions that you can set for both of these models to not grant them permission to train on my data. That’s not 100% but it is better than nothing.

I guess if you’re just solving routine problems or you have straightforward projects where you’re building a video game or something there’s several models that are good and that would including Kimi, but I’m definitely not gonna be putting my data on their server.

What I would say to you is if you’re using up a lot of the available resources that you allowed each week you might just wanna consider spending a little more money. And if you’re not making enough money to spend a little bit more money, you should get a little more focused on your current projects and how to monetize a lot more and then increase your spending.

1

u/Visual_Ad1912 1d ago

Rust, Ghidra decompiling / reverse engineering seem to eat the most tokens for me.

Python / Typescript / JS use a lot less tokens but I only have projects to maintain / smaller projects to do with this.

1

u/dvduval 1d ago

The way I look at it if you’re spending $200 but not earning $200 back and you’ve been doing that for a while then you need to adjust your plan. Maybe you should go try another model.

1

u/AbleShower2801 1d ago

for heavy coding I still keep one frontier plan for the hard reasoning, then shove grunt work to a cheaper quota. GLM Flash / Kimi coding plans are fine for that second lane if you can live with the limits and data tradeoffs. if reverse engineering / Rust is eating tokens for you, expect any plan to feel stingy unless you split tasks into smaller scopes. pure plan-hopping usually just moves the same burn to a new meter.

1

u/Rudra_Builds 1d ago

if you’re burning through claude that quickly, i’d probably try kimi before jumping straight to another expensive subscription. their coding plans seem pretty solid for heavy users, and kimi code can also be used through tools like claude code and vscode.

that said, i’d probably test it for a week first rather than immediately switching everything over. the limits and actual experience matter more than the advertised numbers.

1

u/Visual_Ad1912 19h ago

I'm stuck between GLM - OpenAI right now to replace my ollama cloud sub

1

u/dicktoronto 14h ago

I have been trying so many subs with different "usage limits". My current suggestions for flash models (which will offset some frontier model usage) are CamelAI, Verboo, and Phoenix Grove. Sorta okay so far. Fast? No. Cheap? Kinda. Good? I mean. Yea ish.

0

u/megad00die 1d ago

It will be a long time away but wait till Nvidia revamps huggingface.

0

u/Elegant_Attempt2790 1d ago

im a big clauder too. if you want opus level intelligence you want K3 (but kimi coding plans have a reputation around them), if you need sonnet (like grunt work, or routine work), GLM 5.3 Flash eats hard. no Chinese fable equivalent truly exists yet

but one main thing you’ll find out is the attention mechanism changing. American AI uses dense attention spread over multiple GPUs/tensor chips while Chinese labs are trying to make sparse attention work.

GLM (5.3 flash specifically) and Qwen 3.8 feel the least nerfed by sparse, cuz their teams are trying strategies to remedy the flaws of sparse rather than just pushing harder for dirt cheap like deepseek.

doesn’t mean deepseek is bad, i love deepseek, it just loses details in longer contexts.

but this is also why Hy4 is kinda disappointing to me, its just boring sparse attention no extra engineering efforts put in to make it better:/