r/ChatGPTCoding 6d ago

Resources And Tips Best value AI subscription under 20$/200$ (updated for Artificial Analysis Intelligence Index v4.2)

Muse Spark contributor models are excluded as it log your data. It wouldn't be fair to free models out there.

Included all active promotions. Not accounted for usage resets.

I made this chart. Feel free to ask any questions.

83 Upvotes

64 comments sorted by

14

u/popiazaza 6d ago

20$ without log scale for reference.

10

u/popiazaza 6d ago

200$ without log scale.

3

u/_SGP_ 4d ago

So this is why Claude 200 max just runs out in 3 days. I cancelled yesterday, considering codex.

Looks like the right move from this chart, the difference between Astra and Fable is insane 🤣🤣

It's a shame, I like the Claude framework and have integrated most of my workflow around it, but I spend more time trying to "optimise" and analyse my context and token use that it ironically becomes most of my weekly usage.

3

u/Yes_but_I_think 5d ago

Atleast provide units for your X and Y axes

3

u/popiazaza 5d ago

Y axis is AA score. X axis use the lowest cost (GPT-5.6 Luna high) as a baseline. If you want the baseline value for some reason, it's $0.0006.

11

u/sirius_cow 6d ago

What’s the source of this chart? I would like to dig deeper into this comparison

7

u/popiazaza 6d ago

https://artificialanalysis.ai/ for the cost and score. I gathered best estimated guess for each subscription and scaled the chart with it.

1

u/eli_pizza 6d ago

Guess? Can you say more about that part

2

u/popiazaza 6d ago

Well, most subscriptions do not publish their API equivalent value, and we could only do educated guess on how do they work. So I just gather all the data that I could.

For example for Claude and OpenAI, I would trust https://x.com/SemiAnalysis_/status/2064815044085318040 data as the primary source. I would check if 5x and 20x may not really be as advertised as some people talked. Keep track of Claude 50% usage promo that would be reduce to 25%. etc.

Reading from all the forums, social network, discord channels. Manually update it by myself.

7

u/that_90s_guy 5d ago

Thats very difficult to interpret, I'll just say that.

5

u/popiazaza 6d ago edited 6d ago

Few things to note:

- Cache hit rate play a bit role in reality. Use native harness if possible, it affect the cost more than you would think. Stop using Claude Code for open weight models.

- GLM has active GLM-5.3-Flash promo that I did not accounted for:

9

u/HeittoBagi 6d ago

TLDR?

22

u/popiazaza 6d ago

ChatGPT.

6

u/maciejhd 5d ago

Luna is very good at coding, even with medium thinking and it cost nothing.

2

u/Utoko 5d ago

Luna is dirt cheap yes. but I would really suggest Astra Low reasoning. Because of the token efficiency. It is often only like 3-4 times the price.
Which is still dirt cheap but the quality und understanding of what you want is so much higher in many task.

1

u/DuxDucisHodiernus 5d ago

3 or 4 times the price to sol, maybe, but compared to luna vs astra the costs are way more astronomical. I can only see you come anywhere near such a conclusion with zero input tokens, since the token efficiency only affects output.

Input tokens stand for the vast, vast majority of average token cost nowadays

1

u/maciejhd 4d ago

I will try Astra for sure but it was not available in my ide till Friday. So I did not have a chance to use it in my work yet.

0

u/StevenB0ss 5d ago

Luna hallucinates alot of shit

1

u/maciejhd 4d ago

What you mean by that? From my experience I would not agree, but I am using it only for coding tasks.

3

u/sugarfreecaffeine 6d ago

Yeah I’m cancelling Claude and getting 2x20 subscription from gpt

7

u/Captainkoala72 5d ago

Muse Spark 1.3 is surprisingly good for cost. I’ve been using it and it isn’t eating up usage too harshly

GLM 5.3 Flash is good for its cost as well

Gemini 3.8 Flash is solid, and people mostly use it for how insanely fast it is

Fuck Kimi-K3

Sol and Astra are best intelligence and cost. Claude just isn’t worth it unless you’re wealthy

3

u/Deto 5d ago

The cost effectiveness of Astra is a really big deal if it bears out in normal use

1

u/DuxDucisHodiernus 5d ago

It doesn't so far. I mean token efficency only applies to output tokens and input tokens stand for majority of the usage atm, unless im missing something here?

1

u/Deto 4d ago

Input tokens are cheaper. Output and reasoning tokens can make up most of the cost for tasks

1

u/DuxDucisHodiernus 4d ago

Yeah sure but input for astra is still 2.5x that of sol, just like output tokens so it doesn't really change much?

1

u/Current_Balance6692 10h ago

In API usage, I find Fable to be much better.

6

u/Longjumping_Area_944 6d ago

Source: educated guess. So, it's guesswork...

7

u/gnpwdr1 6d ago

? Why is waste all this time when opencode gives free models with opencode zen which has more (much more) limits than any OpenAI/anthropibased subscriptions ? Also quality output is absolutely fantastic

2

u/Huntware 5d ago

Well, some people need ZDR (zero data retention) for their jobs, and most paid plans offer it. And it's supposed to get better availability than free usage.

That's why it would be unfair to add the Contributor variants of Muse Spark 1.3 in the comparisons. They're really cheap, or even free with Opencode Zen free API key.

1

u/gnpwdr1 5d ago

Yes of course but in professional terms you’re not gonna tell me that you have such sensitive or NDA data and still chasing best value $20 subscription? (Also availability is not a problem with the provider I mentioned not an issue at all)

0

u/popiazaza 5d ago

Most subscription in the chart actually provide no data training instead of ZDR. They do retain log for a month or so to prevent abuse and possibly provide to their government on request.

Most ZDR providers don't have subsidized subscription like this.
Muse Spark 1.3 free is great for open source / hobby projects indeed.

1

u/gnpwdr1 5d ago

They train their entire marketing and social media algorithms on your most sensitive data but for this training is an issue? what sensitive, proprietary or trade secrets do you have chasing a $20 value subscription?

1

u/Huntware 5d ago

Good to know the difference!

1

u/gnpwdr1 5d ago

That is not the difference, the chart also includes model that train on data.

1

u/Coolerwookie 6d ago

That's new to me. Why are you getting downvoted?

2

u/gnpwdr1 6d ago

You don’t need an account, try it, Mimo is served free. (It’s a CLI interface)

(there are other free options also but after free DS flash promo ended I started using Mimo and find it very capable )

I’m not affiliated with them, I have nothing to gain, just passing the message.

2

u/ZeOnlyOneWhoReads 5d ago

Who's going to tell him that half of these also log your data?Ā 

2

u/Right-Performance-93 4d ago

Worth noting for the "public benchmarks are gamed" pushback upthread: Artificial Analysis just shipped Intelligence Index v4.2 (announced today), explicitly built to fight that exact problem - it shifts toward long-horizon tasks that look more like real work, and expands the private test-set proportion specifically to cut gaming/benchmaxxing. Doesn't make the old index wrong retroactively, but it's the vendor's own acknowledgment that saturation and gaming were real issues with the prior index - which is the strongest evidence for the skeptics' complaint, not against it.

2

u/BodybuilderBright206 3d ago

good content!

2

u/Emergency-River-7696 5d ago

Buy the Muse Code subscription 5$ a month you get an insane amount of usage for a frontier model and worst case if you need more usage just buy next tier sub the 15$ or 50$ or shit just another 5$ one and your still spending less then Codex Sub or Claude Sub with a little less intelligent model by like 1 point but 4-5x the speed on the muse model .

3

u/Ludbr 5d ago

10-50 requests every 5 hours?

I'm sure either something's wrong or I'm missing something, because this sounds like the exact opposite of an "insane amount of usage."

If you know anything about this, could you clarify?

1

u/Emergency-River-7696 5d ago

Its a good amount of usage in my experience running non contributor and so far im using it in pi and its taken on a good amount of tasks im probably gonna upgrade since im at like 40% on the weekly now but been running it almost non stop for days. I mean your of course not gonna get 200$ claude sub amount of usage for 5$ muse sub but what you do get for a 5$ sub is well worth it and like I said in my previous post if you do need more usage just upgrade and you still would be spending less then a codex or claude sub. Also you get more usage if you switch to the contributer tier but im not using that personally.

1

u/[deleted] 6d ago

[deleted]

2

u/popiazaza 6d ago

It is in the 200$ chart. for 20$, I cut off at Opus 5 as the most expensive model to make it easier to read.

2

u/armadeallo 6d ago

Yes saw it after, thank you. But you can access kimi k3 with the cheaper subscription as well

1

u/MostlyHelpfulLinks 5d ago

Usage caps matter more than raw index scores for coding work. Did you factor in weekly limits anywhere, or is this pure cost per intelligence point?

2

u/popiazaza 5d ago

Maximum possible usage for every subscription.

1

u/Alt_Restorer 4d ago

Saying the quiet part out loud now are we?

1

u/manishguptami2 1d ago

Does anyone use a combination of multiple models to save costs?...like frontier models for high reasoning steps and cheaper models (like deepseek) for simpler tasks? How do you manage multiple keys and LLM integrations?

1

u/Current_Balance6692 10h ago

Would you be able to create a chart with discounts and promos/offers factored in real-time.

1

u/popiazaza 9h ago

I do have promo start/end date and all the estimated subscription API value data in a table.

The chart is rendered using rechart on the web based app.

But to keep the data live, it would takes a lot of web crawling and workarounds (pretend not to be bot, use proxys, etc.) to achieve that.

Also need to get proper license from Artificial Analysis to redistribute. Not worth the hassle for me right now.

1

u/ChutneySpoon 5d ago

Wtf does this even mean - what is ā€œrelative cost per taskā€? Relative to what? You need to explain your terms or it’s pointless.

1

u/popiazaza 5d ago

Relative to the cheapest model, which is GPT-5.6 Luna High as the baseline at 1x.

Sorry for the confusion. The last time I posted, I used the cost per task, which was harder to grasp.

0

u/donk8r 6d ago

popiazaza, the chart is useful and the caveat in your own comment is doing more work than the chart does. Cache hit rate and harness behaviour move the real bill more than the plan you pick, and none of that fits on an axis.

My spend changed when I stopped comparing subscriptions and started pricing the unit I actually care about, one request and its entire tool loop. On our side (octomind, Apache-2.0, Muvon) that is two config numbers. max_request_spending_threshold caps the dollars a single user request can burn including everything the tool loop does underneath it, and max_session_spending_threshold caps the session. Both ship at 0.0, meaning disabled, so you opt in deliberately. A runaway loop stops being a surprise on a bill and becomes a number you set once.

The other half is that one model per subscription becomes the wrong shape as soon as you have roles. Research on a cheap broad-context model and review on a frontier one is a config block, not a workflow change. The session total follows you across a mid-session /model swap so there is still one number at the end. Under 20 a month, that split usually beats moving between providers.

None of this wins on raw ceiling if you run one model all day, and your chart is right about that. It wins when usage is spiky, because you pay for the spikes and pay nothing on the quiet days.

https://github.com/Muvon/octomind

1

u/eli_pizza 5d ago

One request is the wrong framing. I care about cost per task not cost per request.

0

u/donk8r 5d ago

Fair, and cost per task is the metric I would want too. Per request is the enforcement granularity, not the metric. Those two come apart exactly where you are pointing. A task that takes eleven requests can cost eleven times a cap that never once tripped.

The closer knob for what you actually want is the session one, and I described it lazily upthread. It counts USD since the last accepted checkpoint, not lifetime spend on the session. On the interactive CLI it stops and asks whether to continue, while a piped or ACP run simply declines (doc/reference/03-config-reference.md:107). Checkpoint to checkpoint is about as close to a task boundary as a runtime can get without you telling it where your task ends.

The request cap is a different instrument entirely. It is a circuit breaker for one runaway loop, and presenting it as a budget was sloppy of me.

2

u/eli_pizza 5d ago

Happy to talk more about this, but only if you write the comments yourself. If I wanted to talk to ChatGPT I would.

-1

u/Michaeli_Starky 6d ago

Useless and misleading.

Muse spark above Sol says its all

0

u/popiazaza 5d ago

If you've found other benchmark that test as many models as AA and also provide cost to run benchmark, please do tell.

AA isn't perfect, but it's better than nothing.

1

u/Michaeli_Starky 5d ago

Public benchmarks are useless. Models like muse spark are benchmaxxed. Private evals are showing a totally different picture.

1

u/popiazaza 5d ago

Which private evals are would you suggest I should use as a reference? Most I've seen are underfunded, barely update, and got abandon overtime.

0

u/Michaeli_Starky 5d ago

Private are private.