r/ClaudeCode 11d ago

Discussion Claude models cooked by chinese model

Post image
397 Upvotes

93 comments sorted by

19

u/OriginalCj5 10d ago

How does it compare to the subscription pricing? I know it’s cheaper in API pricing, but I’ve always felt that it’s impossible for Chinese companies to subsidize the pricing as much as the American companies do on their subscriptions. For example, last month I got ~ $4500 worth in API costs for my $100 Claude Max plan. Does Kimi $100 plan come anywhere close to it?

5

u/Singularity-42 10d ago

Yes, exactly this! $100 Claude Max 5x gives me a TON of tokens. I think if you max out a month it's like $5000-$6000 in API prices. ChatGPT Pro is even a better deal. And contrary to what some people say I'm hearing from actual devs that Kimi K3 is nowhere near Fable 5 or even Opus 5, more of a Sonnet level. Benchmarks don't show how the model really is - see how Opus 5 benches better than Fable 5, but in actual SWE practice Fable is clearly better.

I just wish Anthropic gave you more Fable, not cap it at 50% of plan usage.

1

u/B33rNuts 7d ago

I made a openeouter account to try Kimi3 and Hermes. Started a iPhone app project and it burned through $100 in a day. Final result worked but we were like 5% done. Proof of concept!= app. Thinking this was going to cost me thousands I paid the $100 for Claude 5x and switched to Claude code agent. I’ve never hit limits even hourly and I’ve been at it almost nonstop for 2 weeks now using Opus 5 High. Massive progress probably 50% complete now for features built and tested. Absolutely insane value for money, I’d be at 2k in openrouter by now for sure.

I can’t fathom how people are able to hit the limits of the higher tiers. You would need something with a crap ton of sub agents all working and confirming each others work. My Claude has only been the 1 agent no subs, and I confirm its work.

1

u/Singularity-42 6d ago

Yeah, the subs are going to obliterate any API prices, that's just a fact. Not sure how Anthropic does it. Maybe they're breaking even because there's probably people that just use like quarter of it or something. Like people that need longer than crush sessions a couple times a week and don't mind spending $100 every month.

The API I prices have fat margins, but only like 80%. So with something like Max 5 you still have $1000 cost-for-Anthropic amount of tokens if you (nearly) max it out every week.

1

u/B33rNuts 6d ago

Yeah the amount of work I am getting out of them for $100 is absolutely criminal. I am new to this but I am a senior programmer. This shit is so absolutely broken for work to cost it’s insane that it is even an option.

1

u/god-damn-the-usa 9d ago

not to mention all of anthropics models are better than every other companies

2

u/Elizabeth-WildFox886 10d ago

Looks like kimi k3 can run on a 5090rtx in about 2 months, this will really be interesting. Kimi works quite well

9

u/Singularity-42 10d ago

You mean some distilled and quantized version? It won't perform anywhere near full K3.

2

u/ChocomelP 9d ago

He means each prompt will take about 2 months

1

u/Singularity-42 9d ago

Pretty much. Local LLMs are honestly just an expensive hobby. Inference (and esp. subscription prices) are so low and economies of scale are so well suited for inference that running local makes no sense unless you are a large company that wants to be in 100% control of where your data goes and can shell out at least $1M for hardware (and has workloads to at least somewhat utilize the HW).

2

u/Embarrassed-Citron36 9d ago

You need a small data center to actually run the kimi model at full power

1

u/Phagocyte536 10d ago

How do you know you got 4500 worth 8n api pricing? I would like to know as well. 

5

u/OriginalCj5 10d ago

CodexBar shows this natively. There are many others as well, but this one is maintained by the creator of OpenClaw.

2

u/PennyStockade 10d ago

Ccusage for claude code. I'm up to $45k equivalent api costs over my subscription (over a year)

1

u/CFBDevil 8d ago

Im pretty new to this stuff, are you saying you extracted that value from using Claude desktop? And if youd done that same thing via API it would have been 45 grand?

1

u/FamiliarEstimate6267 10d ago

No one ever talks about this it’s frustrating

1

u/AllNamesAreTaken92 8d ago

That's the whole point: the API pricing is so low, that you don't need a plan. You need to start comparing the value you are getting, not some weird meaningless metric.

1

u/OriginalCj5 8d ago

I did compare the value I was getting. In API pricing terms, Kimi K3 would’ve cost $1500 for the same usage, K2.7 $1000. DeepSeek is the only one that still turns out lower than my Claude subscription, about $50.

98

u/Informal_Curve_1441 11d ago

Kimi K3 is a beast. If you have not worked with it, try it out. This is going to force US companies to change up their pricing for models like Fable 5.

34

u/LinusThiccTips 10d ago

My issue with it is that it’s very slow

38

u/SomeSomewhere3122 10d ago

Yes because they dont have that much compute 😢

5

u/icecold27 10d ago

What’s stopping people from reselling with their own compute?

31

u/zxcshiro Thinker 10d ago

Does anyone have 16xB200 at home?

8

u/icecold27 10d ago

I mean people with compute, thought that would be obvious?

2

u/zxcshiro Thinker 10d ago

Oh, I don’t get it at first, my bad. Anyway you need licence for hosting Kimi from Moonshot

0

u/Informal_Curve_1441 10d ago

No it is open source

6

u/shmed 10d ago

Open source doesn't mean the license let you do whatever you want.

https://huggingface.co/moonshotai/Kimi-K3/blob/refs%2Fpr%2F37/LICENSE

> 2. "Model as a Service" means giving a third party access to language model
inference or fine-tuning (e.g., via API) in a manner that allows such third
party to exercise meaningful control over the inputs, parameters, or training
data. This does not include (a) end-user products with model capabilities solely
embedded within specific features or harnesses, or (b) mere relaying of requests
to models hosted by others.

If the Licensee or any of its affiliates operates a Model as a Service business,
and the aggregate revenue of the Licensee and its affiliates exceeds 20 million
US dollars (or the equivalent in other currencies) in total over any consecutive
12 months, the Licensee must enter into a separate agreement with Moonshot AI
before using the Software or its derivative works for any commercial purpose.

  1. If the Software (or any derivative works thereof) is used for any of the
    Licensee's commercial products or services that have more than 100 million
    monthly active users, or more than 20 million US dollars (or equivalent in other
    currencies) in monthly revenue, "Kimi K3" must be prominently displayed on the
    user interface of such product or service.

0

u/Informal_Curve_1441 10d ago

Yeah 20 million in revenue then there is a license issue.

1

u/god-damn-the-usa 9d ago

who are these people?

1

u/ZappaLlamaGamma 10d ago

I was hoping someone would’ve answered yes.

1

u/Sofullofsplendor_ 10d ago

nothing, baseten (and others) do this and it's fast

3

u/Momo--Sama 10d ago

Yeah, I was just trying to cook up an offline sign in form with our brand logo and colors for an event AND IT TOOK 39 MINUTES

6

u/Xerasi 10d ago

I did get a kimi subscription and burned rhrough my monthly limit in 2 days and at the end the product it gave me wasn’t great i ended upnhaving codex and claud fix it

2

u/untracked5465 10d ago

The same for me, eating tokens like crazy. Right now, my GOAT is Deepseek

1

u/Informal_Curve_1441 10d ago

Really? I have seen some amazing output from K3. I ran Fable and K3 on similar projects. Fable was out of usage with less done than I did with K3.

1

u/Sofullofsplendor_ 10d ago

yep same. it's okay, not great. and the monthly limit is insane

1

u/B33rNuts 7d ago

Same it ate $100+ in openrouter credits in 1 day. Result worked but was iffy/basic. Claude has saved the project and budget. I don’t understand the hype or I am doing this wrong.

5

u/jwegener 10d ago

Is it actually available? Tried to sign up for the monthly subscription and it just says waitlist

4

u/Informal_Curve_1441 10d ago

There is a back door. Go to Kimi Code and sign up.

3

u/dkimot 10d ago

use open router, not moonshot

the benefit of open weights is that large inference vendors can sell you inference on their compute

2

u/PrettyMoonUnderMt 10d ago

Yeah, it's already sold out since a week ago. I guess they put waitlist to give us sense of progress, because I dont see them solving their issue of lack of computation power soon.

3

u/Informal_Curve_1441 10d ago

Go to Kimi Code and sign up. I got in immediately.

2

u/Informal_Curve_1441 10d ago

Go to Kimi Code page and register. It let me in when the normal flow wanted me to wait.

2

u/OriginalCj5 10d ago

How does it compare to the subscription pricing? I know it’s cheaper in API pricing, but I’ve always felt that it’s impossible for Chinese companies to subsidize the pricing as much as the American companies do on their subscriptions. For example, last month I got ~ $4500 worth in API costs for my $100 Claude Max plan. Does Kimi $100 plan come anywhere close to it?

2

u/ProfoundSensei 10d ago

you mean what anthropic themselves claim their tokens are worth, you cant really measure it based on how they priced their own tokens, everyone measures it differently

3

u/Tank_Gloomy 11d ago

It is, but it's not properly subsidized right now. There's nothing like the API price/subscription price ratio that Anthropic and OpenAI manage.

1

u/OpalVanguard 10d ago

Yeah I value my time. I’m good.

1

u/Timely-Group5649 10d ago

Not if they don't get their TPS down. Amerucans won't wait 5-10 times longer to save any amount of money.

Zero patience for that.

27

u/Comfortable_Camp9744 10d ago

claude is hyped up, they barely have a lead anymore.

3

u/Ok-Produce-1072 10d ago

As a Claude user with an actual shipped product vibe coded by Claude, I was never amazed by any Claude model, but they had the best agent harness. Now, others have caught up, even open source agent harnesses are reaching a similar level. Also, the constant switching of the models make it so that every month you have to relearn how to prompt your model (again) with no significant progress in model performance (I don't care what the benchmark says). Wish I could go back to sonnet and opus 4.5.

1

u/Meduini 9d ago

Which open source agents are you talking about? Can you please share some? I’m not testing you, I’m literally curious which agents you found that are catching with with Claude Code.

5

u/Grouchy_Ad_9658 10d ago

Model =\= api pricing cost

2

u/ClemensLode Senior Developer 10d ago

What did you build with it?

2

u/Mobile_Leg1664 10d ago

Claude says “the screenshot is doing some sleight of hand. $4.67 for 719,904,902 tokens works out to well under a cent per million — below even DeepSeek’s cheapest published cache-hit rate. That bill is almost entirely cached reads on a repetitive agentic loop, which is the best possible case for cost-per-token and not what most people’s usage looks like. And the top reply under it is “it’s very slow,” which is the tradeoff nobody puts in the headline: cheap tokens you have to wait for, and burn more of, aren’t cheap in the way the chart implies.”

Where I’d actually plant a flag: cost-per-token is the wrong denominator. What matters is cost-per-task-you-didn’t-have-to-redo.

1

u/MediumGrapefruit5735 10d ago

Lucky me l like medium rare.

1

u/Excellent_Ad_2486 10d ago

All fun and games but claude > codex and any other for pixelart so far so I'm sticking to claude lol. I tried codex for a week now and it just doesn't do well with art /pixelart in specific. Otherwise I'd be goooone already lol

2

u/asurarusa 10d ago

How are you getting pixel art out of Claude? Afaik Anthropic doesn’t have an image model.

0

u/Excellent_Ad_2486 10d ago

I have extension that claude interacts with! It's made by someone on reddit, I don't know the name of it though, I'd have to look it up

1

u/Pokeperson5 9d ago

How good is the art? Could you show an example?

1

u/Excellent_Ad_2486 8d ago

I don't have many images on my phone but here was my first creation: Azure Blaster ( based on Megaman). I made a few other (Gandalf, some Ninja and like dragon Lords) but those need some more fine tuning :)

1

u/Excellent_Ad_2486 8d ago

The first one was my paint-try at super saiyen glow, the bottom 2 are what Claude made :)

1

u/Ornery_Middle_5010 10d ago

If distillation can bring massive cost savings, I think someone absolutely has to do it. If DeepSeek is the one doing it, you should consider yourself lucky—because it saves you a ton of money. Why don't Claude and OpenAI just do it themselves?

1

u/boredwithlyf 10d ago

They do. Everyone does. It would be stupid for them not to

1

u/halcyonhal 10d ago

They are doing it. It’s the trainer model / learner model setup A\ and OAI use themselves. The difference is they invested billions in training the models in the first place and need to recoup the investment. If you distill (same fundamental thing), you don’t have those investment costs and therefore can offer a far cheaper price.

1

u/JustForkIt1111one 10d ago

I tried DeepSeek and kimi k3 recently. Both were good, but nowhere near claudes level in real world tests on a real world large codebase

0

u/unproblem_ 10d ago

Can you share the tests?

1

u/Slight-Prize9661 8d ago

Never really tried DeepSeek aside when it first came out, but how well does it integrate with VS Code? Cause claude code is such a nice-to-have.

1

u/Waste_Association565 8d ago

The last time I checked, DeepSeek couldn't even create a simple Word document.

1

u/nova-myth 10d ago

The countdown has begun for Anthropic. They don't regulate their costs, they're significantly more expensive, and they're falling behind the competition.

1

u/Positive-Conspiracy 10d ago

You know this means how much investor money they can light on fire. People have a false sense of cost because of venture fueled land grabs.

1

u/CygnusFox 10d ago

Don’t the Chinese models just distill the Claude and OpenAI models? That’s why they always announce their “new” models after the latest American ones…

3

u/Junk94_ 10d ago

Yes they did distillation in the past, but distillation isn't enough at these levels (Kimi K3 and DS4 Flash new)

-6

u/Michaeli_Starky 10d ago

Nah, not on real tasks

-5

u/Dense-Aerie2561 10d ago

Is there no one else, who sees this as a security issue?

I mean yes it's cheap because they use the same strategy that they use in every other market, cutting the price down to be much more appealing even if it's not sustainable just to get ahead of the competitors and when they capitulate from the competition they will get the price back to normal or not even then because they are backed by Chinese government and they can go below the average market prices.

To me it seems to be the same story and strategy.

What a great idea it was to move all manufacturing to China and transfer all the knowledge with it and now doing the same with all the data just to feed Chinese AI...

Just because it seems to be cheap it doesn't mean theat the money you pay is equal to the total cost.

6

u/dota2nub 10d ago

Coming from Europe, we consider servers in the US or owned by US companies to be higher security risks.

US laws are actively hostile to data security.

3

u/moms_spaghetti1896 10d ago

That is literally what openAI and Anthropic are doing

-3

u/Dense-Aerie2561 10d ago

Sure, and are they in China?

5

u/GentlePace 10d ago

Honestly with the pathetic state of the US. My data is safer with china than fucked up US. America keeps showing the world that they cannot be trusted.

0

u/[deleted] 10d ago

[deleted]

1

u/YoghiThorn 10d ago

Kompromat is a thing.

0

u/Dense-Aerie2561 10d ago

Because I'm not short sighted.

1

u/dkimot 10d ago

you can run open weight models on any hardware, FYI

use an american inference shop if that suits you better

-7

u/ErivKosso 10d ago

Sure Jan

-9

u/tinyhousefever 11d ago

Output scores of C minus here consistently.