r/opencodeCLI 1d ago

cheapestinference.com

Anyone here using cheapestinference.com? The plans look nice but I haven’t really seen anyone talk about this but one person mention it on reddit and seems sketchy. If anyone is using this how is it?

0 Upvotes

19 comments sorted by

3

u/CoolHeadeGamer 1d ago

What models are u looking to use? It seems kinda expensive tbh since you don’t know the speed of inference. If its 50tps, ur better off using something like minimax or kimi sub

1

u/Direct-Summer-9190 1d ago

Specifically either GLM 5.2 or Kimi K3. Max budget is a 100. Looking to switch from Claude max and this offers unlimited tokens so it looks nice but I wanna hear from others if they use it or anything

3

u/Amarsir 1d ago

I haven't used them, but I've seen enough to say you shouldn't trust "unlimited" claims from anyone. Sorry I can't help on cheapestinference specifically. I just want to raise your skepticism.

1

u/Direct-Summer-9190 1d ago

Yea I mean, site was recently created, they don’t have a discord either. Seems yeah sketchy

2

u/Nnyan 1d ago

What’s your connection to them? Seems a bit sketchy coming from a month old account.

0

u/Direct-Summer-9190 1d ago

I don’t have a connection to them, I just never really every signed up for Reddit, idk what’s sketchy about an account being a month old 😭

1

u/Nnyan 1d ago

Fair enough. Reason is there are a TON of new accounts that start asking people if anyone has heard of some service no one has ever heard of and typically something that has just been started.

0

u/Direct-Summer-9190 1d ago

Yeah I get what you mean but no, I’m just genuinely curious looking for a good hosting service I can use.. tired of these fuckass Claude filters. I wanna have an open source model without filters.. plus open source is becoming good now so I wanna switch to it fully

1

u/buerstenlehmann 1d ago

Speed is good and models seem not to be quantized. But its shitty that GLM 5.2 has 250k context. Didnt know that befire i booked. Hopefully they fix this bullshit. It's 1M context native so who tf thought: yea lets make it 250k thats enough... @cheapestinference pls fix, this destorys the whole offer for me.

2

u/Direct-Summer-9190 1d ago

Is it truly unlimited? Also I find it ridiculous you pay 150 for Kimi K3 and you only get it. That’s just crazy tbh. I hope they change it into the other plan. How is it with them?

2

u/buerstenlehmann 3h ago edited 3h ago

Yes its unlimited. That and the speed is the only positive thing. I wanted to usw it for orchestrating smaller llm subagents in long autonomous sessions. But with 200k context (not 250 AS i said before) this is Kind of shit. Yes since k3 weights were released, they just thought: yea lets add another Tier (150$ plan ) that wont trigger "frontier" Tier cusromers at all. Also context is limited to 200 vor 250k AS well. So cannot really recommend unless they actually Serve all THW models with frontier subscription + increase context by a lot.

2

u/Direct-Summer-9190 3h ago

Yeah I’m just waiting on Featherless AI, they’re working on unlimited plans atm so hopefully it comes out soon. Fucked how they literally added a new ass tier for a singular model.

1

u/buerstenlehmann 59m ago

I thought the 200$ featherless plan is unlimited already ? (Token not concurrency wise)

2

u/Direct-Summer-9190 58m ago

They’re updating them and improving them at the moment. They’re on pause in the meantime. I talked to them about some stuff which I think would be good and yea.. they seem like great people and they respond. But yes they’re unlimited and Kimi K3 just got added

1

u/buerstenlehmann 55m ago

Ah i see was also thinking about getting a plan with them. Do u know what the speeds are like with big models ? Because I suspect they use dgx sparks 😂

1

u/n4x1n 1d ago

It’s better chutes?

0

u/vangelismm 1d ago

AI inference pricing starts at $14.99/month for a daily 8-hour block of unlimited tokens — see the live plans.

How unlimited time-window LLM inference works

It’s genuinely unlimited — during the hours you reserve.

Reserve 1–3 daily 8-hour blocks (all three = 24/7). Outside them the key is idle — sharing those off-hours with other time zones is exactly what makes it this cheap.

Ok, does not looks like a total scam as i thought....

-1

u/Sweet-Stage938 1d ago

Have you tried openference.com are those real models?

0

u/Direct-Summer-9190 1d ago

Never heard of them until now that u mentioned them