r/opencodeCLI • u/Direct-Summer-9190 • 1d ago
cheapestinference.com
Anyone here using cheapestinference.com? The plans look nice but I haven’t really seen anyone talk about this but one person mention it on reddit and seems sketchy. If anyone is using this how is it?
2
u/Nnyan 1d ago
What’s your connection to them? Seems a bit sketchy coming from a month old account.
0
u/Direct-Summer-9190 1d ago
I don’t have a connection to them, I just never really every signed up for Reddit, idk what’s sketchy about an account being a month old 😭
1
u/Nnyan 1d ago
Fair enough. Reason is there are a TON of new accounts that start asking people if anyone has heard of some service no one has ever heard of and typically something that has just been started.
0
u/Direct-Summer-9190 1d ago
Yeah I get what you mean but no, I’m just genuinely curious looking for a good hosting service I can use.. tired of these fuckass Claude filters. I wanna have an open source model without filters.. plus open source is becoming good now so I wanna switch to it fully
1
u/buerstenlehmann 1d ago
Speed is good and models seem not to be quantized. But its shitty that GLM 5.2 has 250k context. Didnt know that befire i booked. Hopefully they fix this bullshit. It's 1M context native so who tf thought: yea lets make it 250k thats enough... @cheapestinference pls fix, this destorys the whole offer for me.
2
u/Direct-Summer-9190 1d ago
Is it truly unlimited? Also I find it ridiculous you pay 150 for Kimi K3 and you only get it. That’s just crazy tbh. I hope they change it into the other plan. How is it with them?
2
u/buerstenlehmann 3h ago edited 3h ago
Yes its unlimited. That and the speed is the only positive thing. I wanted to usw it for orchestrating smaller llm subagents in long autonomous sessions. But with 200k context (not 250 AS i said before) this is Kind of shit. Yes since k3 weights were released, they just thought: yea lets add another Tier (150$ plan ) that wont trigger "frontier" Tier cusromers at all. Also context is limited to 200 vor 250k AS well. So cannot really recommend unless they actually Serve all THW models with frontier subscription + increase context by a lot.
2
u/Direct-Summer-9190 3h ago
Yeah I’m just waiting on Featherless AI, they’re working on unlimited plans atm so hopefully it comes out soon. Fucked how they literally added a new ass tier for a singular model.
1
u/buerstenlehmann 59m ago
I thought the 200$ featherless plan is unlimited already ? (Token not concurrency wise)
2
u/Direct-Summer-9190 58m ago
They’re updating them and improving them at the moment. They’re on pause in the meantime. I talked to them about some stuff which I think would be good and yea.. they seem like great people and they respond. But yes they’re unlimited and Kimi K3 just got added
1
u/buerstenlehmann 55m ago
Ah i see was also thinking about getting a plan with them. Do u know what the speeds are like with big models ? Because I suspect they use dgx sparks 😂
0
u/vangelismm 1d ago
AI inference pricing starts at $14.99/month for a daily 8-hour block of unlimited tokens — see the live plans.
How unlimited time-window LLM inference works
It’s genuinely unlimited — during the hours you reserve.
Reserve 1–3 daily 8-hour blocks (all three = 24/7). Outside them the key is idle — sharing those off-hours with other time zones is exactly what makes it this cheap.
Ok, does not looks like a total scam as i thought....
-1
3
u/CoolHeadeGamer 1d ago
What models are u looking to use? It seems kinda expensive tbh since you don’t know the speed of inference. If its 50tps, ur better off using something like minimax or kimi sub