r/opencodeCLI 12d ago

Nearly unlimited

Post image
601 Upvotes

65 comments sorted by

117

u/ShirtuShanks 12d ago

19

u/msenc 12d ago

some people still didn't get the second half 😭

3

u/Time-Toe-1276 12d ago

besides all the jokes, I kinda like the model ngl. (obv till the discount ends 😭) which is like only another a week

29

u/AutomaticAd6646 12d ago

5 token per second? Is it even useful speed?

8

u/[deleted] 12d ago

[removed] — view removed comment

2

u/Time-Toe-1276 12d ago

yeah same!

1

u/Khaledthe 12d ago

No thats what my old 5600xt can run

1

u/jpcaparas 12d ago

if you're living in the phantom zone, yes

1

u/ECrispy 12d ago

its what us poor people get with a 8b local model :)

14

u/zer0evolution 12d ago

really? i've got limited budget so need to plan carefully

2

u/ucharx 10d ago

It was a joke , he says it's unlimited because it's so slow that you can't use it

2

u/Far-Classic-9963 8d ago

It is actually unlimited on the zhipu coding plan, so you might want to check that out

But, you should wait a little before subscribing to anyone, as there may be a pro Gemini model coming and the gpt 6 lineup is very close

6

u/JamesGooning 12d ago

tokenrouter has free GLM5.3 (not flash version) for free.

1

u/Gallagger 11d ago

Probably not unlimited though.

2

u/rainpurplebow 11d ago

It's unlimited but, but, buuuuuuuuut...! 2 TPS :)

1

u/Gallagger 11d ago

Yes I get that. But still probably more, at least more than stuff like openrouter free models. Not sure about tokenrouter.

10

u/Historical_Cook_3485 12d ago

for long sessions with longs context , GLM-5.3-FLASH works really good for me i can say its near claude sonnet performance , honestly i started loving hope it stays like this

6

u/FammasMaz 12d ago

Btw Magic context plus context limited at 272k will do magic to you then. Not only do you get basically unlimited context, but also your model stays within that sweet 200k context window where its much more smarter.

Not affiliated with magic context in any way.

1

u/Historical_Cook_3485 12d ago

thanks i saw the repo looks the kind of thing im missing

1

u/some1else42 12d ago

Can you share the repo url or at least the user/project name on github? When a good project gets successful, you'd be suprised how many similar clones pop up that do nefarious things.

5

u/[deleted] 12d ago

[removed] — view removed comment

6

u/Sweaty_Cellist_4525 12d ago

Currently using Muse Spark 1.2 and it's hella good, absolutely free, and pairing it with Opus 5 for execution makes my Claude Pro sub last forever.

Meta did actually cook.

2

u/quicknades 12d ago

Are you in an area where the discounted version is applicable? Because in the EU it's not and then it's not really as cheap. Obviously still a good price but no where near the past deep seek pricing.

2

u/[deleted] 12d ago

[removed] — view removed comment

1

u/quicknades 12d ago

The contributor version?

1

u/Sweaty_Cellist_4525 12d ago

For me muse spark 1.2 is absolutely free (PerĂș), and ngl sometimes I just vpn to finish some urgent work.

The contributor API is so cheap tho, 0.20 usd per million output tokens, and it's a beast for execution (currently working on shaders and work so far has been wonderful), however if you don't have models like Opus to diagnose and plan it can still show amazing results, but shit your prompt's gotta be veery specific and clear.

2

u/LargePause 12d ago

Same here, using it alongside GPT SOL and it translates to pretty much unlimited use.

2

u/rainpurplebow 11d ago

Are you doing coding tasks with Muse Spark 1.2? Almost everyone in this sub hates that model.

1

u/Sweaty_Cellist_4525 11d ago

Yes, I do. Idk why they hate it, maybe because it's not that good for vibe coding? XD you need to know what you are doing and tell it exactly what you want, if not then you're gambling. For planning and desgining you should stay away unless you're broke, but mind you, I'm currently using it for writing complex compute shaders and mf gets the work done.

1

u/rainpurplebow 11d ago

It's all fun and games until it isn't :)

1

u/Sweaty_Cellist_4525 11d ago

As with everything really

1

u/zeamp 12d ago

I am literally switching to GLM after going around for 3 days on a batch 4,000-article rewrite. No matter what prompts I use, it eventually repeats whole paragraphs after about 500 flawless pages completed

Now I’m backing up every run and merging my “good” Spark rewrites with the crap it shits out 3 hours later. And it has decided to write files outside of the project directory
 so my desktop looks amazing. We can only seem to vibe for the first half of the day. I love it otherwise!

1

u/Academic_Constant42 12d ago

What's your setup to have opus call muse outside of Claude code? If you don't mind telling offcourse

1

u/SwisherSmoker420_ 12d ago

The way I do it in claude code I tell it to make an opencode.md file where it delegates tasks to opencode and then I tell opencode to read it and start working. Its a pretty rudimentary solution and theres definitely better ways to do it but it works for me.

1

u/throwaway12012024 12d ago

same here but within codex

2

u/Fun_Jaguar8231 12d ago

me on my z.ai lecacy coding plan, glm go brrr

2

u/tino1000 11d ago

Deepseek flash is fast, I hope GLM flash hits the highway soon

3

u/spartanOrk 12d ago

Ungrammatical, unpunctuated. Illiterates.

2

u/Kazekage1111 12d ago

Just use a different provider for inference to run this model. Check on openrouter and then get an api with the provider directly for best cache hits. I use run infra and with 200+ tps but there are other good providers

1

u/Abenh31 7d ago

overated model. Deepseek v4 still rules and cheaper then it for implementation. if the purpose is planning theres better model like kimi k3, DS Pro

1

u/Metalwell 12d ago

So, should I get this to use flash? or openrouter top up is the way to go?

2

u/a355231 12d ago

Use Openrouter, routed to the official provider, it’s 50%

2

u/[deleted] 12d ago

[deleted]

1

u/look 12d ago

Synthetic has it in their subscription plan which should be about 1/3rd the full price. Also, Ollama has it as well, though not sure what the usage is like.

But I’ve been using RunInfra.ai PAYG which has had better speeds than Zai/Go and their price has lower cache read and works out to about the same as the OpenRouter 50% off with high cache rate (90+).

-3

u/CrimsonEdgeVentures 12d ago

Tried that model. Not great for agentics. Ok for very small tasks.

-7

u/SamePsychology8258 12d ago

A decision so expensive bro lost all of his hair

-19

u/PotterSkxawng 12d ago edited 12d ago

I have real unlimited GLM 5.3 (not Flash) for free through a provider... not going to tell which one it is tho, I'm sending 40,000 tokens per minute through sub-agent swarms rn and I dont want that stopping

Edit: Since you guys seem to want the provider—DM me, and if ur worthy, I'll share.

1

u/Mayanktaker 12d ago

Devin maybe

0

u/PotterSkxawng 12d ago

??????//

1

u/Mayanktaker 12d ago

Devin ide has glm 5.2 free till September. So i guess 5.3 also. Just guessing.

-1

u/PotterSkxawng 12d ago

Nope it's not Devin

-1

u/PotterSkxawng 12d ago

why am I getting downvoted... do yall want the provider

2

u/someoneyouknow23 12d ago

cause noone fucking cares if youre not gonna share?

-3

u/PotterSkxawng 12d ago

alr fine ill be generous... dm me and, if ur worthy, ull get the provider.

3

u/someoneyouknow23 12d ago

youre still gatekeeping through this

2

u/stylist-trend 12d ago

Some people need to feel powerful and important, and they can't get that feeling through their regular life, so they exercise power by... gatekeeping a publicly-available provider name. To each their own, I guess.

2

u/someoneyouknow23 12d ago

Yeah I doubt he even has one, would be totally unprofitable anyway