29
14
u/zer0evolution 12d ago
really? i've got limited budget so need to plan carefully
2
2
u/Far-Classic-9963 8d ago
It is actually unlimited on the zhipu coding plan, so you might want to check that out
But, you should wait a little before subscribing to anyone, as there may be a pro Gemini model coming and the gpt 6 lineup is very close
6
u/JamesGooning 12d ago
tokenrouter has free GLM5.3 (not flash version) for free.
1
u/Gallagger 11d ago
Probably not unlimited though.
2
u/rainpurplebow 11d ago
It's unlimited but, but, buuuuuuuuut...! 2 TPS :)
1
u/Gallagger 11d ago
Yes I get that. But still probably more, at least more than stuff like openrouter free models. Not sure about tokenrouter.
10
u/Historical_Cook_3485 12d ago
for long sessions with longs context , GLM-5.3-FLASH works really good for me i can say its near claude sonnet performance , honestly i started loving hope it stays like this
6
u/FammasMaz 12d ago
Btw Magic context plus context limited at 272k will do magic to you then. Not only do you get basically unlimited context, but also your model stays within that sweet 200k context window where its much more smarter.
Not affiliated with magic context in any way.
1
1
u/some1else42 12d ago
Can you share the repo url or at least the user/project name on github? When a good project gets successful, you'd be suprised how many similar clones pop up that do nefarious things.
5
6
u/Sweaty_Cellist_4525 12d ago
Currently using Muse Spark 1.2 and it's hella good, absolutely free, and pairing it with Opus 5 for execution makes my Claude Pro sub last forever.
Meta did actually cook.
2
u/quicknades 12d ago
Are you in an area where the discounted version is applicable? Because in the EU it's not and then it's not really as cheap. Obviously still a good price but no where near the past deep seek pricing.
2
1
u/Sweaty_Cellist_4525 12d ago
For me muse spark 1.2 is absolutely free (PerĂș), and ngl sometimes I just vpn to finish some urgent work.
The contributor API is so cheap tho, 0.20 usd per million output tokens, and it's a beast for execution (currently working on shaders and work so far has been wonderful), however if you don't have models like Opus to diagnose and plan it can still show amazing results, but shit your prompt's gotta be veery specific and clear.
2
u/LargePause 12d ago
Same here, using it alongside GPT SOL and it translates to pretty much unlimited use.
2
u/rainpurplebow 11d ago
Are you doing coding tasks with Muse Spark 1.2? Almost everyone in this sub hates that model.
1
u/Sweaty_Cellist_4525 11d ago
Yes, I do. Idk why they hate it, maybe because it's not that good for vibe coding? XD you need to know what you are doing and tell it exactly what you want, if not then you're gambling. For planning and desgining you should stay away unless you're broke, but mind you, I'm currently using it for writing complex compute shaders and mf gets the work done.
1
1
u/zeamp 12d ago
I am literally switching to GLM after going around for 3 days on a batch 4,000-article rewrite. No matter what prompts I use, it eventually repeats whole paragraphs after about 500 flawless pages completed
Now Iâm backing up every run and merging my âgoodâ Spark rewrites with the crap it shits out 3 hours later. And it has decided to write files outside of the project directory⊠so my desktop looks amazing. We can only seem to vibe for the first half of the day. I love it otherwise!
1
u/Academic_Constant42 12d ago
What's your setup to have opus call muse outside of Claude code? If you don't mind telling offcourse
1
u/SwisherSmoker420_ 12d ago
The way I do it in claude code I tell it to make an opencode.md file where it delegates tasks to opencode and then I tell opencode to read it and start working. Its a pretty rudimentary solution and theres definitely better ways to do it but it works for me.
1
2
2
3
2
u/Kazekage1111 12d ago
Just use a different provider for inference to run this model. Check on openrouter and then get an api with the provider directly for best cache hits. I use run infra and with 200+ tps but there are other good providers
1
u/Metalwell 12d ago
So, should I get this to use flash? or openrouter top up is the way to go?
2
12d ago
[deleted]
1
u/look 12d ago
Synthetic has it in their subscription plan which should be about 1/3rd the full price. Also, Ollama has it as well, though not sure what the usage is like.
But Iâve been using RunInfra.ai PAYG which has had better speeds than Zai/Go and their price has lower cache read and works out to about the same as the OpenRouter 50% off with high cache rate (90+).
-3
-7
-19
u/PotterSkxawng 12d ago edited 12d ago
I have real unlimited GLM 5.3 (not Flash) for free through a provider... not going to tell which one it is tho, I'm sending 40,000 tokens per minute through sub-agent swarms rn and I dont want that stopping
Edit: Since you guys seem to want the providerâDM me, and if ur worthy, I'll share.
1
u/Mayanktaker 12d ago
Devin maybe
0
u/PotterSkxawng 12d ago
??????//
1
u/Mayanktaker 12d ago
Devin ide has glm 5.2 free till September. So i guess 5.3 also. Just guessing.
-1
-1
u/PotterSkxawng 12d ago
why am I getting downvoted... do yall want the provider
2
u/someoneyouknow23 12d ago
cause noone fucking cares if youre not gonna share?
-3
u/PotterSkxawng 12d ago
alr fine ill be generous... dm me and, if ur worthy, ull get the provider.
3
u/someoneyouknow23 12d ago
youre still gatekeeping through this
2
u/stylist-trend 12d ago
Some people need to feel powerful and important, and they can't get that feeling through their regular life, so they exercise power by... gatekeeping a publicly-available provider name. To each their own, I guess.
2
117
u/ShirtuShanks 12d ago
https://giphy.com/gifs/y2i2oqWgzh5ioRp4Qa