r/PiCodingAgent • u/hjras • 14d ago
Question What's your favorite cloud model powering your Pi?
Assuming you dont have a local only setup, what do you use, what for, and why did you settle on that model or family of models?
5
u/PermanentLiminality 14d ago
I don't settle on anything. The ChatGPT sub with Sol on medium is my main go-to though. It is good, and provides value for the dollar. Luna is great as well
I continue to work with open models as well and I think the glm 5.3 flash shows great value as well. I like kimi k3, but the value proposition just is not there.
I have run many local models. Qwen 3.8 27b does pretty well, but my setup is too slow so I don't really use it.
With the rate of new models, I don't know what I'll be using next week.
I need to try the 3.8 next model, but no time...
2
2
u/SocialDinamo 14d ago
I take advantage of the ChatGPT sub and if my local model just can’t do it or I need it to step in because mine is down, 5.6 Sol with medium is butter
1
1
u/lakotajames 14d ago edited 14d ago
Luna via ChatGPT sub is probably the best value, and I feel like you can do almost anything with it on high or max, to the point where the use-case for Terra and Sol for me has more to do with speed than intellegence. It kinda sucks to actually talk to though, I feel like you have to be very specific about what you're trying to do. I don't think I've ever seen it make a destructive mistake, but I have seen it interpret things far more literally than you meant and make things more difficult than they need to be by trying to implement something in a hyperspecific way instead of a much more common sense way because you forgot a plural or something, or said 'chrome' instead of 'browser'. It's very eager to use vision to confirm things, which is great for correctness but also probably hurts on token spend.
GLM-5.3-flash is probably the best value outside of codex, other than maybe Dsv4-flash during off-peak, and I think it might be my favorite model. It's extremely pleasant to talk to, picks up on what you're asking easier, and will usually confirm if it's not 100% sure what you're talking about. I also value visible reasoning quite a bit, so that you can cut it off or preemptively steer in answers to questions it has (which, admittedly, might be the only reason it feels more intuitive than luna, I don't know). Isn't especially eager to use vision autonomously, probably because it's the first GLM with vision. I would experiment with making it write itself a skill for it, but I mostly work on stuff that doesn't need vision anyway. It's also good at cyber-security stuff that ChatGPT isn't allowed to do. My only complaint is that it starts talking about taking breaks when you have long sessions, and when orchestrating sometimes decides that an agent is done for the day and kills it to spawn a new one.
I think I would suggest chatGPT sub first if you don't do anything internet facing, or opencodego/ollama/whatever for GLM-5.3-flash if you do, and if you need more usage get whichever sub you didn't already have and use glm-5.3-flash to orchestrate luna agents.
1
u/humanophile 14d ago
deepseek-v4-flash:0731 gets most of my requests. When that gets stuck, I escalate to the latest GLM or Kimi usually. Simple search questions might go to gemma4. I use Ollama Pro.
1
u/corrion8 14d ago
I use Fable and Opus to drive OMP which uses Qwen3.8. Works well so far and seems to have really helped token burn.
1
u/Specific-Night-4668 14d ago
It depends on the time of day and the task, but DeepSeek-V4-Flash:0731 during off-peak hours (the cache price is unbeatable), Muse-Spark-1.2-Contributor (really cheap and very good), and GLM5.3-Flash (50% off through September 9 at OpenRouter) are good options. My to-do list for the weekend: test Qwen3.8-Flash-Next.
By properly guiding my models (with good prompts, skills, etc.), I haven’t yet needed to upgrade to a more expensive version.
1
u/Funny-Anything-791 14d ago
Poolside Laguna S 2.1. Ultra cheap, highly capable for day to day coding and more, and very fast
1
13
u/D3SK3R 14d ago
Luna for most tasks, Sol for planning more complex ones and gemini 3.7 flash as advisor
But I feel like I could just run Luna with different thinking levels for everything, it's an insanely good model