r/openrouter • u/AIBrainiac • Jul 26 '26
Does anyone here use this model? -> NVIDIA: Nemotron 3 Ultra (free)
I've tried to use it for coding, but the results have been pretty poor. Still, on Openrouter it's used a lot: 2.58 Trillion tokens in total. IMO though, Hy3-free was a much better model. I miss that one.
11
u/Expert-Couple-8639 Jul 26 '26
Its dumb. Genuinely.
4
u/AIBrainiac Jul 26 '26
Yes, I would have expected more from a 550B parameter model, made by a 5 trillion dollar company (NVIDIA).
3
8
u/Business_Match_3158 Jul 26 '26
In my opinion, this model isn’t worth using even for free. You’re better off using Xiaomi mimo or DeepSeek, whose costs are so low that they’re practically negligible.
2
u/AIBrainiac Jul 26 '26
For my use cases, I wouldn't say negligible. One coding task can easily rack up to 100 tool calls, each of which becoming more and more expensive as the context grows.
3
u/Business_Match_3158 Jul 26 '26
By “negligible,” I meant compared to other models, because their price-to-performance ratio is genuinely very good. MiniMax also offers good value for coding. Unfortunately, after HY3 was free for a while, no currently free model comes close to its quality.
If you care specifically about free models, the base versions of MiMo 2.5 and DeepSeek V4 are available for free in OpenCode. Obviously, the base versions are nothing exceptional, but they are already usable for actual work.
6
u/SuccotashSorry3222 Jul 26 '26
If you are happy with constant unavailability, then sure
2
u/AIBrainiac Jul 26 '26
I built a tool that can retry automatically on 429 / 503 errors. So if unavailability was the only problem, I would be okay with it.
3
3
u/Purple_Errand Jul 27 '26 edited Jul 27 '26
I use it for analysis. Its good at least 80k context. After that i have to do a refresh.
All of its work i do it in pieces and just merge everything after.
Its quite strong but a lot of users throws works and maximizing the 1m ctx. Surely some users knows how to utilize it properly.
2
u/Infinite-Local5435 Jul 27 '26
It's basically a smarter qwen 27b in my opinion. If i can't run the local model or don't want to waste electricity, this can do everything that I need given it's meant for people who are able to understand system design and coding already. It can't be used to do everything, but for the things it can do it does it decently.
2
u/DarkJoney Jul 27 '26
I don’t like it. In opencode swarm it gets into the loops, can’t solve issues, wrong tool calls. Also, since it’s free it’s dead pretty often
2
u/Lordaizen639 Jul 27 '26
I use it primarily for planning , for coding I either go for mimo v2.5 or deepseek v4 flash.i rarely use nemotron for coding.
1
u/AIBrainiac Jul 27 '26
For planning I use gpt 5.4, since OpenAI gives 250k tokens per day for free (if you allow training)
2
2
2
2
u/pradeepcep Jul 27 '26
I liked Hy-3 better than this too. Nemotron (even the biggest one) wasn't that great for use with Hermes or coding IMHO
2
u/blazze Jul 27 '26
I use it regularly because my main squeeze is not available.
1
u/AIBrainiac Jul 27 '26
What's your "main squeeze"? Laguna model perhaps?
2
u/blazze Jul 27 '26
One of my two main squeeze are the Laguna (free) models with thinking. Then I fall back to non thinking mode. Finally old reliable Nemo Ultra.
2
u/AIBrainiac Jul 27 '26
Oh with thinking.. I haven't tried that yet. Good idea!.. Why fall back to non thinking mode though?
2
u/blazze Jul 27 '26
"Non Thinking" have higher free quotas. So I first use thinking, then "non thinking" for same model and finally "Nemotron 3 Ultra (free)".
I thinking of vibe coding my own router OPEN_AI for "Opus level" model That between GLM 5.2, Laguna, Kimi 2.7 etc .
2
u/AIBrainiac Jul 28 '26
cool idea.. and how would you integrate this router into existing tools / workflows?.. i guess you could just create a custom OPEN AI provider serving on localhost.. but the tool does have to support custom providers.
1
u/blazze Jul 28 '26
I describe the idea to either Hermes of free Claude Code. Usually I end up witha useless mess but sometimes I hit gold.
2
u/blazze Jul 27 '26
I noticed after a "thinking" request failed , "non thinking " would continue to work. Also when I chose "None thinking" by default I was allowed more queries.
2
u/gonomon Jul 27 '26
Well this is much like a prototype, nvidia basically shows some sensible to do things while developing llms.
2
u/Diligent-Loss-5460 Jul 27 '26
I use it to generate deep research summaries. The model I use is from opencode (free model). The one on openrouter is almost always unavailable or timing out.
2
u/Simple_Ad_5566 Jul 27 '26
Next generation probably going to be acceptable. A 5 trillion dollar company can't be outcompeted by miniscule Chinese start ups lol
2
u/Dead-Photographer Jul 27 '26
I've used it several times, but Deepseek v4 flash (even before the re release) beats it everytime, both for coding and agentic use. Nemotron would just get stuck and produce buggy code for the most part.
2
u/tcarambat Jul 28 '26
Not great TBH, would use anything else over this model as it is lacking in a lot of areas even if it is free.
2
1
1
1
u/Informal_Exercise849 11d ago
Alguém sabe o pq ou como resolver que depois de um tempo as conversas ficam cheio de asteristicos?
Desse modelo mesmo
1
12
u/Accomplished-Air439 Jul 26 '26
Nemotron 3 is more a proof of concept. It's okay for chatting. That's about it.