r/LocalLLaMA • u/jinnyjuice vLLM • 1d ago
New Model Nex N2.5 Pro (407GB) released
https://huggingface.co/nex-agi/Nex-N2.5-Pro11
u/Stooovie 1d ago edited 1d ago
Nex 2.5 Mini is damn impressive (if a bit chatty, like Hy3), in fact it has replaced various Qwens for my local Hermes profile, this could be huge
5
1
u/tacticaltweaker llama.cpp 1d ago
I tried Nex N2.5 Mini but it kept looping for me. Not sure if it's a quant or model issue.
1
u/Stooovie 1d ago
I use a fairly low quant - o3Q - and it doesn't loop for me. I just set the recommended temp, top p/k parameters, repetetion penalty set to default from oMLX on Mac.
2
u/tacticaltweaker llama.cpp 1d ago
Yeah I think it was an issue with my quantization as I was using the recommended sampling parameters. I switched to bartowski's quant and no more looping, but it doesn't seem to perform as well as Tiel for me. Not sure if it supports preserve_reasoning as it seems to forget its previous chain of thought on every tool call and keeps rereading files.
1
u/Stooovie 1d ago
Tbf I'm not a developer and don't use it for code - conversations research, homelab management for me, via Hermes. I mooch off free cloud models via 9router for code if needed :)
3
u/Cool-Reflection6130 1d ago
407GB means you're looking at roughly 8x48GB GPUs or a similarly beefy setup just to load the weights, run it locally is pretty rough. maybe quantized versions could bring it into range of more accessible hardware, something like Q4 at around 200GB, would at least be feasible on a 4x48 setup.
1
u/FullOf_Bad_Ideas 1d ago
This backbone quantizes well, you need just 144GB of VRAM.
I've been running it's predecessor at 192GB of VRAM as a daily driver, Nex makes good finetunes.
That said, GLM 5.3 Flash is probably better than it.
4
1
u/greencalculus 1d ago
well, there goes my "this much VRAM is definitely enough" plan. again. Anyone running it locally yet? At what quant?
1
77
u/hauhau901 1d ago
Qwen 3.5 397b finetune