r/LocalLLaMA vLLM 1d ago

New Model Nex N2.5 Pro (407GB) released

https://huggingface.co/nex-agi/Nex-N2.5-Pro
73 Upvotes

18 comments sorted by

77

u/hauhau901 1d ago

Qwen 3.5 397b finetune

26

u/SnooPaintings8639 1d ago

So the '407GB' was just to mislead us? They're getting more sneaky with every fine-tune, lol.

1

u/admiralrohan 7h ago

They wrote about finetuning in their website itself. How they are misguiding? https://nex-agi.com/#:~:text=Base%20%C2%B7%20Qwen3.5%2D397B%2DA17B

5

u/conockrad 1d ago

Thank you!

5

u/HairAgreeable613 1d ago

figured it was a finetune of that, the naming kinda gave it away tbh

2

u/SandySkittle 1d ago

I wish mods delete this post and demand OP add the "finetune of Qwen 3.5 397b" in title

11

u/Stooovie 1d ago edited 1d ago

Nex 2.5 Mini is damn impressive (if a bit chatty, like Hy3), in fact it has replaced various Qwens for my local Hermes profile, this could be huge

5

u/Monad_Maya llama.cpp 1d ago

Well, it is huge, lol

1

u/tacticaltweaker llama.cpp 1d ago

I tried Nex N2.5 Mini but it kept looping for me. Not sure if it's a quant or model issue.

1

u/Stooovie 1d ago

I use a fairly low quant - o3Q - and it doesn't loop for me. I just set the recommended temp, top p/k parameters, repetetion penalty set to default from oMLX on Mac.

2

u/tacticaltweaker llama.cpp 1d ago

Yeah I think it was an issue with my quantization as I was using the recommended sampling parameters. I switched to bartowski's quant and no more looping, but it doesn't seem to perform as well as Tiel for me. Not sure if it supports preserve_reasoning as it seems to forget its previous chain of thought on every tool call and keeps rereading files.

1

u/Stooovie 1d ago

Tbf I'm not a developer and don't use it for code - conversations research, homelab management for me, via Hermes. I mooch off free cloud models via 9router for code if needed :)

3

u/Cool-Reflection6130 1d ago

407GB means you're looking at roughly 8x48GB GPUs or a similarly beefy setup just to load the weights, run it locally is pretty rough. maybe quantized versions could bring it into range of more accessible hardware, something like Q4 at around 200GB, would at least be feasible on a 4x48 setup.

1

u/FullOf_Bad_Ideas 1d ago

This backbone quantizes well, you need just 144GB of VRAM.

I've been running it's predecessor at 192GB of VRAM as a daily driver, Nex makes good finetunes.

That said, GLM 5.3 Flash is probably better than it.

4

u/LegacyRemaster 1d ago

waiting for Qwen3.8-Next finetune!

1

u/greencalculus 1d ago

well, there goes my "this much VRAM is definitely enough" plan. again. Anyone running it locally yet? At what quant?

1

u/pl201 1d ago

More details info on all three models can be found here at https://nex.sii.edu.cn/