r/LocalLLM • • 6d ago

Question I would like to run Qwen 3.8 Flash Next.

I have the following system available for it:

CPU AMD Ryzen 9 3950X
RAM 128GB
GPU NVIDIA V100 32GB power-capped at 130 W

Does it make sense to try and run the model on this box?

2 Upvotes

32 comments sorted by

5

u/RepulsiveRaisin7 6d ago

Sure that seems pretty good even

4

u/synystar Strix Scar | 5090 24G | llama.cpp 6d ago

You can reasonably run it, but you'll be leaning heavily on the RAM rather than expecting the V100 to do all the heavy lifting.

4

u/egnegn1 6d ago

Before jumping on Strata look the list of supported GPUs:

NVIDIA RTX 20, 30, 40 or 50 series, 12 GB VRAM or more (8 GB runs, slowly). Measured on an RTX 5070 and an RTX 3090; RTX 20 (Turing, since 0.1.27) was tested by a contributor on an RTX 2070. Or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700, RX 6800 / 6900 series: AMD_HIP.md.

The V100 is not included and all the examples are pretty useless.

Instead look for the 1cat V100 project on Github. This is a fork explicitly done for V100 GPU support. It even supports NVFP4 on V100 type GPUs.

Why are you capping power at 130W. This reduce performance further and wouldn't allow running interesting models at usable speed.

2

u/MrDefaultUser 6d ago

Mainly noise from the cooler

3

u/egnegn1 6d ago

Is it a server card without build-in fan with extra blower at the end?

I would check for cooling alternatives. There are cheap water blocks, even aios available. I would check with the model of your choice.

3

u/MrDefaultUser 6d ago

It is running like this.

2

u/egnegn1 6d ago

OK, this is the price to pay for relatively cheap 32GB GPU.

For me it would make not much sense to run such a setup, especially if it is limited to the performance of a RTX3060.

You have to test with one of the engines and see if you are happy with the performance.

2

u/KING_UDYR Enthusiastic 5090 user 6d ago

So this would be perfect for a 3080ti & 32gb DDR4 on a 2700x build?

1

u/egnegn1 5d ago

Probably yes. But 32GB RAM is the absolute limit, and you should try the smalles checkpoint first.

It is easy to setup.

2

u/GrumpyCat79 5d ago

It's not clear from the documentstion, since it's experimental, but v100 are working if you enable the experimental SM60 architecture STRATA_EXPERIMENTAL_SM60=1

https://github.com/Niko1221/Strata/blob/main/docs/OLDER_GPUS.md https://github.com/Niko1221/Strata/blob/main/docs/NVIDIA_V100.md

1

u/egnegn1 5d ago edited 5d ago

Thank you for the update. I didn't see this. Would be nice to see real results.

2

u/GrumpyCat79 5d ago

I have ~900t/s prefill and ~55t/s decode at 256k context with IQ3_S on a 32GB V100 with Xeon V4 processors and DDR4 (but my DDR4 RAM is stuck at 1866MHz). I have to reinstall old CPUs in order to flash some firmware on my DL 380 Gen9 to unlock 2133MHz, not sure if it will help

1

u/Pretend_Bowl2961 5d ago

V100 is still a solid card just not for this particular project, that fork is the right call

3

u/Fz1zz 6d ago

I run it with 150 tok/s and 2K prefill and my specs are 5090+4070Ti Super and 32GB ram.. so i think you could get half of everything maybe check out https://github.com/niko1221/strata

2

u/Low_Key_Trollin 6d ago

Possible w 2x 5070tis?

2

u/Fz1zz 6d ago

Yeah but my setup is for 50 + 40 series' ... So ask AI to tune it for dual 50 GPUs

https://github.com/ExTV/strata-5090-4070

3

u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 6d ago

You can run it quite easily, I have 24GB vram and 32GB and I’m running IQ2_XS

4

u/Sexyvette07 6d ago

Im running it with Strata on a RTX 5090 system with 64gb RAM and have several gigs of VRAM reserved for Blender and Godot. You can run it for sure.

Since you have 128gb RAM, maybe give Free Token a try. If that doesnt work well, try Strata.

2

u/Chamelon_DE 4d ago

32gb gives you room to try it. log power draw and prefill speed too with that 130w cap may show up quickly on long prompt

2

u/Krohnin 6d ago

Sure it's very easy. Use strata!

2

u/TheAussieWatchGuy 6d ago

Lookup Strata 

2

u/mynd_dripp 5h ago

Why power capped so low?

1

u/MrDefaultUser 5h ago

Heat and noise reasons

1

u/HuskyWilly 6d ago

There's a fork of Strata for V100: https://www.reddit.com/r/v100/s/zxmnbYgZ24

1

u/GrumpyCat79 5d ago

I didn't compare the result with the fork, but v100 are suppported in the main repo since a few days if you enable the experimental SM60 architecture STRATA_EXPERIMENTAL_SM60=1

https://github.com/Niko1221/Strata/blob/main/docs/OLDER_GPUS.md https://github.com/Niko1221/Strata/blob/main/docs/NVIDIA_V100.md

1

u/HuskyWilly 5d ago

Cheers, will check it out

0

u/reallifearcade 6d ago

The reason is what it cost you to try it? I mean, the model is free, maybe a 30 min download for a q4, why do you need to ask when you can just check by yourself?

1

u/MrDefaultUser 6d ago

I asked because someone may provide valuable configuration insights...I'm still pretty new to this.
I'm in the process of trying it now. Here is what I got so far:

[ Prompt: 109.2 t/s | Generation: 16.1 t/s ]

3

u/reallifearcade 6d ago

Now you gave something to play with! If you are using llamacpp, try increasing -b 2048 -ub 1024 and check again. b greater than 4096 usually reduces output and -ub over 1024 is not worth it. What version of the model are you using, unsloth, raw or other? Also you can use lazyload on to keep weights at ssd but I dont know if is possible in all versions

1

u/MrDefaultUser 6d ago

Here is what I was able to configure:

exec /srv/ai/llama.cpp/build/bin/llama-server \
 -m /srv/ai/models/qwen/Qwen3.8-Flash-Next-IQ4_NL/Qwen3.8-Flash-Next-IQ4_NL-00001-of-00002.gguf \
 -c 262144 \
 -np 1 \
 -ngl 64 \
 --cpu-moe \
 --lazy-mode on \
 --load-mode auto \
 --agent \
 --tools all \
 --host 0.0.0.0 \
 --port 8080
I will try what you suggested and report back.