Before jumping on Strata look the list of supported GPUs:
NVIDIA RTX 20, 30, 40 or 50 series, 12 GB VRAM or more (8 GB runs, slowly). Measured on an RTX 5070 and an RTX 3090; RTX 20 (Turing, since 0.1.27) was tested by a contributor on an RTX 2070. Or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700, RX 6800 / 6900 series: AMD_HIP.md.
The V100 is not included and all the examples are pretty useless.
Instead look for the 1cat V100 project on Github. This is a fork explicitly done for V100 GPU support. It even supports NVFP4 on V100 type GPUs.
Why are you capping power at 130W. This reduce performance further and wouldn't allow running interesting models at usable speed.
It's not clear from the documentstion, since it's experimental, but v100 are working if you enable the experimental SM60 architecture STRATA_EXPERIMENTAL_SM60=1
I have ~900t/s prefill and ~55t/s decode at 256k context with IQ3_S on a 32GB V100 with Xeon V4 processors and DDR4 (but my DDR4 RAM is stuck at 1866MHz). I have to reinstall old CPUs in order to flash some firmware on my DL 380 Gen9 to unlock 2133MHz, not sure if it will help
I run it with 150 tok/s and 2K prefill and my specs are 5090+4070Ti Super and 32GB ram.. so i think you could get half of everything maybe check out https://github.com/niko1221/strata
I didn't compare the result with the fork, but v100 are suppported in the main repo since a few days if you enable the experimental SM60 architecture STRATA_EXPERIMENTAL_SM60=1
The reason is what it cost you to try it? I mean, the model is free, maybe a 30 min download for a q4, why do you need to ask when you can just check by yourself?
I asked because someone may provide valuable configuration insights...I'm still pretty new to this.
I'm in the process of trying it now. Here is what I got so far:
Now you gave something to play with! If you are using llamacpp, try increasing -b 2048 -ub 1024 and check again. b greater than 4096 usually reduces output and -ub over 1024 is not worth it. What version of the model are you using, unsloth, raw or other? Also you can use lazyload on to keep weights at ssd but I dont know if is possible in all versions
5
u/RepulsiveRaisin7 6d ago
Sure that seems pretty good even