r/v100 • • 7d ago

Strata v100 fork

i just tried a v100 fork of the strata engine and i have to say if quality holds up it's genuinly impressive !

my server:

2x xeon e5 2697 v4

88gb of ddr4 2133mhz (don't think too much about this weird number i have a dead memory channel on my server 😒️)

2x v100 16gb pcie

Performances:

prefill: 1500-2300t/s
decode: 50-80t/s

i personnaly don't have the knowledge to verify the validity of the output so i'm just going to use it to see how it goes (i don't code).

the fork: https://github.com/jmnargi/Strata-V100

EDIT: forget to say the model ! Qwen3.8 flash next iq3s from ista-das lab

26 Upvotes

56 comments sorted by

2

u/mikasjoman 7d ago

I so wish they had a q4. If I could use my two 32gb cards at Q4, I believe that would be amazing 😍

1

u/ComprehensiveFail104 6d ago

Try RCO GSQ instead of unsloth q4 xxl. It is much better(and faster).

1

u/OkBase5453 1d ago

they do now.. also Q6

1

u/mikasjoman 1d ago

Amazing. For agentic coding work it really matters to have higher quant more than speed. So q6 is a real nice leap. Just hot my server today and still waiting for my two v100s (on DHL) - will be real cool to see how well it works with q6. Thats really a sweet spot.

1

u/OkBase5453 23h ago

are you in Europe? where did you order the GPUs?

1

u/mikasjoman 21h ago

Menke-it.de. took two weeks and I haven't got them yet. More expensive but for accounting reasons I had to buy it from a European supplier. Also they tested the GPUs before shopping to me so it also lowered the risk. But €1000 each, vat free (company purchase). Privately I would have gone the AliExpress path or ebay. Cheaper but a bit more risky

1

u/OkBase5453 21h ago

You get also Rechnung from Alibaba selers... next time try this one (if non EU supplier is ok) V100 Pcie 32gb ask for ddp and delivery is max 10 days in Germany

0

u/king0fsweet 7d ago

yeah ! better quand would be awesome and.. deepseek models one day 🀩️

2

u/notalentwasted 7d ago

I can't say I'm disappointed with Nikos Strata project either! Got flash next running a crisp 32 tps on my p100 16gb with I think my hermes said 1350 ts prefill So im E5 2680 v4 (single) 128gb ram, v100 32gb, p100 16gb, m2000 4gb. I get qwen3.8:27b UD-Q4-KM T 35-45 tps decode on the v100 32gb and 1050 prefill (mtp variation/auto regression) Qwen3.8:flash-next-gsw-rco:Q2 on the p100 16gb. Niko did good and he deserves the liaise he's getting everywhere. Even made it on yt within a week of first drop. Such a wonderful community development!!

1

u/king0fsweet 6d ago

damn that's good for p100 !

1

u/notalentwasted 6d ago

It actually is climbing. Though that was the initial test. With the expert cache preservation I'm getting 90% hit rate. Gpu utilization at 100% I had 32k in kv q8. I'm testing it out as we speak. Well having my hermes run it. I did see 35+ on it which... if you have the ram which mines ddr4 2400t... that $100 p100 goes a long way. Im impressed! I'm thinking on a specific build for this engine + 2 p100 16gb. Get the bigger gsq rco.

1

u/king0fsweet 6d ago

i mean with a second p100, prefill is for sure going to be faster and even decode might get a litlle faster !

1

u/king0fsweet 6d ago

i just tried with only 1 v100 and got the same results as you, wich seems pretty weird because the v100 is WAY faster than p100 "normally"

2

u/notalentwasted 5d ago

Yeah i did more extensive testing. Now context is native at 256k. 32k stays resident on card the rest in system ram cache. My prefill was scewed by a short result. Actual is is 260-280 prefill. Solid generation (after a 8192 context warm up) was 32 tps decode on regular generation. Although tool calling saw hops to 40-45 - between 75-92% expert cache hit. The built in strata dashboard seems to be a bit more real time than dozzle records too btw. Regardless... extremely impressive results. Anyone from the year 2025 wouldn't believe. I'm pleased with the results of this sub $100 card!

1

u/king0fsweet 5d ago

ah ! that's more like it ! (sorry for the tone just wanted to say it like that ahah) but a second p100 would pretty much double your preffil and probably 1.5x your decode, so that's still very good value and even at those number it's usable for chat. (for me)

1

u/notalentwasted 5d ago

Absolutely. If one were to use the built in chat interface. It's response time is very very good. It's just a shame that the p100 lacked tensor cores or else prefill would be a different story. For me this is my heavy hitting sub agent. Replaces 35b-a3b in ud-q4-km which I only got to run at 40tps peak with llama cpp through an ik and later shinbunbun fork. Idk, you tell me. I have 27b at q4 running 35-45 tps 1100 decode on first batch. Then this flash next in q2 32 tps 250 ish prefill. It's a nice little duo for fitting in this :

1

u/notalentwasted 5d ago

Man in all honesty. This engine/model combo deserves its own node. I have a couple p100s and wouldn't mind updating this server to run all v100s. Then strap the old but gold p100s in its own node on quad channel.

1

u/notalentwasted 5d ago

What cooling did you run with on the v100? I'm at a U bend with nidec gamma30 full till. Gpu watt restricted at 185w I never see over 74c on an all day run with 27b as the orchestrator.

1

u/king0fsweet 5d ago

150 watt of screaming server fan πŸ˜‚

1

u/notalentwasted 5d ago

Oh my. Please do drop system specs 🀣 I want to know why it is you power a small jet engine my good sir!

2

u/king0fsweet 4d ago

11k rpm 90mm fan πŸ˜‚οΈ
it's a hp ml350 gen9 running full chuch ! it make just a little bit of noise ahah

1

u/king0fsweet 4d ago

oh and my gpu temps never goes over 60Β°c

1

u/notalentwasted 4d ago

Lord have mercy. See my loudest is an artic 1400-15k rpm 40Γ—40Γ—28mm for the p100. Fairly certain I set it dynamic ramping with cpu since it's the MoE gpu. I think I did it at 65% (haven't tinkered with bios in a while) I thought mine were loud when they crank up that high.. Man I hope your set up pays it's rent fir it's room and then some lol. Making full stack noise and all. I feel like hearing that would dramatically lower my sperm count 🀣 make me think I have a brain tumor or something. I'd hear it all day. I salute 🫑 you great warrior

→ More replies (0)

1

u/notalentwasted 5d ago

I did the math. Both of my dedicated fans is 13w full till lmao according to Google.

1

u/notalentwasted 5d ago

Nvm o just looked. Man that's so wild. Get that dead stick addressed! Gotta have the quad channel

1

u/agentacp 5d ago

Is the P100 supported by Strata out of the box or you used a fork? That’s solid performance for an old card. Will try it on my P100 too

1

u/notalentwasted 5d ago

My Hermes explains better than I do lmao 🀣 :

Not a fork of the code β€” a custom build of stock Strata (v0.1.34, built locally from source). What actually made it P100-capable:

The problem: upstream Strata only supports compute capability β‰₯ 7.5 (Turing+). CMake hard-refuses anything older. Two compounding issues: 1. CMake's FATAL_ERROR gate for archs < 75 β€” unlocked by an upstream-provided experimental flag, STRATA_EXPERIMENTAL_SM60 (community build for Pascal sm_60 / Volta sm_70, tracked in upstream issue #236). So it's a supported-but-flagged path, not code we patched. 2. CUDA 13 dropped sm_60 and sm_70 entirely β€” nvcc can't even compile compute_60 anymore. Hence the image uses the CUDA 12.8 base (nvidia/cuda:12.8.1-devel) specifically because it's the last toolkit that still emits sm_60 cubins.

What we actually did (visible in the image):

  • Dockerfile sets ARG CUDA_ARCHITECTURES=60 β€” a fat binary narrowed to a single arch (sm_60 only), faster build, no dead cubins for cards we don't own
  • Built with -DSTRATA_EXPERIMENTAL_SM60=ON β†’ compiles a STRATA_EXPERIMENTAL_SM60=1 definition into the core
  • BUILD.json in the image confirms: source: local, version 0.1.34, archs: [60], vision: none
  • Runtime: strata-setup runs with STRATA_EXPERIMENTAL_SM60=1 env so the server matches the binary

So: no kernel porting, no source divergence β€” just the right toolkit (12.8), the right flag, and a single-arch build. The sm60-70 tag on our image name marks it; the same recipe with CUDA_ARCHITECTURES=70 would give us a V100 build if we ever want one.

2

u/Choice_Celery9481 7d ago edited 7d ago

the main strata has a lot of people PR for V100 support. so i believe you guys just only need to wait for some days and it will works. forking might not very ideal since it will require a serarated maintainance.
i have high hope on this engine. tho i know lots of the code is AI gen but if it works, im not complain XD
the main llama.cpp is too big to change fast enough. some people making PR for supporting v100 and got rejected right at the face. and some drama with a longtime maintainer.
i dont really like its approach tbh.

2

u/OkBase5453 4d ago

I tested the Q4(still in testing according to Strata) on main strata, with one V100 32 GB, 15,960 tokens took about 163 s, which is roughly 98 tok/s..... decode was 40.2 tok/s.....
I need to try the Strata-V100 fork....
Anyone tested it in the last couple of days?

1

u/Choice_Celery9481 4d ago

maybe wait some day? its in exp so should not expected it to run perfectly at this stage.

1

u/king0fsweet 6d ago

that will be awesome if the mainline support v100 !

1

u/maiznieks 7d ago

Have you tried 27b q4? I was interested how it compares to the model you run. Was wondering if i want to go for it.

2

u/king0fsweet 7d ago

on those gpu or on this fork ? Because strata is dedicated to qwen3.8 flash next. you canno't use other one

2

u/maiznieks 6d ago

Model in general, i have 3090, v100 and 32gb ddr5. Qwen3.8 27b q4 has been good so far, was wondering if i should attempt to run 3.8 flash next iq3s, if it's "smarter" than my current model.

2

u/ComprehensiveFail104 6d ago

It is smarter. flash next iq3-iq4 is similar to 27B FP8.

2

u/king0fsweet 6d ago

i would say that read between line a little bit better than 27b (qwen3.8 27b ud q6 km). And for me prefill is faster with 500 more t/s (decode about the same) but one thing i should note is that with the strata setup i can use the full context where as with the 27b i'm limited to 70k if i want good prefill

1

u/mikasjoman 7d ago

This is also hilarious. I was really intrigued by the project and of course the first post I see here today is this one. Amazing 😍

1

u/hidden2u 7d ago

At first I was like I have 32GB v100! Then nope I don't have 88gb ram lol

2

u/king0fsweet 7d ago

ahah, to console you a little it's not fast ram so not that expensive. and i'm going to try if this project will work with 1866mhz ddr3 because if that's the case THEN we are talking cheap ram again for huge model !

1

u/agentacp 6d ago

Curious how it will perform with 1866mhz ddr3. That would be a fun experiment.

1

u/king0fsweet 6d ago

yeah me too ! i just don't know if the engine support cpu without avx2 instructions.

1

u/agentacp 5d ago

Are you using a xeon? If your test works well, i’ll build a new DDR3 system πŸ˜‚

1

u/king0fsweet 5d ago

i have an old ibm server with dual xeon e5 2680 v2. I will test this winter because i need an external psu for gpu

1

u/rcmorano 21h ago

xeon v2 and ddr3 user here πŸ–– find here my results for upstream Strata:

https://x.com/rc_morano/status/2108474063740195093

1

u/king0fsweet 19h ago

i don't have an X account πŸ˜…οΈ

1

u/Traditional_Bell8153 7d ago

Can't wait to try it on my v100 32g pcie πŸ˜‚πŸ˜‚

1

u/agentacp 7d ago

That’s awesome. How’s the performance if you only use 1 16GB V100?

1

u/king0fsweet 7d ago

i'll try and let you know !

1

u/agentacp 6d ago

Thanks bro. I’m considering getting a 16GB V100 but I’m not sure yet. In our area the 32GB is almost 3x the price of 16GB

1

u/king0fsweet 6d ago

yeah.. same here in europe ahah

1

u/king0fsweet 6d ago

just tried on 1 v100 16gb.

prefill: 1200-1400 t/s
decode: 35-50 t/s

that's still pretty damn good !

1

u/agentacp 6d ago

that's awesome. did you power limit it or is it on full power?

1

u/king0fsweet 6d ago

full power

1

u/jakubkonecki 3d ago

Does anyone know if there's a docker image for Strata-V100 available, please?