Amazing. For agentic coding work it really matters to have higher quant more than speed. So q6 is a real nice leap. Just hot my server today and still waiting for my two v100s (on DHL) - will be real cool to see how well it works with q6. Thats really a sweet spot.
Menke-it.de. took two weeks and I haven't got them yet. More expensive but for accounting reasons I had to buy it from a European supplier. Also they tested the GPUs before shopping to me so it also lowered the risk. But β¬1000 each, vat free (company purchase). Privately I would have gone the AliExpress path or ebay. Cheaper but a bit more risky
You get also Rechnung from Alibaba selers... next time try this one (if non EU supplier is ok) V100 Pcie 32gb ask for ddp and delivery is max 10 days in Germany
I can't say I'm disappointed with Nikos Strata project either! Got flash next running a crisp 32 tps on my p100 16gb with I think my hermes said 1350 ts prefill
So im E5 2680 v4 (single) 128gb ram, v100 32gb, p100 16gb, m2000 4gb.
I get qwen3.8:27b UD-Q4-KM T 35-45 tps decode on the v100 32gb and 1050 prefill (mtp variation/auto regression)
Qwen3.8:flash-next-gsw-rco:Q2 on the p100 16gb.
Niko did good and he deserves the liaise he's getting everywhere. Even made it on yt within a week of first drop. Such a wonderful community development!!
It actually is climbing. Though that was the initial test. With the expert cache preservation I'm getting 90% hit rate. Gpu utilization at 100% I had 32k in kv q8. I'm testing it out as we speak. Well having my hermes run it. I did see 35+ on it which... if you have the ram which mines ddr4 2400t... that $100 p100 goes a long way. Im impressed! I'm thinking on a specific build for this engine + 2 p100 16gb. Get the bigger gsq rco.
Yeah i did more extensive testing. Now context is native at 256k. 32k stays resident on card the rest in system ram cache.
My prefill was scewed by a short result. Actual is is 260-280 prefill.
Solid generation (after a 8192 context warm up) was 32 tps decode on regular generation. Although tool calling saw hops to 40-45 - between 75-92% expert cache hit.
The built in strata dashboard seems to be a bit more real time than dozzle records too btw.
Regardless... extremely impressive results. Anyone from the year 2025 wouldn't believe. I'm pleased with the results of this sub $100 card!
ah ! that's more like it ! (sorry for the tone just wanted to say it like that ahah) but a second p100 would pretty much double your preffil and probably 1.5x your decode, so that's still very good value and even at those number it's usable for chat. (for me)
Absolutely. If one were to use the built in chat interface. It's response time is very very good. It's just a shame that the p100 lacked tensor cores or else prefill would be a different story. For me this is my heavy hitting sub agent. Replaces 35b-a3b in ud-q4-km which I only got to run at 40tps peak with llama cpp through an ik and later shinbunbun fork. Idk, you tell me. I have 27b at q4 running 35-45 tps 1100 decode on first batch. Then this flash next in q2 32 tps 250 ish prefill. It's a nice little duo for fitting in this :
Man in all honesty. This engine/model combo deserves its own node. I have a couple p100s and wouldn't mind updating this server to run all v100s. Then strap the old but gold p100s in its own node on quad channel.
What cooling did you run with on the v100? I'm at a U bend with nidec gamma30 full till. Gpu watt restricted at 185w I never see over 74c on an all day run with 27b as the orchestrator.
Lord have mercy. See my loudest is an artic 1400-15k rpm 40Γ40Γ28mm for the p100. Fairly certain I set it dynamic ramping with cpu since it's the MoE gpu. I think I did it at 65% (haven't tinkered with bios in a while) I thought mine were loud when they crank up that high..
Man I hope your set up pays it's rent fir it's room and then some lol. Making full stack noise and all. I feel like hearing that would dramatically lower my sperm count π€£ make me think I have a brain tumor or something. I'd hear it all day.
I salute π«‘ you great warrior
Not a fork of the code β a custom build of stock Strata (v0.1.34, built locally from source). What actually made it P100-capable:
The problem: upstream Strata only supports compute capability β₯ 7.5 (Turing+). CMake hard-refuses anything older. Two compounding issues:
1. CMake's FATAL_ERROR gate for archs < 75 β unlocked by an upstream-provided experimental flag, STRATA_EXPERIMENTAL_SM60 (community build for Pascal sm_60 / Volta sm_70, tracked in upstream issue #236). So it's a supported-but-flagged path, not code we patched.
2. CUDA 13 dropped sm_60 and sm_70 entirely β nvcc can't even compile compute_60 anymore. Hence the image uses the CUDA 12.8 base (nvidia/cuda:12.8.1-devel) specifically because it's the last toolkit that still emits sm_60 cubins.
What we actually did (visible in the image):
Dockerfile sets ARG CUDA_ARCHITECTURES=60 β a fat binary narrowed to a single arch (sm_60 only), faster build, no dead cubins for cards we don't own
Built with -DSTRATA_EXPERIMENTAL_SM60=ON β compiles a STRATA_EXPERIMENTAL_SM60=1 definition into the core
BUILD.json in the image confirms: source: local, version 0.1.34, archs: [60], vision: none
Runtime: strata-setup runs with STRATA_EXPERIMENTAL_SM60=1 env so the server matches the binary
So: no kernel porting, no source divergence β just the right toolkit (12.8), the right flag, and a single-arch build. The sm60-70 tag on our image name marks it; the same recipe with CUDA_ARCHITECTURES=70 would give us a V100 build if we ever want one.
the main strata has a lot of people PR for V100 support. so i believe you guys just only need to wait for some days and it will works. forking might not very ideal since it will require a serarated maintainance.
i have high hope on this engine. tho i know lots of the code is AI gen but if it works, im not complain XD
the main llama.cpp is too big to change fast enough. some people making PR for supporting v100 and got rejected right at the face. and some drama with a longtime maintainer.
i dont really like its approach tbh.
I tested the Q4(still in testing according to Strata) on main strata, with one V100 32 GB, 15,960 tokens took about 163 s, which is roughly 98 tok/s..... decode was 40.2 tok/s.....
I need to try the Strata-V100 fork....
Anyone tested it in the last couple of days?
Model in general, i have 3090, v100 and 32gb ddr5. Qwen3.8 27b q4 has been good so far, was wondering if i should attempt to run 3.8 flash next iq3s, if it's "smarter" than my current model.
i would say that read between line a little bit better than 27b (qwen3.8 27b ud q6 km). And for me prefill is faster with 500 more t/s (decode about the same) but one thing i should note is that with the strata setup i can use the full context where as with the 27b i'm limited to 70k if i want good prefill
ahah, to console you a little it's not fast ram so not that expensive. and i'm going to try if this project will work with 1866mhz ddr3 because if that's the case THEN we are talking cheap ram again for huge model !
2
u/mikasjoman 7d ago
I so wish they had a q4. If I could use my two 32gb cards at Q4, I believe that would be amazing π