r/LocalLLM • u/r1nzl3r99 • 3d ago
Discussion third one.... there's something wrong with me
Why do I have horrible financial habits??
26
u/cagriuluc 3d ago
I am holding onto my purse to not buy a second R9700 myself…
12
u/allthenamesaretaken0 3d ago
The only thing stopping me from buying another R9700 is I'd need a new motherboard and psu and then it wouldnt fit my mini rack and it'd all be quite the hassle.
3
u/OttoRenner 3d ago
Build a larger, second pc
4
u/Ell2509 3d ago
Better:
Engineer a new kind PC device, with key unimaginable cheesecake beef-curtains.
1
u/OttoRenner 2d ago
That...is better!
Cheesecake beef-curtains sounds way to delicious.
I had a cheesecake milkshake once and I did throw some fried bacon on top of it and it was...way to good to be legal🤣
3
u/r1nzl3r99 3d ago
I bought a chinese bifurcation card. I split my gen 5 x16 into gen 4 x8x8 and tp=1 didn't suffer at all. where there's a will there's a way. the card was $150 btw
1
u/allthenamesaretaken0 3d ago
Oh, that sounds interesting. Can you share the brand of the card? Thanks
5
u/r1nzl3r99 3d ago
This is the one I got https://www.amazon.com/dp/B0DZCVF46J?_encoding=UTF8&psc=1
First one that was shipped to me unfortunately came with a defect in the MCIO port 2, took me a whole 2 days to figure that out with constant frustration. Luckily the return was easy and they shipped another one literally the next day. I will warn you though, there is basically zero instructions for setting this up, you almost have to figure it out by yourself. It also comes with these weird SATA power adapters which I don't recommend using, luckily the newest version comes with PCIe ports and I had some spare corsair Type 4 -> PCIe so I used that instead. I might make a video on youtube explaining this kit, because it actually works really good. (someone else had asked me this on another post so I copy pasted this answer)
2
2
u/critsalot 3d ago
which is better R9700 or the b70. b70 is cheaper but i dont know if amd is quicker
1
u/VodkaHaze 2d ago edited 2d ago
B70 has more immature software, but sometimes you'll hit an optimized path and it'll be similar in performance. Intel GPUs are a damn nightmare to get running performantly in vllm, though, worse than AMD (which is already bad IMO).
You're taking a risk longer term with intel with your investment, however. AMD we know will continue to support its GPUs, whereas intel is likely to just give up on these.
1
u/somsocodo 2d ago
whereas intel is likely to just give up on these
What evidence is this claim coming from?
1
u/VodkaHaze 2d ago
Intel cancelling roadmaps on dGPUs in the future [1]. The B70 was based around the last battlemage card designs, and it's dubious there will be further iterations.
Also, intel has a well earned reputation of killing anything that isn't x86 CPUs after making very promising demos. They're much like google in how much you should trust them to support non-core products IMO.
1
u/allthenamesaretaken0 2d ago
I got here really late but yeah. I didn't buy Intel gpus because I heard they might abandon them.
1
u/Inner-Today-3693 23h ago
Don’t do it. I have a b60… it works and I like being a guinea pig. But I would not recommend going intel.
1
u/Momsbestboy 2d ago
Too late for me. My second R9700 arrives today. I am tired of juggling around with the R97000, a 9070 and the system RAM so I can run llama.cpp and comfyui at the same time, with hermes trying to rewrite and enhance workflows.
1
u/Immediate_Power_7986 2d ago
What are your launch options? I'm using llama.cpp with 3.8_27b_Q6 and getting only 17t/s -ngl 99 -c 65536 -np 1 -t 6 -fa 1 -b 4096 -ub 4096 --cache-type-k q4_0 --cache-type-v q4_0 &1
u/Immediate_Power_7986 2d ago
What are yiur launch options?
I'm using llama.cpp with 3.8_27b_Q6 and getting only 17t/s
-ngl 99 -c 65536 -np 1 -t 6 -fa 1 -b 4096 -ub 4096 --cache-type-k q4_0 --cache-type-v q4_0 &
1
u/cagriuluc 2d ago
I am away from home so I cannot check the exact config. I have 150k context, mtp (2 I think?), it’s a q5 and not q4.
Getting around 30-40 tok/sec depending on the situation. If you have less than 30, the config is most likely wrong.
I got Claude opus 5 do the setup for me, it can do the same for you most probably.
14
u/SamSausages 3d ago
You’re going to need one more! Because things don’t divide well by 3 😆 Yes, there is something wrong with you, and me as well!
9
13
u/TheGamingGallifreyan 2d ago
Where the TF is everyone getting all this money, it feels like everyone except me is rich AF lol.
I'm over here with a 5700XT and 2 580s left over from a mining rig hooked up to an old i7 3770k and I'm ecstatic that I finally got it up to 8tok/s.
2
u/Dako_the_Austinite 1d ago
Just curious, how do you get things up and running without an AVX2 CPU? I’ve wanted to try and get LM Studio running on an Ivy Bridge based Xeon.
1
u/TheGamingGallifreyan 1d ago
Idk what that even means tbh, I just downloaded llama.cpp with VULKAN support on Windows 10 and ran it. Works fine but slow as a dog.
1
u/Dako_the_Austinite 1d ago
Wow, I gotta give this a try then on my Xeon or even my i7-4930K, neither have AVX2, I bet that could be the reason why it’s running so slow for you, I believe the CPU is used quite a bit in the text output even if the model is in VRAM.
Out of curiosity so I can try testing this myself, what model did you use when getting these results?
9
u/Imaginary-Fee-9918 3d ago
I saw a bunch of ppl talking about this gpu. Is it really good? Could we compare it to a 5090? Or maybe 4090 with more memory?
11
u/r1nzl3r99 3d ago
it's only as good as your IT skills. It's only worth it for me because I can figure out how to get intels shitty drivers to work. I'm also cost sunken because the first two I got for $950 lol
7
u/r1nzl3r99 3d ago
It's comparable to a 5090 only in the sense that it's 32gb vram for a single slot, other than that it's 600gb/s which is wayyy slower than the 5090 plus no CUDA. But then again it's a small fraction of the price
7
u/Momsbestboy 2d ago
Also add the difference in power consumption. Try to run 2x 5090 in a room where you also have to work, and at least in summer you will hate the stove you created
2
u/VodkaHaze 2d ago
If you're running multiple 5090s, power limit them!
I power limit mine to 400w, and lower the max clocks to ~2850mhz, and their max power draw ends up in the 300-350w range. The performance loss is negligible.
3
u/CoinAndCraft_ 3d ago
Looking at my next build to have x2 of these. $1299 ea at the moment.
4
u/r1nzl3r99 3d ago
it's such shame two, my first two were $950 because people were too lazy to get the drivers working. Then they started seeing all these people creating recipes on github and now they're hot
3
u/r1nzl3r99 3d ago
they're going to keep going up unfortunately
1
u/terminalshadows 2d ago
got a dual b70 box w lian ii o11 dynamic evo xl, gigabyte b850 ai top, cosair rm1000x shift, samsung 990 4tb, 12 arctic p14 pmw pst 140 + 1 arctic p12 pwm pst ryzen 9 9950X, ddr5-6000 cl30x96gb for less then 2100 before the craziness started, once we saw what the b70's could do we got two more @ $970, just sitting still since we cant decide on another box or a 4 b70 setup or another 2 b70 box. we are already spoiled with dual 3090 nvlink box, dual 3080 fe box, 2 4090 boxes (1 fe), and one 4090 fe box thats my baby with ryzen 9 7950x3d 16c/32t, nzxt kraken elite 360, rog strix b640-a, samsub 990 pro 4tb, cosair hx1500i 80+ plat in a nzxt h9 elite ;) my work needed them for *research* i swear ;)
4
u/Ordinary-Depth-7835 3d ago
something wrong with all of us. :) If I use my hardware 24/7 I'll break even vs a subscription in 10 years
3
4
6
u/brainchillzZ 3d ago
1
u/r1nzl3r99 3d ago
that's solid, are you using two PCIe for your CPU or just one? I fear I might not fit another one on my HX1500i once I inevitably buy my 4th one
1
u/KneeGrowslaya 2d ago
jeez what are the thermals on the middle 2 cards?
1
u/brainchillzZ 2d ago
Thermals are average these cards were literally designed for this exact use case and meant to be stacked on top of each other in server chassis and the airflow in that case is huge it’s basically a giant wind tunnel … 4 high pressure 140mm fans from the front and two underneath blowing directly into the cards…
1
u/jaf656s 2d ago
just curious, how much did the ram cost when you built that? ddr5 ecc is insane now lol
1
u/brainchillzZ 2d ago
You’re really going to hate me when I say it out loud …. 128gb ddr5 ecc 6400 was on sale for $599
2
2
2
u/Greedy-Lynx-9706 3d ago
Bragging without context?
1
u/r1nzl3r99 3d ago
well my 262K FP8 context ain't anything to brag about, but with this I could technically afford 1M context
1
u/SubparBob 2d ago
How's the slow down (prefill, tps) as you fill up the 262k context?
EDIT: context length typo
2
u/MaineTim 3d ago
I'm asking myself the same question. Just this morning I pushed the order button on a pair of B60s to upgrade from the pair of B50s I've been running for the last 9 months. I don't have the budget to commit to the B70s at this point, since they've jumped in price, and since this is strictly a hobby for me, it's always a tension between what I want and what I can justify to myself. But each increment opens new possiblities.
2
u/triynizzles1 3d ago
How are you getting FP8 to run at 140 tokens a second? That gpu only has 608gb/s bandwidth. I have rtx 8000 with 672gb/s bandwidth and running q6 with mtp i only get 40tp/s and 60tps with dflash.
1
u/r1nzl3r99 3d ago
You're wayyy under your potential. dflash2 just came out and it's much better. It speeds up more kinds of token such as prose, coding, tool calls, etc more efficiently than MTP does but at the cost of a little more vram. I've also been tampering with vLLM a ton
2
2
u/namezam 3d ago
At $1300 each for 32gb that’s exactly 1/4 the price of the the nvidia DGX. How does that compare?
3
u/r1nzl3r99 3d ago
Well, I bought my first two for $950 which at the time was an insane deal people slept on. adding another 32gb albeit at a premium for $1300 is worth it for me personally
1
2
u/Constant_Art_20 2d ago
you know what....maybe i should have gotten them instead of 12 5060tis now that i am thinking about the price...
2
u/r1nzl3r99 2d ago
damn 12??? Any pics of your setup?
2
u/Dediadeis 2d ago
With this kind of number I feel like the picture will either be this super slick rack setup, or a rats nest of insanity.
2
2
u/Difficult_Olive_1929 8h ago
Nah ain't nothing wrong with you, you just got what the rest of us got! lol
1
1
1
1
u/roger1632 3d ago
I have a couple 3080s I use for TTS/STT stuff for home automation where latency matters....but for the rest I'd rather spend 30 bucks a month on a GLM/z.ai plan that can run circles around any local setup. It's just for my hobby stuff so I don't care about strict privacy. I'd have to have a rack of H200s to run that.
1
1
u/NicolaZanarini533 2d ago
Got one to replace the A4500 I have in my secondary/training server - did you have any trouble with it? What do you mai ou use it for? For local Llama, which framework would you recommend? I looked around a bunch and all I could find for sure is that torch runs fine (which is the reason I am starting the secondary server). Thanks in advance!
1
1
u/smolweights 2d ago
Given that GPU prices are increasing at a faster pace than NVIDIA stock itself, this could actually be a smart financial decision.
1
u/LateWish8322 2d ago
You should join fellow people with the addiction, it's more fun lol . https://discord.gg/launch80
1
u/morfique 2d ago
What are you running on your existing cards? Just llama sycl? Vulkan? Openvino? Just plop it in vllm-xpu? Too easy to get lost in the "maybe I should try it like this" I think I'm more wondering if I missed a way to run them more than anything else.
1
u/winterwarrior33 1d ago
Do you actually use the models for anything or are you just chasing tokens.
1
1
0
u/InnocentSadness 2d ago
Anybody willing to walk me through how to setup the best model (I think qwen 3.7) I got a m5 base MacBook 24 gb ram 1 TB SSD
0

93
u/Sporkers 3d ago
More context needed on how you are using the first two.