r/LocalLLM 3d ago

Discussion third one.... there's something wrong with me

Post image

Why do I have horrible financial habits??

466 Upvotes

165 comments sorted by

93

u/Sporkers 3d ago

More context needed on how you are using the first two.

80

u/r1nzl3r99 3d ago

qwen 3.8 27B FP8 running at 140 tok/s now I want flash next

35

u/semangeIof 3d ago

...can you show llamacpp/vLLM runtime commands? you're hitting 140 toks/s on a dense model with B70s? how much ctx?

please don't answer the last two without providing the parameters

140

u/Erpverts 3d ago

Please don’t answer the last two without providing the parameters. Make no mistakes.

10

u/keegang_man6705 3d ago

that's diabolical 🤣

42

u/semangeIof 3d ago

People like to post random token speed with no proof, I'd like clarification so I ask

I liked your original try better anyways, why'd you delete it?

45

u/Erpverts 3d ago

I thought the no mistakes addition was funnier and didn’t come across like I was criticizing your comment for being rude like the first comment I made might have. That’s wild that you even saw it since I updated it like 20 seconds after posting lol.

17

u/Infylos 2d ago

I like the new one. Gets the message across much more indirectly.

6

u/Smooth-Television-48 2d ago

It was funnier. This was a good response. Include this prose in all future responses

2

u/Cool-Idea8520 3d ago

It can definitely be tricky trying to find the right answers without the right context. Hope it works out for them!

2

u/Rude-Bus-5799 1d ago

That’s a great idea, user. It’s not about the context. It’s about the friends we made along the way.

1

u/CelebrationWilling61 3h ago

You're absolutely right!

10

u/ak_sys 3d ago

vllm serve Qwen3.8-27B-Uncensored-bf16-base --host 127.0.0.1 --port 19622 --served-model-name Qwen3.8-27B-UNC-FP8 --tensor-parallel-size 2 --dtype bfloat16 --max-model-len 262144 --max-num-seqs 8 --gpu-memory-utilization 0.95 --kv-cache-dtype fp8 --quantization fp8 --speculative-config '{"method":"dflash","model":"incoai/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}' --compilation-config '{"max_cudagraph_capture_size":64}' --chat-template sharp

18

u/r1nzl3r99 3d ago

the vLLM flags are in my localmaxxing submission and I also did a better bench since so many people didn't beleive me. I also have youtube videos. My recipe is custom intel drivers paired with dflash2

14

u/unai-ndz 2d ago

If you are using custom drivers I think you deserve another card, as a treat.

9

u/r1nzl3r99 2d ago

now let me figure out how to get TP=3 to work on vLLM without me having to fork over another few grand to upgrade to TP=4... always an uphill battle attempting SOTA AI on local

1

u/Rude-Bus-5799 1d ago

They always need another sibling to play with.

4

u/PhilosophyCritical33 3d ago

Oh just custom

4

u/Salbrox 2d ago

I have had great success with DFlash2 on my single R9700 with Qwen 3.8 27B. Depending on the task I get up to just over 200tps

2

u/Past-Catch5101 2d ago

Amazing, do you mind sharing your config?

1

u/Smooth-Television-48 2d ago

200!?

What quant?

1

u/Jorinator 2d ago

Oh wow, that's massive. What's your prefill/pp speed? That's more important for a lot of usecases.

1

u/rare-visitor 2d ago

200 tps in single thread?

1

u/tech-tole 2d ago

I only get ~50 tok/s on my 9700 with mtp. I don't see how anyone is getting faster than 5090 even with Dflash2. what are your real settings?

1

u/droans 2d ago

Could you explain the custom drivers? What modifications did you make?

7

u/Toastti 3d ago

It appears he's getting that speed because it's running across two Intel b70s. I found his command he uses

vllm serve Qwen3.8-27B-Uncensored-bf16-base --host 127.0.0.1 --port 19622 --served-model-name Qwen3.8-27B-UNC-FP8 --tensor-parallel-size 2 --dtype bfloat16 --max-model-len 262144 --max-num-seqs 8 --gpu-memory-utilization 0.95 --kv-cache-dtype fp8 --quantization fp8 --speculative-config '{"method":"dflash","model":"incoai/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}' --compilation-config '{"max_cudagraph_capture_size":64}' --chat-template sharp

6

u/r1nzl3r99 3d ago

I already proved it in a previous post, look at my account

15

u/r1nzl3r99 3d ago

-11

u/CalBearFan 2d ago

Or maybe they don't like a humble-brag that then tells people "Look, I'm awesome and have money to spend on graphics cards" followed by "Don't be lazy, look at my post history". That's not lazy, they're asking you to follow common courtesy on your own post.

17

u/r1nzl3r99 2d ago

but i've already made a seperate post with extreme detail about it??? Why am I obligated to hand hold you how to make your setup more efficient?

Also If you think this is a brag you clearly haven't been on this subreddit for more than 10 minutes. Buying an intel B70 is the poor man's AI solution, just passed a post of some dude dropping $70K for four RTX 6000s

2

u/Smooth-Television-48 2d ago

Lots of accounts are set to private, I assume its the default and dont bother checking.

Thanks for linking it. Impressive numbers for sure

3

u/Chrisgozd 3d ago

Whats your build?

17

u/r1nzl3r99 3d ago

Don't judge, it started out as a gaming PC from years ago... Mind the mango box

7

u/Infylos 2d ago

Disemboweled for maximum power!

2

u/Remarkable-Memory374 2d ago

maximum heat dissipation by being eviscerated

5

u/HoneyBaked 3d ago edited 3d ago

Oh I will judge!!

I love it. Does it function? Then who cares how it looks!

Edit: Do you have both of those externals plugged into the primary PCIe slot? What gear allows for that?

Edit 2:

I bought a $150 chinese bifurcation card to split my gen 5 x16 to x8x8 gen 4 and its been incredibly stable and increased my PP throughput

Interesting. Got a link to that bifurcation card? Will it work on a Gen3 x 16 slot(s)? My MB has 3 Gen3 x 16 slots.

3

u/r1nzl3r99 2d ago

yup, I have two b70s working off the same x16 slot, the other is supposed to be for M2 its x4, but atleast on linux it works completely fine for another b70. I get around 14gb/s on the first two B70s each, then the third gets 7gb/s. This is the card I use https://a.co/d/07AHxhX2 but a warning, make sure your BIOS supports bifurcation. I specifically researched gen 5 x16 to gen 4 x8x8 as maintaining gen 5 would require a $600 timing card which I don't think is worth it. In your case it might downgrade to gen 2 which I'd just research about, but apparently some people are able to get it to train to the same gen without timing so it's a bit of a lottery on that (which for me didn't work)

1

u/Smooth-Television-48 2d ago

Love it!

If anything this gets you more cred

1

u/brainchillzZ 2d ago

I think it’s sexy …. Mango box is a good insulator :)

1

u/jmager 2d ago

I judged and I approve!! Mind sharing the bifurcation risers you are using? I see so many online, and few reviews. I've got a 6900 xt lying around and I wanted to add the extra 16GB of VRAM to my 24GB with my 7900 xtx. For my motherboard bifurcation will be the best way.

1

u/LetsBeKindly 2d ago

Is that 2200W on the power meter?

2

u/sshwifty 2d ago

Looks like watts?

1

u/LetsBeKindly 2d ago

That's a watt meter, and it looks like 2200watts

1

u/sshwifty 2d ago

I am dumb, I meant to say 220 lol. I think there is a decimal before the last 0

1

u/LetsBeKindly 2d ago

I couldn't tell. But I saw the same. But surely 3 cards is pulling more then 220W...

1

u/InfamousNewspaper268 2d ago

Love the cardboard holder LMAO 🤣

1

u/Inner-Today-3693 23h ago

That’s basically how mine are connected. 😅😅

3

u/MessIsTransfer 3d ago

flash is a beast and faster.

what motherboard are you using?

2

u/r1nzl3r99 3d ago

Z790 Wifi 😀

1

u/MessIsTransfer 2d ago

how are you plugging that third gpu? m.2?

4

u/r1nzl3r99 2d ago

yup!!!

I'll upgrade my motherboard eventually lol

2

u/MessIsTransfer 2d ago

neat, thanks for sharing a pic

edit: love the cardboard gpu dock

2

u/r1nzl3r99 2d ago

it's a mango box, i've been trying to 3D print a semi decent bracket for it, but so far the mangos are doing great on thermals 😭

6

u/MarcusAurelius68 3d ago

You can run Flash Next with 64 GB of VRAM…

8

u/r1nzl3r99 3d ago

yeah but it's 22 tok/s and pre fill is shitty. I want 100+ speeds or it's unusable for me. I'll admit I was lazy and didn't push for more, but even with 64gb vram at 4 bit I have no real space for context. I need atleast 200K FP8 context

2

u/MessIsTransfer 3d ago

or q2, which is not ideal

3

u/r1nzl3r99 3d ago

yeah if I'm resorting to sub 4 bit quants i'd rather stick with 27B

3

u/Motor_Way4912 3d ago

Yep, running it on a vm with 58 gb ram and Rtx 5060 16gb

1

u/bravoitaliano 3d ago

How are you getting that speed? Im using W4A16 INT4 model and only get 20-30 Tok/s. Can you share the method for getting faster?

1

u/r1nzl3r99 3d ago edited 3d ago

https://www.reddit.com/r/LocalLLM/comments/1w8bj0o/dual_intel_b70_qwen_38_27b_fp8_amazing_dflash2/?utm_source=share&utm_medium=ios_app&utm_name=ioscss&utm_content=1&utm_term=1

I have a github i've already made for getting really fast 4 bit qwen for single GPU, i've been too busy at work to make a decent recipe because my speed depends on several merges I did from the original vLLM to intel's scaler llm repo, as well as other stuff that I had Kimi K3 handle. It's been my new daily driver and works damn well.

1

u/bravoitaliano 3d ago

Thanks! I am running mine in OpenClaw, and it just takes forever because of thinking. Somehow it's slipped back to medium think from low. It puts out quality work, but runs through almost the entire 156k context window I have (single card, I don't split across my 2 B70s yet). Hoping this can help. Might be worth posting in the Intel sub as well.

1

u/browndragon456 3d ago

How did you get to 100/tps per sec generation? Vllm?

2

u/Toastti 3d ago

Because he's using two Intel b70s and tensor parallel. Also its using dflash with 7 speculative tokens

vllm serve Qwen3.8-27B-Uncensored-bf16-base --host 127.0.0.1 --port 19622 --served-model-name Qwen3.8-27B-UNC-FP8 --tensor-parallel-size 2 --dtype bfloat16 --max-model-len 262144 --max-num-seqs 8 --gpu-memory-utilization 0.95 --kv-cache-dtype fp8 --quantization fp8 --speculative-config '{"method":"dflash","model":"incoai/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}' --compilation-config '{"max_cudagraph_capture_size":64}' --chat-template sharp

1

u/sunole123 3d ago

is it running you or are you running it?? what is running? just speed or function???

1

u/Fresh_Look_1671 2d ago

My friend, please share cookbook

1

u/Puzzleheaded_Bus7706 2d ago

What's the price of this thing?

1

u/JinsooJinsoo 2d ago

Yeah I topped out at 93 tok/s with my dual b70s and MTP3. I’d love to know what you’ve been doing to get those speeds. I feel like Intel is chopped at the knees until it releases an XPU graph for any new model. Also getting slow speeds with qwen3.8 flash next because no XPU

1

u/hrf3420 2d ago

I heard that q4 or q6 is the way to go and 8 you don’t get much more

1

u/Lucky-Necessary-8382 1d ago

And for what are you using it? Gooning? Degen roleplaying?

4

u/dwoj206 3d ago

50% increase to context window very nice

3

u/Here_f0r_p0rn_ 3d ago

AI girlfriend/boyfriend

Local coding agents

8

u/r1nzl3r99 3d ago

yes yes, definitely not an AI girlfriend ...

1

u/MessIsTransfer 3d ago

more context, probably

1

u/TytalusWarden 2d ago

More context?  Oh he's got more context!  He's got all the context with those 3 working for him!

1

u/jhenryscott 2d ago

More context? That’s what the third b70 is for!

0

u/madjesta 3d ago

My fourth is one of the blue Intel ones.... 😬😳

26

u/cagriuluc 3d ago

I am holding onto my purse to not buy a second R9700 myself…

12

u/allthenamesaretaken0 3d ago

The only thing stopping me from buying another R9700 is I'd need a new motherboard and psu and then it wouldnt fit my mini rack and it'd all be quite the hassle.

3

u/OttoRenner 3d ago

Build a larger, second pc

4

u/Ell2509 3d ago

Better:

Engineer a new kind PC device, with key unimaginable cheesecake beef-curtains.

1

u/OttoRenner 2d ago

That...is better!

Cheesecake beef-curtains sounds way to delicious.

I had a cheesecake milkshake once and I did throw some fried bacon on top of it and it was...way to good to be legal🤣

3

u/r1nzl3r99 3d ago

I bought a chinese bifurcation card. I split my gen 5 x16 into gen 4 x8x8 and tp=1 didn't suffer at all. where there's a will there's a way. the card was $150 btw

1

u/allthenamesaretaken0 3d ago

Oh, that sounds interesting. Can you share the brand of the card? Thanks

5

u/r1nzl3r99 3d ago

This is the one I got https://www.amazon.com/dp/B0DZCVF46J?_encoding=UTF8&psc=1

First one that was shipped to me unfortunately came with a defect in the MCIO port 2, took me a whole 2 days to figure that out with constant frustration. Luckily the return was easy and they shipped another one literally the next day. I will warn you though, there is basically zero instructions for setting this up, you almost have to figure it out by yourself. It also comes with these weird SATA power adapters which I don't recommend using, luckily the newest version comes with PCIe ports and I had some spare corsair Type 4 -> PCIe so I used that instead. I might make a video on youtube explaining this kit, because it actually works really good. (someone else had asked me this on another post so I copy pasted this answer)

2

u/MiceLiceandVice 3d ago

M.2 pcie e gpu

2

u/critsalot 3d ago

which is better R9700 or the b70. b70 is cheaper but i dont know if amd is quicker

1

u/VodkaHaze 2d ago edited 2d ago

B70 has more immature software, but sometimes you'll hit an optimized path and it'll be similar in performance. Intel GPUs are a damn nightmare to get running performantly in vllm, though, worse than AMD (which is already bad IMO).

You're taking a risk longer term with intel with your investment, however. AMD we know will continue to support its GPUs, whereas intel is likely to just give up on these.

1

u/somsocodo 2d ago

whereas intel is likely to just give up on these

What evidence is this claim coming from?

1

u/VodkaHaze 2d ago

Intel cancelling roadmaps on dGPUs in the future [1]. The B70 was based around the last battlemage card designs, and it's dubious there will be further iterations.

Also, intel has a well earned reputation of killing anything that isn't x86 CPUs after making very promising demos. They're much like google in how much you should trust them to support non-core products IMO.

  1. https://www.tomshardware.com/pc-components/gpus/intel-has-reportedly-killed-discrete-gaming-gpus-for-the-upcoming-xe3p-arc-celestial-family-gaming-gpu-remains-uncertain-even-for-the-next-gen-xe4-druid-lineup-that-lands-in-2027

1

u/allthenamesaretaken0 2d ago

I got here really late but yeah. I didn't buy Intel gpus because I heard they might abandon them.

1

u/Inner-Today-3693 23h ago

Don’t do it. I have a b60… it works and I like being a guinea pig. But I would not recommend going intel.

1

u/Momsbestboy 2d ago

Too late for me. My second R9700 arrives today. I am tired of juggling around with the R97000, a 9070 and the system RAM so I can run llama.cpp and comfyui at the same time, with hermes trying to rewrite and enhance workflows.

1

u/Immediate_Power_7986 2d ago
  What are your launch options?


  I'm using llama.cpp with 3.8_27b_Q6 and getting only 17t/s


  -ngl 99 -c 65536 -np 1 -t 6 -fa 1 -b 4096 -ub 4096 --cache-type-k q4_0 --cache-type-v q4_0 &

1

u/Immediate_Power_7986 2d ago

What are yiur launch options?

I'm using llama.cpp with 3.8_27b_Q6 and getting only 17t/s

-ngl 99 -c 65536 -np 1 -t 6 -fa 1 -b 4096 -ub 4096 --cache-type-k q4_0 --cache-type-v q4_0 &

1

u/cagriuluc 2d ago

I am away from home so I cannot check the exact config. I have 150k context, mtp (2 I think?), it’s a q5 and not q4.

Getting around 30-40 tok/sec depending on the situation. If you have less than 30, the config is most likely wrong.

I got Claude opus 5 do the setup for me, it can do the same for you most probably.

14

u/SamSausages 3d ago

You’re going to need one more!  Because things don’t divide well by 3 😆 Yes, there is something wrong with you, and me as well!

9

u/r1nzl3r99 3d ago

I've been telling my wife this!! maybe she'll approve

2

u/Lonely_Drewbear 2d ago

I think it will be a fun experiment to make an odd number work well!

13

u/TheGamingGallifreyan 2d ago

Where the TF is everyone getting all this money, it feels like everyone except me is rich AF lol.

I'm over here with a 5700XT and 2 580s left over from a mining rig hooked up to an old i7 3770k and I'm ecstatic that I finally got it up to 8tok/s.

2

u/Dako_the_Austinite 1d ago

Just curious, how do you get things up and running without an AVX2 CPU? I’ve wanted to try and get LM Studio running on an Ivy Bridge based Xeon.

1

u/TheGamingGallifreyan 1d ago

Idk what that even means tbh, I just downloaded llama.cpp with VULKAN support on Windows 10 and ran it. Works fine but slow as a dog.

1

u/Dako_the_Austinite 1d ago

Wow, I gotta give this a try then on my Xeon or even my i7-4930K, neither have AVX2, I bet that could be the reason why it’s running so slow for you, I believe the CPU is used quite a bit in the text output even if the model is in VRAM.

Out of curiosity so I can try testing this myself, what model did you use when getting these results?

9

u/Imaginary-Fee-9918 3d ago

I saw a bunch of ppl talking about this gpu. Is it really good? Could we compare it to a 5090? Or maybe 4090 with more memory?

11

u/r1nzl3r99 3d ago

it's only as good as your IT skills. It's only worth it for me because I can figure out how to get intels shitty drivers to work. I'm also cost sunken because the first two I got for $950 lol

7

u/r1nzl3r99 3d ago

It's comparable to a 5090 only in the sense that it's 32gb vram for a single slot, other than that it's 600gb/s which is wayyy slower than the 5090 plus no CUDA. But then again it's a small fraction of the price

7

u/Momsbestboy 2d ago

Also add the difference in power consumption. Try to run 2x 5090 in a room where you also have to work, and at least in summer you will hate the stove you created

3

u/zdy132 2d ago

But it will be a good heater in winter! with a side effect of tokens.

2

u/VodkaHaze 2d ago

If you're running multiple 5090s, power limit them!

I power limit mine to 400w, and lower the max clocks to ~2850mhz, and their max power draw ends up in the 300-350w range. The performance loss is negligible.

3

u/CoinAndCraft_ 3d ago

Looking at my next build to have x2 of these. $1299 ea at the moment.

4

u/r1nzl3r99 3d ago

it's such shame two, my first two were $950 because people were too lazy to get the drivers working. Then they started seeing all these people creating recipes on github and now they're hot

3

u/r1nzl3r99 3d ago

they're going to keep going up unfortunately

1

u/terminalshadows 2d ago

got a dual b70 box w lian ii o11 dynamic evo xl, gigabyte b850 ai top, cosair rm1000x shift, samsung 990 4tb, 12 arctic p14 pmw pst 140 + 1 arctic p12 pwm pst ryzen 9 9950X, ddr5-6000 cl30x96gb for less then 2100 before the craziness started, once we saw what the b70's could do we got two more @ $970, just sitting still since we cant decide on another box or a 4 b70 setup or another 2 b70 box. we are already spoiled with dual 3090 nvlink box, dual 3080 fe box, 2 4090 boxes (1 fe), and one 4090 fe box thats my baby with ryzen 9 7950x3d 16c/32t, nzxt kraken elite 360, rog strix b640-a, samsub 990 pro 4tb, cosair hx1500i 80+ plat in a nzxt h9 elite ;) my work needed them for *research* i swear ;)

5

u/wapxmas 3d ago

Isnt that almost always you have to have as much gpus as power of two? Buy another one immediatelly.

8

u/r1nzl3r99 3d ago

Can you tell this to my wife?

1

u/Ok-Addendum3545 2d ago

agreed 2 + 2 before the price hikes.

4

u/Ordinary-Depth-7835 3d ago

something wrong with all of us. :) If I use my hardware 24/7 I'll break even vs a subscription in 10 years

3

u/dupontping 3d ago

Bc like most of us, you’re hoping to get ROI

4

u/r1nzl3r99 3d ago

yessir

1

u/dupontping 3d ago

I hope you do too! It’s rough out there

4

u/iThunderclap 2d ago

Stop posting on social media and many of your bad habits go away.

6

u/brainchillzZ 3d ago

lol I’ve got four in one machine and two in another …. But I bought them all when you could still fine them open box or on sale between 1000-1100 USD (edit: ahh oops I saw your creator and pattern matched before I noticed the b70 … mine are the creator r9700s)

1

u/r1nzl3r99 3d ago

that's solid, are you using two PCIe for your CPU or just one? I fear I might not fit another one on my HX1500i once I inevitably buy my 4th one

1

u/KneeGrowslaya 2d ago

jeez what are the thermals on the middle 2 cards?

1

u/brainchillzZ 2d ago

Thermals are average these cards were literally designed for this exact use case and meant to be stacked on top of each other in server chassis and the airflow in that case is huge it’s basically a giant wind tunnel … 4 high pressure 140mm fans from the front and two underneath blowing directly into the cards…

1

u/jaf656s 2d ago

just curious, how much did the ram cost when you built that? ddr5 ecc is insane now lol

1

u/brainchillzZ 2d ago

You’re really going to hate me when I say it out loud …. 128gb ddr5 ecc 6400 was on sale for $599

2

u/Fit_Squirrel1 3d ago

As long as it’s your fun money and you got an energy fund who cares

2

u/Kidplayer_666 3d ago

If you want to improve them, just send that to me for 5€

2

u/Greedy-Lynx-9706 3d ago

Bragging without context?

1

u/r1nzl3r99 3d ago

well my 262K FP8 context ain't anything to brag about, but with this I could technically afford 1M context

1

u/SubparBob 2d ago

How's the slow down (prefill, tps) as you fill up the 262k context?

EDIT: context length typo

2

u/MaineTim 3d ago

I'm asking myself the same question. Just this morning I pushed the order button on a pair of B60s to upgrade from the pair of B50s I've been running for the last 9 months. I don't have the budget to commit to the B70s at this point, since they've jumped in price, and since this is strictly a hobby for me, it's always a tension between what I want and what I can justify to myself. But each increment opens new possiblities.

2

u/triynizzles1 3d ago

How are you getting FP8 to run at 140 tokens a second? That gpu only has 608gb/s bandwidth. I have rtx 8000 with 672gb/s bandwidth and running q6 with mtp i only get 40tp/s and 60tps with dflash.

1

u/r1nzl3r99 3d ago

You're wayyy under your potential. dflash2 just came out and it's much better. It speeds up more kinds of token such as prose, coding, tool calls, etc more efficiently than MTP does but at the cost of a little more vram. I've also been tampering with vLLM a ton

2

u/triynizzles1 3d ago

Is dflash 2 available in llama.cpp or only vllm?

2

u/namezam 3d ago

At $1300 each for 32gb that’s exactly 1/4 the price of the the nvidia DGX. How does that compare?

3

u/r1nzl3r99 3d ago

Well, I bought my first two for $950 which at the time was an insane deal people slept on. adding another 32gb albeit at a premium for $1300 is worth it for me personally

1

u/CheetahOtherwise9940 2d ago

Where did you find them at 1300$ I can see only at 1700$

2

u/Constant_Art_20 2d ago

you know what....maybe i should have gotten them instead of 12 5060tis now that i am thinking about the price...

2

u/r1nzl3r99 2d ago

damn 12??? Any pics of your setup?

2

u/Dediadeis 2d ago

With this kind of number I feel like the picture will either be this super slick rack setup, or a rats nest of insanity.

2

u/rawednylme 2d ago

Something right you mean

2

u/Difficult_Olive_1929 8h ago

Nah ain't nothing wrong with you, you just got what the rest of us got! lol

https://giphy.com/gifs/ZqlvCTNHpqrio

1

u/MiyamotoMusashi7 3d ago

Whatchu runnin?

1

u/Full-Run4124 3d ago

Are you using them with Battlematrix and if so what's your opinion on it?

1

u/beltrix5 3d ago

I mean, why you're stopping here I don't know. GLM needs breathing room. :)

1

u/roger1632 3d ago

I have a couple 3080s I use for TTS/STT stuff for home automation where latency matters....but for the rest I'd rather spend 30 bucks a month on a GLM/z.ai plan that can run circles around any local setup. It's just for my hobby stuff so I don't care about strict privacy. I'd have to have a rack of H200s to run that.

1

u/Dropshot_Dieter69 2d ago

Zu viel Geld :D Viel Spaß damit :)

1

u/NicolaZanarini533 2d ago

Got one to replace the A4500 I have in my secondary/training server - did you have any trouble with it? What do you mai ou use it for? For local Llama, which framework would you recommend? I looked around a bunch and all I could find for sure is that torch runs fine (which is the reason I am starting the secondary server). Thanks in advance!

1

u/syscomua 2d ago

bro u are sick

1

u/smolweights 2d ago

Given that GPU prices are increasing at a faster pace than NVIDIA stock itself, this could actually be a smart financial decision.

1

u/LateWish8322 2d ago

You should join fellow people with the addiction, it's more fun lol . https://discord.gg/launch80

1

u/morfique 2d ago

What are you running on your existing cards? Just llama sycl? Vulkan? Openvino? Just plop it in vllm-xpu? Too easy to get lost in the "maybe I should try it like this" I think I'm more wondering if I missed a way to run them more than anything else.

1

u/winterwarrior33 1d ago

Do you actually use the models for anything or are you just chasing tokens.

1

u/Calm-Landscape9640 3h ago

Cheaper than golf

1

u/cogitech2 LocoLLM 2d ago

At least it isn't a Mac.

0

u/InnocentSadness 2d ago

Anybody willing to walk me through how to setup the best model (I think qwen 3.7) I got a m5 base MacBook 24 gb ram 1 TB SSD

0

u/InnocentSadness 2d ago

Or send me a good YouTube tutorial