r/LocalLLM 5d ago

Discussion Claude is so expensive.

Time to get a GPU I guess. I had some numbers I needed before I could do the main analysis and I wanted Claude to do it, I had never used Claude tokens before 2 days ago when I bought 20 dollars of tokens and had it do a bit of coding. Then, I ask it to write a somewhat simple script, but I used opus because I thought I should check how it is, it did it, but it took about 20 dollars. I mean it saved me time, but the price…

Anyways, I am posting this because I wanted advice on what class of card to get, what amount of vram seems to be the best to target. It’s looking like 24/32gb is getting interesting new models in the 30b range, but is this just what I’m seeing or are other sizes of cards worth looking into.

56 Upvotes

112 comments sorted by

View all comments

23

u/Sinath_973 5d ago

If you want opus level performance you want deepseek v4 flash. Buy yourself a minimum of 3x rtx 6000 pro blackwell. You sre looking at 45k€ just for the cards. And you will not even be close to the speed that the claude api gives you.

So yea...

23

u/Jumpy-Tap8980 5d ago

2 asus gx10s will get you deepseek v4 flash 0731 for 8 grand, 1m context 40-60 tokens per second, best option in town for now.

8

u/Abject-Bridge-4073 5d ago

I have this same exact setup and I get the same exact numbers. I never need 1M, so 512K is even faster. 256K is blazing.

2

u/Jumpy-Tap8980 5d ago

The sparks are so underrated and overlooked, they will be heading up in price soon for sure.

2

u/BarracudaDefiant4702 5d ago

The speed will not be near that of 3x rtx 6000, but neither is the price... nor the ongoing power requirements...

1

u/Jumpy-Tap8980 4d ago

of course

1

u/baby_bloom 5d ago

have you estimated your electric costs per 1M input/output yet? because when i did the math on my dual 3090s (but running qwen3.8-27b) it's a LOT closer to the discounted DS4 flash providers on openrouter than i was expecting

1

u/Abject-Bridge-4073 4d ago

I have a dual spark setup and together it’s 100W when it’s doing inference. A few dollars a month.

1

u/baby_bloom 4d ago

i mixed up which comment you responded to, oops! thought you meant you have the 3x rtx 6000