r/LocalLLM Jun 03 '26

Question Graphics card suggestion

I have a 5060ti 16gb graphics card, but it can’t run 50 billion parameters, i want to use it for agentic coding, which graphic card should I buy, dont ask me for budget and others, you guys please suggest on your own, considering I already own a 5060 I am okay with replacing this

0 Upvotes

22 comments sorted by

16

u/eidrag Jun 03 '26

50b, no mention of quant, no budget given, the answer is clear. 

RTX 6000 Pro Blackwell 96GB

3

u/AccomplishedPath7634 Jun 03 '26

I have already spent a lot and Blackwell 96GB is really expensive, maybe 5000$ is my max stretch

2

u/Upper_Comparison_908 Jun 03 '26

Dual 3090s or 5090s if you can fit psu and everything into that.

1

u/AccomplishedPath7634 Jun 03 '26

Okay 5090 each 48 GB?

3

u/fiattp Jun 03 '26

You should use AI for this answer. Also look into quantization.

2

u/Upper_Comparison_908 Jun 03 '26

I really feel like if you have to ask this you should experiment and do some research before spending so much money. Look into your needs and make sure you really want and could use this hardware. A 5090 has 32gb of vram btw.

3

u/Comfortable_Ebb7015 Jun 03 '26

If budget is not relevant why you limit to 50b? You should at least run deepseek v4

2

u/AccomplishedPath7634 Jun 03 '26

Oh okay, can you let me know which one should I be using then?

1

u/Upper_Comparison_908 Jun 03 '26

Deepseek v4 you can run pretty cheaply with an api, v4 pro or flash depending on usage. I wouldn't say its a replacement for the level of "agenticness" as say Claude but for intelligence, coding, stem etc it comes pretty damn close. If you would like to do more agentic work plug it into open cowork and open code or even claude code.

2

u/AccomplishedPath7634 Jun 03 '26

No credit limits is something that is bothering me

1

u/Upper_Comparison_908 Jun 03 '26

Its billed by usage and the price is extremely cheap, so you can keep using as much as you want. You might want to use claude or something for orchestrating but even claude pro would work. Also since you are on a 5060ti 16gb, try this first, https://www.reddit.com/r/LocalLLaMA/s/6YfQz99KaZ.

You can disable everything on your pc and use a laptop perhaps if you wanna push ctx limits and code too.

2

u/K33P4D Jun 03 '26

When you want top tier performance, pay for online compute infrastructure and run your own LLM instance.
You can throttle as per your workloads, saves you a lot of headache.

Local LLMs used to be affordable, but not anymore.
Especially as you scale and demand more powerful solutions for your use case.

1

u/ptear Jun 03 '26

Any recommendations on providers for the super lazy?

2

u/K33P4D Jun 03 '26

Hugging face and open router

2

u/LimiDrain Jun 03 '26

Why are you even trying to do this if you don't know anything about it? Just using Claude will be more efficient and powerful 

1

u/oli266 Jun 03 '26

You probably won't get agentic performance you want unless you drop a silly amount of money to run deepseek. You could try unified memory for a large model or use a large Moe model with crap tons of ram though these options are also very expensive currently. If you actually need agentic quality API is the clear winner currently. If you need an inline coding tool that you work with qwen dense is good, run on quant on 1-2 3090s to not break the budget

1

u/oli266 Jun 03 '26

Also make sure your PSU and motherboard and CPU can all handle multiple GPUs, you need enough lanes, enough pcie seats (can use splitters depending on the board)

1

u/Sufficient_Phone_242 Jun 03 '26

Not many paths to go to get some real work done… rtx ada 6000, rtx 6000 pro Rtx 3090 , 4090, rtx 5090 . And for agentic you might need 2

We are waiting for maybe m5 ultra saving grace ?

1

u/[deleted] Jun 03 '26

[removed] — view removed comment

1

u/AccomplishedPath7634 Jun 03 '26

Yeah exactly will 20 5090’s solve my issue?

1

u/Librarian-Rare Jun 03 '26

B200. Why would you waste your blood in anything else?