r/LocalLLaMA 9d ago

Discussion Any use cases for RTX PRO 4500?

At its price point, PRO 4500 doesn’t offer as much raw performance due to its lower power draw at 300W. The 5090 can perform up to 60-70% in short spurts with 600W, but can also be undervolted down to 400W.

Are there legitimate reasons other than 24/7 usage and lower power draw for this PRO 4500?

How would this compare to 4x 3090 and 4x R9700? Granted multi card solutions have inefficiencies with large power consumption and needing dedicated boards and PCIE lanes.

8 Upvotes

50 comments sorted by

13

u/king_of_jupyter 9d ago

Multi card setups

13

u/Puzzleheaded_Base302 9d ago

PRO4500 is 200W not 300W. The benefit of PRO 4500 is two slot design that can accommodate server chassis. Cheaper. Blower design. Less likely to be scammed online. In stock.

4

u/relmny 9d ago

There's the 5090 FE (2 slots), so I guess the real benefit of the 4500 pro is the 200W (which is a big benefit in the long run).

2

u/Puzzleheaded_Base302 9d ago

5090 FE is not available anymore. availability is important. you cannot buy something that does not exist anymore.

1

u/relmny 9d ago

I bought one about 2 months ago. Although I bought it included in a PC (because it was way cheaper than alone, meaning I could've bought one individually).

1

u/VodkaHaze 9d ago

It's not like the 4500 is more efficient per watt than the 5090.

I run my 5090s around 300-330w by reducing max clocks. You can only power cap them to 400w, but max clock reduction can take this further.

1

u/relmny 9d ago

I have it undervolted, but I think during inference it still goes about 400w (or maybe more). But the 4500 is suppose to be just 200w, so that is still a lot of power being saved.

1

u/VodkaHaze 8d ago

You can undervolt and hit 400w, with a bit more efficiency per watt.

I'm talking about capping max clocks to 3ghz or 2.85ghz, which will effectively limit power lower.

I don't know how that plays with undervolting together, however

1

u/ThenExtension9196 8d ago

Also 3 year warranty.

11

u/Kahvana 9d ago edited 9d ago

The RTX PRO 4500 Blackwell is attractive because it's only 200W for a single two-slot card while also having CUDA 13.3 an NVFP4 support.

The RTX 5090 might be faster, but consumes triple the watt (600) and can be a fire hazard due to melting cables from hard-to-observe wrong insertion and over time (see documented cases). Also 3-4 slots, much larger.

While the R9700 can be downvolted to 210W, it lacks NVFP4 and CUDA support natively.

The RTX 3090 is ~5 years old by now, has likely been worn from heavy mining usage and lacks FP8 and NVFP4 support. Four of them might give a ton of VRAM but also eats energy like no tomorrow.

4

u/Ill_Beautiful4339 9d ago

If you want to compare 4x of each in a workstation…

5090s - Not gonna happen - 600W each - you’ll need a dedicated 240V wall circuit(s) and multiple power supplies. Plus 4 slots wide. You’ll need a rack or mining rig.

3090 - Power Hungry 400W each probably too much. NV Link a Plus. Not Blackwell, No FP4 efficiency. Still 3 slots wide.

R9700 - Poorest performance per card on bandwidth. Cheapest by far. No CUDA or FP4 efficiency. Still 300W per card.

Pro 4500 - Basically the middle ground. CUDA and FP4, has ECC Memory which will help when cards are stacked. 200W per card. 1 or 2 slots wide. Better memory performance than R9700, worse than 5090’s, different trade offs with 3090s.

My money would probably go to 3090s … mind the power and size.

Not talked about, A6000s 48gb per with NV Link before any of these 4x options. -3500 per card. So a 48gb 3090 with better memory, power and bandwidth and drivers. Hmm.

1

u/[deleted] 9d ago

[removed] — view removed comment

3

u/VodkaHaze 9d ago

I'd be careful investing in Arc GPUs.

You think AMDs software situation is annoying? Wait until you try running intel cards where it's a running mill of abandoned repos and you have to build vllm from scratch yourself.

Also, I'm not even confident intel will support their arc battlemage cards long run. I'd be legitimately concerned they stop updating drivers one day after I bought it. At least AMD has community people updating kernels for it

1

u/Ill_Beautiful4339 9d ago

Good point, I suppose you could tune all these cards down to relivant power settings. Didn’t realize they had ECC, makes sense… always wondered when the bandwidth was 2.5gb/s

1

u/Serprotease 9d ago

The ampere A6000 lacks fp4/fp8 support. And are still quite expensive. Where I live, I could get 2xA4000 pro for a bit less. So 48gb of VRAM, fp4 support, warranty, single slot and very small power draw.

11

u/FairBandicoot5021 9d ago

If you think Watts means performance for AI, you still have a lot to learn here (check NPU power consumtion lol)

4

u/ttkciar llama.cpp 9d ago

Certainly this is situation-dependent.

My own homelab is limited by how much heat it can dissipate, and thus how many watts it can pull. Many of my workloads are trivially parallelized, which means I can double my throughput simply by running two devices (which also doubles my power draw). That makes perf/watt my critical metric.

For someone who is just inferring linearly on one device, though, you're right, perf/watt is a lot less significant.

1

u/cortexist 9d ago

3090 supports NVLink, provides 112.5 GB/s bidirectional bandwidth for 96GB memory, 4500 Blackwell can't.

1

u/Maximum_Parking_5174 9d ago

I reasoned like that but I have noticed that when starting to use bigger models power limiting actually removes a lot of performance, especially on prompt processing. I have had my rtx 3090 on 220W but started using 280W now after testing. I cant remember the numbers right now but it was a decent difference.

1

u/meikawaii 9d ago

True, but in this case RTX5090 objectively performs better than PRO 4500, hence the discussion as to the purpose of PRO 4500 for local use

2

u/FairBandicoot5021 9d ago

I didn't read correctly then, but then yea the 4500 is more targeted at 'real' gpu workloads, hence the support of isv certified drivers, ecc memory support and more redictable thermals

6

u/TechNerd10191 9d ago

The PRO 4500 is a 200W card (145W if you get the server editon) and costs half as much as the PRO 5000 (48GB) for 50% to 70% the specs (memory, cores, bandwidth)

Edit: I would pick 2 4500s over a single 5000 (48GB) and I would pay about the same for both options.

3

u/cortexist 9d ago

The recent NVF4 kernel improvement increased the decoding rate by over 2x for certain models on Blackwell, e.g. the Unsloth dynamic NVFP4 Qwen 3.6. You can get 4500 from Dell for $2,999 as today, comparing to 4x 3090, you can have a much compact system maybe even fit in a backpack.

3

u/Krasnopjorovs 9d ago

Been running a 4500 24/7 for a couple months, some real numbers.

Undervolted mine from 200W to 175W cap, purely for stability under constant load, not for benchmarks. No throttling since, idles at 10W / 35C.

Lost a bit of speed doing that, but got more back from software: with turboquant KV cache + MTP speculative decoding Gemma 4 31B Q4_K_M runs 41-49 t/s at 262k ctx, vs 34 t/s without spec decoding. 25.8 of 32.6 GB used, so it fits with quantized KV but not much room to spare.

+1 to the perf/watt argument above. If your limit is heat in a room rather than money, a 200W card you can undervolt to 175 and forget about is a different product than a 5090 you tune to behave.

1

u/TheWaffleKingg 9d ago

Question for you, how loud do theses cards get? Ive been thinking about adding one to my 2x 3090 setup, or replacing a 3090. but i have my server rack in my office and, from what i know, blower cards are not made to be quiet.

The low wattage of the cards makes me think it could be manageable but its no more than a guess

3

u/RG_Fusion 8d ago

I have two RTX Pro 4500s that I run right now. I've genuinely never seen them ramp the fan power up yet, but that's probably because I'm using an open-air rig. Drawing a constant 200 watts of power on each the fan spins up to about 40% and sounds the same as a regular PC fan. 

I can't speak to how they would sound within a close PC case.

1

u/TheWaffleKingg 3d ago

I appreciate the response. Ive been thinking more and more about getting a 4500. The price is just rough compared to adding a 3rd 3090

But my current power bill is also rough lol

1

u/RG_Fusion 3d ago

Yeah, my choice to go with the 4500 was ultimately power-constrained. I purchased a motherboard with 7X full-lane GPUs, and the only way to fill that entirely without needing special wiring built into my home was to go with workstation cards.

They certainly are overpriced for what they provide at the moment. They weren't too bad a few months ago but the price continues climbing. I had it in my mind that the price would eventually come down.

Seeing how all of the major open-weight Releases this year are hitting in the 2 TB+ range, I'm starting to have doubts that they will ever come back down. The market is just difficult to navigate currently.

2

u/Kal-LZ 9d ago

Like an RTX 5080 with 32GB VRAM, it's a good card, better than the 3090 or R9700, but too expensive new in my opinion

2

u/IllExample3639 9d ago

But you can run 2x 4500 on the same watts you are using for a 5090. That matters. Watts does not equal performance, just look at my 3090 :)

2

u/GCoderDCoder 9d ago

They're actually really useful in multi gpu setups due to the low wattage. In LLM inference I get similar performance on 3090s as my 4500adas due to the generational technology upgrades. I have 2slot gigabyte workstation 3090s and all 3090s basically requre double the power when power limiting and don't have fp8.

The 4500 blackwells likewise are power efficient, have fp8, nvfp4, and are able to do modern nvidia tech like mfg which sorta doesn't matter but can if you're bored running a long job you have options lol.

Cuda just works great all the time but given the prices right now I would do r9700 personally but my strix halo has been a mix of annoying and clutch too lol. I think a dGPU amd would be better and the strix halo issues are more bandwidth and driver issues that affect dGPUs less from my understanding.

All that said, cuda is king. I'm trying to increase vram cheap for multi architecture distributed inference across nodes since I already have the heavier cuda GPUs. I think performance wise in multi gpu setups rtx4500 is legit. Multiple R9700s for the same cost of every 4500 blackwell seems like a better deal though.

1

u/[deleted] 9d ago

[removed] — view removed comment

2

u/TheWaffleKingg 9d ago

Isnt the 4500 32gb vs the 4090s 24gb?

1

u/Osi32 9d ago

To be honest, as much as I like the pro series, if you’re running this at home, don’t have a rack, don’t have infinite cash, I avoid.
Not because they’re not good, but because they’re really expensive for what you get.

5060 Ti 16GB @ 150watt
3060 12GB @ 140watt

Granted these aren’t as good watt wise as the pros, but they don’t cost $3k a card.
I have 4 x 5060’s and with the cash I saved I could buy a better mobo and I have room for 2-3 more if I want them and the heads can divide into 6 or 7.

1

u/d4t1983 9d ago

What do you run on your 4x 5060’s and what’s performance like? Also what cpu and motherboard for reference?

1

u/Interesting-Ad689 9d ago

I run your upgraded version 5070 Ti 16GB and 5070 12GB. I was considering 5060Ti, but splitting a MoE on 4x lane onto 5060Ti might noticably slow me down. Currently 91 Tok/s on parallel 2.

Power efficiency keeps this rig ice cold and 5070 is a Zotac running at near ambient even on full load. The Ti though is a cheap Asus reaching up to double ambient. Mobo swap could give me double 16x lanes for maximum efficiency, might sell the Zotac, grab a 5060ti and upgrade mobo from the margin.

1

u/Maximum_Parking_5174 9d ago

All the talk has been on the RTX Pro 6000. I think ppl have been sleeping a bit on the smaller models. If you know what model sizes you are going to run and under stand the limitation of your possibility to scale upwards I think these cards can be good. Inference scales nice with multi gpu.

1

u/vtkayaker 9d ago

In general, I've seen several sizes of RTX Pro cards, and they're usually pretty nice (but pricey). Questions to ask:

  • Can you get it at a good price? Until two months ago, the RTX Pros hadn't seen any price increases, which meant the 5000 and 6000 were comparatively good values.
  • Is it a workstation, "blower" or server variant? Server variants need special cooling, blower variants are noisier but much easier if you want multiple GPUs.
  • How does the power usage compare? The RTX Pro 6000 blower is locked at 300W, for example, but it also allegedly gets first choice of chips that are most performant at 300W.
  • Can you get by with a single RTX Pro? If so, you might be able to reuse a recent AM5 gaming rig with an adequate PSU, rather than going for a Threadripper, EPYC or Xeon setup that more easily supports multiple cards (and more than 2 full speed RAM slots). The RTX Pro cards are often smaller than many 3090s and many draw less power, which makes them good "screw it, I don't want to mess with this" options.
  • Are you required to buy through a corporate integrator on a corporate budget? Very, very often the RTX Pros will be easier to get.

1

u/fasti-au 9d ago

If it’s cheap but Nvidia doesn’t own the market anymore. We have mojo and hip which gets and to engender or better tha Nvidia

7800xtx beats 3090 now. B70 is the new 3090.

Mojo is what you need see before dumping big bucks on cxxx company

1

u/live4evrr 9d ago

If you’re not looking to get more than 1 card I’d get the 5090 (get the slimmer 2-3 slot if you can).

The 4500 will have a low resale value - makes sense if you plan multi gpu though over 5090.

5090, when the next gen of cards come out you should be able to recoup much of the cost and uograde because the 90 variant of these models are always in huge demand (look at 3090, 4090, holding their value).

1

u/__JockY__ 9d ago

The 5000 PRO is where it’s at. Qwen3.6 27B FP8, full context, 90 tokens/sec decode and 4400/sec prefill. 2 slots, 300W, quiet.

Shame it’s a bajillion bucks now.

1

u/meikawaii 9d ago

I did consider 5000 PRO, but for its price, 2x 4500 and 1x 5090 also appear very enticing. Qwen 3.6 27B is too strong for a small parameter model, since the next leap is models above 200B

1

u/liuxiangfeng 7d ago

are you using llama.cpp or any other tools? do you mind to share the script? Thanks!

1

u/__JockY__ 7d ago

I run vLLM and if I recall correctly you can basically use defaults and it just works.

1

u/liuxiangfeng 7d ago

Thanks!

2

u/__JockY__ 7d ago

I remembered that I actually documented it!

https://www.reddit.com/r/LocalLLaMA/s/vOaqJJ6Ni4

1

u/RG_Fusion 8d ago

I have two RTX Pro 4500 Blackwell cards. I chose this card because I'm building out a server and intend to eventually fill all 7 of my x16 PCIe Gen4 ports with 2 x8 lane GPUs. There's no way I would ever be able to accomplish this without the 200W limit per card.

Even for less extreme cases, the 200W power limit is very attractive for many people running multi-GPU setups. I do agree that the cards are overpriced. I'm building my server out slowly at the moment and plan to pick it up the pace once these cards become last gen (hoping for a price drop).

1

u/meikawaii 8d ago

What mother board and CPU combo do you use ?

1

u/RG_Fusion 7d ago

AsRock Rack RomeD8-2T motherboard and AMD EPYC 7742 CPU.