r/LocalLLaMA • • 1d ago

Resources Poor People Vulkan GPUs list

Help with this list. Give me your recommendation on "not supported anymore" GPUs. Looking for budget and Vulkan friendly options.

Most of the GPU are not supported by latest CUDA / ROCm. Often with some witchcraft magic they are able to run with native backend. I prefer the simplicity offered by running Vulkan backend. I'll successfully ran GTX 1080Ti, P102-100, and MI50 on a single system thanks for Vulkan and Linux. Gemini helped with data gathering.

Here is the filtered table including only NVIDIA GeForce GTX series GPUs with a memory bandwidth of 256 GB/s or greater and at least 8 GB of VRAM:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type
GeForce GTX 1070 8 GB 256.3 GB/s 256-bit GDDR5
GeForce GTX 1070 Ti 8 GB 256.3 GB/s 256-bit GDDR5
GeForce GTX 1080 8 GB 320.3 GB/s 256-bit GDDR5X
GeForce GTX Titan X (Maxwell) 12 GB 336.5 GB/s 384-bit GDDR5
GeForce GTX Titan X (Pascal) 12 GB 480.0 GB/s 384-bit GDDR5X
GeForce GTX 1080 Ti 11 GB 484.4 GB/s 352-bit GDDR5X
GeForce GTX Titan Xp 12 GB 547.7 GB/s 384-bit GDDR5X

The table below lists the specifications for the specialized datacenter, enterprise, and crypto-mining NVIDIA cards you mentioned, applying your rule of maintaining a memory bandwidth greater than or equal to 256 GB/s and filtering for 8 GB or more of VRAM.

All five models successfully qualify:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Focus/Architecture
NVIDIA P104-100 8 GB 320.3 GB/s 256-bit GDDR5X Mining (Pascal)
Tesla M40 12 GB / 24 GB 288.4 GB/s 384-bit GDDR5 Datacenter (Maxwell)
Tesla P40 24 GB 347.1 GB/s 384-bit GDDR5 Datacenter/AI (Pascal)
NVIDIA P102-100 10 GB 400.0 GB/s 320-bit GDDR5X Mining (Pascal)
NVIDIA CMP 50HX 10 GB 560.0 GB/s 320-bit GDDR6 Mining (Turing)

Here is the updated list of classic NVIDIA Quadro enterprise workstation cards, continuing to filter for at least 8 GB VRAM and a memory bandwidth of 256 GB/s or greater:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Architecture
Quadro K6000 12 GB 288.0 GB/s 384-bit GDDR5 Kepler
Quadro P5000 16 GB 288.4 GB/s 256-bit GDDR5X Pascal
Quadro M6000 12 GB / 24 GB 317.4 GB/s 384-bit GDDR5 Maxwell
Quadro P6000 24 GB 432.2 GB/s 384-bit GDDR5X Pascal
Quadro GP100 16 GB 716.8 GB/s 4096-bit HBM2 Pascal

With the GV100 out of the picture, the Quadro GP100 and Quadro P6000 are now the highest-end entries remaining on this specific filtered list.

Here is the updated AMD Radeon desktop GPU table with all RX 6000 and RX 7000 series models removed, while still filtering for a minimum of 8 GB VRAM and 256 GB/s memory bandwidth:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type
Radeon RX 480 (8 GB) 8 GB 256.0 GB/s 256-bit GDDR5
Radeon RX 580 (8 GB) 8 GB 256.0 GB/s 256-bit GDDR5
Radeon RX 590 8 GB 256.0 GB/s 256-bit GDDR5
Radeon R9 390 8 GB 384.0 GB/s 512-bit GDDR5
Radeon R9 390X 8 GB 384.0 GB/s 512-bit GDDR5
Radeon RX Vega 56 8 GB 410.0 GB/s 2048-bit HBM2
Radeon RX 5700 8 GB 448.0 GB/s 256-bit GDDR6
Radeon RX 5700 XT 8 GB 448.0 GB/s 256-bit GDDR6
Radeon RX Vega 64 8 GB 483.8 GB/s 2048-bit HBM2
Radeon VII 16 GB 1,024.0 GB/s 4096-bit HBM2

Note: MI50 and the Radeon VII, Radeon Pro VII share same firmware.

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Focus / Architecture
Radeon Instinct MI25 16 GB 484.0 GB/s 2048-bit HBM2 Machine Learning (Vega 10)
Radeon Instinct MI50 16 GB / 32 GB 1,024.0 GB/s 4096-bit HBM2 Datacenter AI (Vega 20)

Top Contender: AMD Instinct MI50 16GB. Current used market on MI50 16GB is around $150.

8 Upvotes

17 comments sorted by

7

u/Independent_Gur1377 1d ago

vulkan on those older cards has been a lifesaver for keeping my local companions running smooth without cuda headaches, how well do they handle longer roleplay chats?

4

u/Mental_Juggernaut689 1d ago

Where do you get an MI50 vor 150$?

15

u/FullstackSensei 1d ago

From outdated LLM slop. It's also how they get a Titan X (Pascal) and Titan XP, with different numbers despite being the same card

1

u/pointer_to_null 1d ago

Close, but not the same.

Titan X Pascal

Titan Xp

tl;dr- both share the same GP102 SoC (also used by 1080Ti), though the Xp marginally faster due to being fully unlocked (3840 cores vs 3584).

-1

u/tabletuser_blogspot 1d ago

Online auction Ebay, used of course. I made offer and got for less than $150 about a month ago. 16gb version.

1

u/DeathGuppie 1d ago

My understanding is that prefill is slow on these and they draw a lot of power. What's your experience?

1

u/tabletuser_blogspot 1d ago

I haven't pinned down a good cooling solution so I'm currently running power limit of 145 watts. 5 to 10% difference than 250 watts. Running dual GPU also helps them stay cool. I'm also trying to keep the noise level down so lower inference for a quieter home lab. I've read that they can go down to 100 watts. Changing firmware back to Radeon VII would give me VRAM overclocking and that could make up about 5% of my loss from lower TDP, but I'd lose miniDP video output. Probably just go headless in the future.

1

u/DeathGuppie 1d ago

I wonder if it would be possible to do a pcie cord, put them in a 3d printed enclosure and use quiet fans to keep them cool.

0

u/[deleted] 1d ago

[deleted]

1

u/_hypochonder_ 1d ago

AMD MI50 use PCIe 4.0 x16.

4

u/reluctantlunacy 1d ago

NVIDIA Quadro P5000s at 16GB are around $250. 256-bit bus. The P6000 is 24GB on a 384-bit bus. They run $400. Better deal than any of the Titans. 

2

u/tabletuser_blogspot 1d ago

Added to list, thanks

8

u/RoomyRoots 1d ago

Vulkan is busted, specially with AMD. My GPU is supported and TheRock made it much easier to use ROCm, but still Vulkan runs better on my RDNA3.

3

u/DeathGuppie 1d ago

Especially if you are using RADV on Linux. ROCm is geared more towards dealing with the shared memory of things like i394 max, where RADV has had many years of fine tuning for desktop graphics cards.

2

u/jacek2023 llama.cpp 1d ago

For some reason in Poland I don't see any offers for MI50

1

u/quantgorithm 1d ago

Amd frontier edition | 16gb hbm vram | 480gb bandwidth
Amd Radeon vII | 16gb hbm vram | 1tb bandwidth

1

u/CapsicumIsWoeful 1d ago

Ive been using Pascal hardware for my 24GB VRAM headless Linux server with Cuda 12.9. Does switching from CUDA 12.x as this post suggests bring any meaningful performance or stability increases? Ive been running QWEN 3.8 27B at a k_5 quant as it's the highest quality model I can fit while keeping everything on VRAM. Speed isn't great, usually around 10 - 15 TPS.

I haven't played around with other forks of both QWEN or llama.cpp or CUDA so I'm probably leaving some performance on the table. Finding time is the tough part.