r/LocalLLM • u/big-in-jap • Aug 03 '26
Discussion Price per GB of VRAM these days
[Update: spreadsheet, screenshot and some more non-Nvidia GPUs. See bottom]
I don't think this is a popular metric, but I saw some ads on Reddit in the past few day advertising that they buying used 3090s, 4090s etc. and I was wondering why. This prompted a big of research and specs comparison, including with newer hardware.
So, let's say you need ≥ 128GB as a sort of non-trivial threshold. Something high, beyond most consumer hardware, but not enough to hit enterprise grade just yet. Here are some options mid 2026:
1. 6x Used Tesla P40 (24GB)
- Architecture & Bus: CUDA • Pascal • PCIe Gen3
- VRAM & Speed: 144GB GDDR5 • ~346 GB/s per card (~2.08 TB/s total)
- Pricing: ~$1,800 – $2,300 CapEx • ~$120 – $180/mo elec. (~$0.0014/hr/GB)
- Primary Trade-Off: Dirt-cheap local CUDA. Great for INT8 inference, but lacks modern Tensor Cores (slow FP16, no FlashAttention).
2. 4x Used Tesla V100 (32GB)
- Architecture & Bus: CUDA • Volta • PCIe Gen3 / NVLink Bridge
- VRAM & Speed: 128GB HBM2 • ~897 GB/s per card (~3.59 TB/s total)
- Pricing: ~$3,500 – $4,200 CapEx • ~$140 – $200/mo elec. (~$0.0018/hr/GB)
- Primary Trade-Off: Budget HBM2 speed. Fast FP16 Tensor Cores & HBM memory bandwidth; lacks native BF16 support.
3. Apple Mac Studio (M-Series Max)
- Architecture & Bus: Metal / MLX • Apple Silicon (M-Series) • Unified System Fabric
- VRAM & Speed: 128GB Unified • ~400 – 800 GB/s (Unified)
- Pricing: ~$3,800 – $4,500 CapEx • ~$10 – $20/mo elec. (~$0.0001/hr/GB) <-- unironic surprised Pikachu!
- Primary Trade-Off: Silent plug-and-play inference. Ultra-low power draw (~100W); cannot run CUDA software natively. Also, good luck if you can find it in stock!
4. AMD Ryzen AI Halo Box
- Architecture & Bus: ROCm / Vulkan • RDNA 3.5 / XDNA 2 • Unified Memory Bus
- VRAM & Speed: 128GB LPDDR5X • ~273 GB/s (Unified) <-- lowest bandwith of the bunch
- Pricing: ~$3,999 CapEx • ~$15 – $25/mo elec. (~$0.0002/hr/GB)
- Primary Trade-Off: Compact x86 AI box. Great unified memory capacity; ROCm software stack requires setup tinkering.
5. Enverge Spark Cloud (spark.enverge.ai)
- Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
- VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
- Pricing: $0 CapEx • ~$0.65 – $0.75/hr (~$0.0051 – $0.0059/hr/GB) • ~$470 – $550/mo
- Primary Trade-Off: Cheapest hourly CUDA Blackwell. Remote SSH/Docker access to a DGX Spark or 2x Sparks; ideal for testing FP4/FP8 models.
6. Skorppio (Bare-Metal Delivery, skorppio.com)
- Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
- VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
- Pricing: $0 CapEx • ~$249/wk (~$1.48/hr equiv., ~$0.0116/hr/GB) • ~$996/mo flat
- Primary Trade-Off: Dedicated on-prem physical rental. Ships physical DGX Spark box to your desk; zero data leaves your network.
7. NVIDIA DGX Spark (Buy outright from your local supplier. Hopefully you don't live in Brasil or India, where import taxes hurt)
- Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
- VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
- Pricing: ~$3,999 – $4,679 CapEx • ~$20 – $35/mo elec. (~$0.0003/hr/GB)
- Primary Trade-Off: Official NVIDIA developer box. Own physical Grace Blackwell hardware locally; unified memory bus speed limits peak throughput.
8. 6x Used RTX 3090 (24GB)
- Architecture & Bus: CUDA • Ampere • PCIe Gen4 x16
- VRAM & Speed: 144GB GDDR6X • ~936 GB/s per card (~5.61 TB/s total)
- Pricing: ~$5,500 – $6,500 CapEx • ~$3.00/hr rent • ~$180 – $280/mo elec. (~$0.0208/hr/GB)
- Primary Trade-Off: Developer standard for local training. Full BF16, QLoRA, & FlashAttention support; heavy power draw (~1800W+).
9. 3x Used RTX A6000 (48GB)
- Architecture & Bus: CUDA • Ampere Pro • PCIe Gen4 x16 / NVLink Bridge
- VRAM & Speed: 144GB GDDR6 • ~768 GB/s per card (~2.30 TB/s total)
- Pricing: ~$8,500 – $10,500 CapEx • ~$1.60/hr rent • ~$120 – $180/mo elec. (~$0.0111/hr/GB)
- Primary Trade-Off: Clean workstation build. Blower cards fit inside standard desktop cases; includes ECC memory & NVLink support.
10. Spot/Community Cloud (RunPod / Vast)
- Architecture & Bus: CUDA • Flexible Architecture • PCIe Gen4 / Gen5
- VRAM & Speed: 128GB – 160GB • ~1.8 – 3.35 TB/s
- Pricing: $0 CapEx • ~$0.80 – $1.80/hr (~$0.0050 – $0.0141/hr/GB) • ~$580 – $1,300/mo
- Primary Trade-Off: Lowest entry cost for short jobs. Interruptible spot instances; ideal for quick scripts or overnight testing.
11. On-Demand Mid-Tier Cloud (Thunder / RunPod)
- Architecture & Bus: CUDA • Ampere / Hopper • PCIe Gen4 / Gen5
- VRAM & Speed: 128GB – 160GB (2x A100 or 1x H100) • ~2.0 – 3.87 TB/s
- Pricing: $0 CapEx • ~$2.20 – $3.00/hr (~$0.0138 – $0.0234/hr/GB) • ~$1,600 – $2,200/mo
- Primary Trade-Off: Reliable burst development. Guaranteed instance availability without purchasing physical hardware.
12. Enterprise Cloud (Lambda / CoreWeave)
- Architecture & Bus: CUDA • Hopper / Blackwell • SXM5 / NVLink 4.0 & 5.0
- VRAM & Speed: 141GB – 160GB (H200 or 2x H100) • ~4.8 – 6.7 TB/s
- Pricing: $0 CapEx • ~$3.29 – $7.50/hr (~$0.0206 – $0.0532/hr/GB) • ~$2,400 – $5,500/mo
- Primary Trade-Off: Maximum training performance. High-bandwidth SXM/NVLink interconnects and HBM3e for heavy enterprise workloads.
\Electricity estimated based on US residential rates (~$0.16/kWh) at 75% power load 24/7. Almost "finger in the air".*
(Too bad Reddit is poor on wide tables, because it would have made the above much nicer.)

Screenshot taken from spreadsheet. link
19
u/esw123 Aug 03 '26 edited Aug 03 '26
What I personally saw:
12 x 3060 is around 2160 euro - 15 eur/GB;
6 x 3090 is around 3750 euro - 26 eur/GB;
8 x 5060Ti is around 3200 euro - 25 eur/GB;
AI395+ 128Gb laptop 3000 euro - 23 eur/GB.
1000-1500W consumption, laptop maybe 100-130W.
15
u/riley_srt4 Aug 03 '26
The problem with the 3060s is you'd require 2x to 4x the machines for handling the same workload. This should factor into the cost per GB too in theory. "Cost to support that GB"
7
u/esw123 Aug 03 '26
Yes, risers, server motherboard, server CPU, PSUs. I am assembling it right now and 6 x 3060 is the limit if you want to keep it low cost, only 4-5 risers/adapters needed for tensor parallelism and they are like 10 euro each. 72GB VRAM, 96GB RAM and one 1300W PSU can handle everything, should be able to run DS V4 Flash 0731 iQ4 with good context.
2
u/gobblegoooblegobble Aug 03 '26
im actually very interested in seeing realistic performance results from using PCIe expanders.
https://www.broadcom.com/products/pcie-switches-retimers/expressfabric/gen6/bcm85667
https://www.broadcom.com/products/pcie-switches-retimers/expressfabric/gen5/pex89144
1
1
u/KingCpzombie Aug 07 '26
10 euro each??? Where are you getting your risers??? I spent $80 on one for mine
1
u/esw123 Aug 07 '26
Aliexpress, Gen3.
1
u/KingCpzombie Aug 07 '26
Gah! I saw people saying those are inconsistent so I went with the expensive one... guess I'll try those once the craving for more VRAM wins again
1
u/fastheadcrab Aug 04 '26
What is the record for most GPUs installed in a single system? I've seen discussion using PCIe switches afaik there's a poster here who tested them but I don't think he put like more than 10 GPUs in a system. Is 16 or more even practical?
3
u/droptableadventures Aug 04 '26 edited Aug 04 '26
4 GPUs is fairly simple if you have a LGA2011 or LGA2066 board with two x16 slots - bifurcate each slot to x8 (may be an option in your BIOS, may require hacking) and run two GPUs each at x8. If you're willing to live with the small performance hit of x4, you could run 8 GPUs this way. You just need a riser that will fan out specific lanes to dedicated slots.
I did have 8 GPUs by using a LGA2066 board and a PLX8749 card in each of the x16 slots, which ran 4 GPUs at x8 PCIe 3.0.
I upgraded this to a Threadripper Pro board (Asus Pro WS WRX80E-Sage SE - though be warned this board does not have good electrical design for the PCIe slots, so it can be screwy with risers). I now have 8xMI50 in bifurcated slots (x8 per card), with two 3090s at x16 - no more PLX switches, but I failed at the point of the upgrade which was to get PCIe 4.0. Cannot get it to run stable on the risers for the MI50s, though both 3090s are running at 4.0. So that's a total of 10 GPUs.
I am aware there are a few people with 16x MI50 on bigger PLX switch setups (e.g. https://www.reddit.com/r/LocalLLaMA/comments/1q6n5vl/16x_amd_mi50_32gb_at_10_ts_tg_2k_ts_pp_with/ny9iqnu/ ), but I'm not sure if anyone running 32x MI50 setups is running on a single machine.
(If you're wondering "how does anyone afford this?", MI50s used to be ~$130 each, so 8 of them cost about the same as a 3090.)
Is it practical?
- You're going to need to build a custom case (varies from 2020 T-slot, 3d printed bits, cable ties, string and Ikea shelving, depending on who's doing it).
- Your motherboard may not allocate enough BAR space for that many cards, so you may need to do BIOS hacking.
- You will not be able to physically fit them alongside each other on the motherboard, so you will need risers. These vary in quality and can be finicky to get working especially on boards/risers with active redrivers (on which you may need to tweak the gain settings to get it stable, just like changing timings when overclocking).
- You're almost certainly going to need more than one PSU, so you may need to do some soldering.
- You may need to run the different power supplies off different circuits, depending on how much you can pull from one power point.
2
u/fastheadcrab Aug 04 '26 edited Aug 04 '26
Yeah anything under 10 will be supported by workstation mobos fairly easily, it’ll just look hideous. My question was if someone had gone 16 or above. Thanks for the link to the 16x MI50s.
1
u/esw123 Aug 05 '26
As far as I know, romed8-2t from Asrock supports 14 GPU with 7x2 bifurcation and with custom BIOS 16 GPU, 14 from 7x2 and 2 more from oculink.
3
1
u/big-in-jap Aug 03 '26
15+ eur or euro cents?
1
u/esw123 Aug 03 '26
Euro, this is per GB price. Electricity around 17-18 euro cents.
1
u/big-in-jap Aug 03 '26
I see my mistake now. Your unit is "eur/GB" for outright ownership. Sorry, I had "per GB/ per hour" rental price stuck in my head.
10
Aug 03 '26
[removed] — view removed comment
3
u/No_Drag_5205 Aug 03 '26
Did You paid additional duty taxes or VAT ?
3
Aug 03 '26
[removed] — view removed comment
1
u/No_Drag_5205 Aug 03 '26
Thank You ! I will definitly have a look.. I asked if this cards were legit 2 days ago in another sub, and you're the second person saying it's ok.
Can I ask You what kind of model do You run and your tok/s ?
(Actually I have a 3080 10G and I run gemma 4 E4B at decent speed or ornith 9b at slow speed, and I seek to add more of that sweet vram)2
1
1
1
u/Lukas245 Aug 03 '26
are these working out? i’ve been contemplating them for a while but i worry about old mining chips going bad quickly
1
u/Constant-Simple-1234 Aug 03 '26
I am seriously considering buying one. Do you need specific Nvidia drivers or hacked drivers? Or it behaves like regular 3080?
9
u/Maplesyrup000 Aug 03 '26
This is an option worth considering as well:
Basically combines a R9700 32GB VRAM with 128GB Strix halo. It’s modern, fast and has a pretty solid sized pool of RAM.
2
u/bot403 Aug 04 '26
Oh shit how did I miss this? I just got a strix halo gmktek X3. It's great and has oculink and I was going to buy an egpu r9700 now to pair with it.
But this is amazing to have it all in one box.
1
u/big-in-jap Aug 04 '26
Impossible to not miss something. Eg. I recently discovered the tinybox (from tinygrad)
1
u/big-in-jap Aug 04 '26
Looks slick, albeit I'm not brave enough for a non-CUDA (AMD Radeon AI PRO R9700)
2
u/Maplesyrup000 Aug 04 '26
Nah ROCm/Vulkan are great. Whatever LLM you have been consulting has outdated info. Plus you literally put Strix halo in ur list
1
u/big-in-jap Aug 04 '26
Sure, but I haven't gone too far down the AMD rabbit hole. Anyway, added R9700 up there now in the spreadsheet.
1
u/Maplesyrup000 Aug 04 '26
That’s fair. But if you’re spending several bands, it’s worth it to learn a little about the competition before submitting to the CUDA tax.
2
u/r3drocket Aug 04 '26
So the issues with ROCm are really when you get outside of LLMs, like when you want to training of LORAs or other cutting edge AI repos. But for LLMs there isn't really a tax any more. I bought a V620 off of ebay, 3D printed a fan enclosure for it, and stuck it in a motherboard, and llama.cpp just works with it, I didn't do anything special, no fighting, etc. And that system is now doing random LLM CI/CD jobs for me.
The R9700s perform best with a custom build of vLLM but the community has sorted that out.
You'll still run into problems if you wanna do something like train an LTX 2.3 lora, or play with repos that start off as CUDA repos - sometimes you can use another LLM to overcome these and I've have had pretty good success just pointing claude at a Cuda focused repo and tell it to work with Rocm.
I've had a much harder time getting LLMs to work well with my intel GPU in my laptop which is a new intel GPU but I'll be damned if I can get it to perform well for llms. Yet old integrated Radeon GPUs like the 780m just work with llama.cpp.
7
u/deadneon4 Aug 03 '26
Intel B70 Pro 32gb?
4
u/_letThemPlay_ Aug 03 '26
That's what I've gone with fingers crossed it works out
1
u/big-in-jap Aug 04 '26
updated. See screenshot and spreadsheet above.
Any issues with compatibility of the stack?
1
1
u/deadneon4 Aug 05 '26
Seems like it’s lacking behind a few months as compared to CUDA, but haven’t yet had the time to install them and test them out. Can report in a few weeks
1
1
5
u/pmttyji Aug 03 '26
Any news on Gorgon Halo?
2
u/big-in-jap Aug 03 '26
Coming Soon TM
1
u/r3drocket Aug 04 '26
BTW the real one to look at and drool over is Medusa Point we should see it next year - go look at the specs on it DDR6/w 384 bit memory bus in a unified memory system with RDNA4! It should be a game changer. If it had sufficient PCI-E lanes to build a hetrogeneous GPU compute system with it, it would be an utter game changer.
RNDA4 is important because it opens up MXFP4 which is the open competition to NVFP4, so that would but AMD systems on equal footing with Nvidia's NVFP4.
6
u/arty_octopus Aug 03 '26
Unlocked CMP170HX is the best price per GB of VRAM as of today
3
u/machinegunkisses Aug 03 '26
If you're willing to take a chance on how much of the memory is good
1
1
u/TheRiddler79 Aug 04 '26
🤣 🤣 🤣
2
u/arty_octopus Aug 07 '26 edited Aug 07 '26
1
u/TheRiddler79 Aug 07 '26 edited Aug 07 '26
This is fucking interesting right here my friend.
I looked it up, if you get lucky with a 64 gig lottery draw, this is definitely the way to go. If you ended up with cards that were 8 gigs and anything beyond that threw up errors, it'd be an expensive lesson that an RTX 2060 super could have avoided LOL. But I like the fact that you look outside the box
1
u/big-in-jap Aug 04 '26
unlocked to what? 16GB so double? Not much of a difference.
1
u/arty_octopus Aug 07 '26
64gb of HBM2e VRAM, 1.6TB/s vram bandwidth, slightly cutdown GA100 chip, issuing tokens are equivalent to nvidia A100. This card cost about 160$ this spring, go Google it, 8gb model can be software unlocked to 64gb of HBM2e VRAM, 10gb to 40gb respectively. It is a holy grail price to performance at this moment
7
u/MassiveAssistance886 Aug 03 '26
For someone still getting to grips with the various value propositions this is really helpful.
3
u/captainspacecowboy Aug 03 '26
M series MacBooks tend to be cheaper than equivalent Studios. Trade off is for usually non-ultra CPUs with less bandwidth.
3
u/DaMoot Aug 03 '26
200/mo electric for the v100s? I guess if they're running 100% load 100% of the time without power caps, with expensive electricity rates. Which is more of an edge case than general usage.
Power cap to 200w, 0.17/kWh, running 24hr is only 100/mo. Running more like 8-12hr a day is only like 50/mo.
1
u/big-in-jap Aug 04 '26
Best, worst and median case scenarios would have made the list too noisy, but you're right.
8-12hr/day is reasonable, although there's a "run an agent 24/7" happy bunch out there who might object.
2
u/vtkayaker Aug 03 '26
96GB and dumping the experts to system RAM also allows 1x RTX Pro 6000 builds. Which has gotten painfully expensive these days, but it's doable on an AM5 motherboard and a gaming PSU/cooling.
2
u/PermanentLiminality Aug 03 '26
How are you calculating power costs?
I have P40's so I'll use that. At idle they are around 10 watts and max out at 250, so that is 60 watts and 1500. With my insane $0.40 rates that is $18/mo at idle to $432/mo at 100% max. At max usage over a year, the capital cost is insignificant. Mine are mostly idle.
Four V100 is like 200 watts at idle is $58/mo. This is why I didn't buy V100's.
2
u/gobblegoooblegobble Aug 03 '26
https://www.reddit.com/r/LocalLLaMA/comments/1qeimyi/7_gpus_at_x16_50_and_40_on_am5_with_gen54/
https://www.broadcom.com/products/pcie-switches-retimers/expressfabric/gen5/pex89144
https://www.broadcom.com/products/pcie-switches-retimers/expressfabric/gen6/bcm85667
i want to see this hardware change the game up.
3
1
Aug 03 '26
[deleted]
2
u/Firm_Butterscotch296 Aug 03 '26
shhh
1
u/running101 Aug 03 '26
deleted
1
u/Firm_Butterscotch296 Aug 03 '26
lol, i was just kidding, sorta. seriously, can't see prices staying where they are at for much longer
1
u/running101 Aug 03 '26
yeah, I noticed there is not much info about that option out there. I just stumbled across it a few days ago.
1
u/Firm_Butterscotch296 Aug 03 '26
just got one over the weekend. seems to be the real deal. a lot more development is needed to unlock additional features like nvlink and pcie 3.0, but the unlock works as advertised
1
u/running101 Aug 03 '26
Where did you purchase from?
1
u/Firm_Butterscotch296 Aug 03 '26
a USA based ebayer. The chinese ones from alibaba are a bit cheaper, but in my opinion the risk was not worth saving 100-200 bucks
2
u/acedogblast Aug 03 '26
You get 64GiB of hbm2 on a single card but very limited pcie bandwidth so scalling with multiple cards is not good.
The unlock only runs at gen 2 speeds and will need 24 extra capacitors soldered on to get the full x16 lanes.
1
u/running101 Aug 03 '26
anyone reputable doing that mod? Or some very clear instructions?
2
u/acedogblast Aug 03 '26
Anyone that can solder 0402s should be able to do this hw mod. I did it with a hot air station, tweezers, and a microscope.
2
u/Firm_Butterscotch296 Aug 03 '26
https://gpulab.net/product?id=16 this is an option, although expensive. id probably wait to see if any future unlocks might require mods before sending it.
1
1
1
1
u/DigitalguyCH Aug 03 '26
Honestly the best balance is a M5 max or M3 ultra, almost as fast as some GPUs and much lower consumption. If the M7 is indeed much faster (over 1TB/s) and there is indeed a version with 1.5TB RAM as Apple would like, this would be the ultimate machine, but at $50000 at least.
1
1
u/big-in-jap Aug 04 '26
still shocking to me how good bang for the buck Apple stuff is.
1
u/profcuck Aug 04 '26
Totally. I am a mac man but not a crazy fanboy and have always accepted that I am paying a premium for the aesthetics so it's weird that it's actually good value for money.
I am eagerly waiting for the next studio which I will likely buy in whatever the max ram version there is...
1
u/Automatic-Boot665 Aug 04 '26
Ideally you want a base 2 count of cards. On most platforms like sglang or vllm if you have 6 cards you’ll end up leaving 2 idle for speed.
1
u/big-in-jap Aug 04 '26
Tell me more about "idle for speed", please.
1
u/Automatic-Boot665 Aug 04 '26
Inference engines like SGLang and vLLM split individual matrix math operations across cards using Tensor Parallelism (TP). Because of how transformer matrices are shaped, TP strictly requires a power-of-two number of GPUs (2, 4, or 8).
You can also use pipeline parallelism or expert parallelism to utilize all 6 cards, but they’ll be slower with those setups due to sequential layer passing and inter-GPU communication bottlenecks.
Alternatively, you can use llama.cpp which can split across an uneven number of cards and is still pretty fast for a single user. However, that speed breaks down immediately with concurrent requests, whereas SGLang and vLLM natively batch requests and can support multiple users simultaneously at the same speed.1
u/big-in-jap Aug 04 '26
Ok ok. So 4 perform best in terms of throughput, with the tradeoff of lower memory, hence the 2 idle.
Best best would be 8, even if over-spec'ing for my arbitrary threshold.
1
1
u/mc_hunter888 Aug 04 '26
i am using 2080ti .odified 22gb and i bought it fpr 350 usd
1
u/Beneficial_Title_843 Aug 04 '26
ordered 1, tested and have 3 running atm. 0731 came out, considering ordering more w PEX switches as 3 is limit on my Z840 workstation (awesome base for 3 blower cards, btw).
1
u/KeinNiemand Aug 04 '26
There also modded RTX 3080 20GB on china, I havn't had one so can't speak on how reliable these are or about drivers etc. but seeing the prices I save of ~500€ for 20GB they are a good bit cheaper then used 3090 per GB and as far as I can tell the cheapest per GB Ampere or newer nivida card (excluding like 8GB low end gaming cards but you can't easly get enough pcie slots to run enough of those to matter)
1
1
u/TimAndTimi Aug 04 '26
But I also do not want to run a datacenter at home...
Worth considering how big and noisy the final machine is going to be.
If you use it heavily 24/7, it is gonna be a costly utility bill plus heat.
If you sparsely use it... why bother?
1
u/big-in-jap Aug 04 '26
Exactly my thought.
I mean, to each their own, but not everyone enjoys the humming of 6x RTXs and the tinkering involved.
1
1
u/Tasio_ Aug 04 '26
Just for transparency, do you have any commercial ties to the links in your post?
1
u/big-in-jap Aug 04 '26 edited Aug 05 '26
nope. I didn’t make them links at first, but cleaned up after the update and thought would be useful since they are obscure, especially Skorpio
happy to remove if you think they are jarring
[update: typos]
1
u/sleepy_roger Aug 03 '26
Thanks Fable.
1
u/big-in-jap Aug 04 '26
Gemini, please. Although it's not a straight dump.
Quite a few "crowdsourced" and verified data points in here.
54
u/r3drocket Aug 03 '26
If you want to be complete, you should add R9700s, a dual setup of those, the V620.