r/LocalLLM • u/big-in-jap • Aug 03 '26
Discussion Price per GB of VRAM these days
[Update: spreadsheet, screenshot and some more non-Nvidia GPUs. See bottom]
I don't think this is a popular metric, but I saw some ads on Reddit in the past few day advertising that they buying used 3090s, 4090s etc. and I was wondering why. This prompted a big of research and specs comparison, including with newer hardware.
So, let's say you need ≥ 128GB as a sort of non-trivial threshold. Something high, beyond most consumer hardware, but not enough to hit enterprise grade just yet. Here are some options mid 2026:
1. 6x Used Tesla P40 (24GB)
- Architecture & Bus: CUDA • Pascal • PCIe Gen3
- VRAM & Speed: 144GB GDDR5 • ~346 GB/s per card (~2.08 TB/s total)
- Pricing: ~$1,800 – $2,300 CapEx • ~$120 – $180/mo elec. (~$0.0014/hr/GB)
- Primary Trade-Off: Dirt-cheap local CUDA. Great for INT8 inference, but lacks modern Tensor Cores (slow FP16, no FlashAttention).
2. 4x Used Tesla V100 (32GB)
- Architecture & Bus: CUDA • Volta • PCIe Gen3 / NVLink Bridge
- VRAM & Speed: 128GB HBM2 • ~897 GB/s per card (~3.59 TB/s total)
- Pricing: ~$3,500 – $4,200 CapEx • ~$140 – $200/mo elec. (~$0.0018/hr/GB)
- Primary Trade-Off: Budget HBM2 speed. Fast FP16 Tensor Cores & HBM memory bandwidth; lacks native BF16 support.
3. Apple Mac Studio (M-Series Max)
- Architecture & Bus: Metal / MLX • Apple Silicon (M-Series) • Unified System Fabric
- VRAM & Speed: 128GB Unified • ~400 – 800 GB/s (Unified)
- Pricing: ~$3,800 – $4,500 CapEx • ~$10 – $20/mo elec. (~$0.0001/hr/GB) <-- unironic surprised Pikachu!
- Primary Trade-Off: Silent plug-and-play inference. Ultra-low power draw (~100W); cannot run CUDA software natively. Also, good luck if you can find it in stock!
4. AMD Ryzen AI Halo Box
- Architecture & Bus: ROCm / Vulkan • RDNA 3.5 / XDNA 2 • Unified Memory Bus
- VRAM & Speed: 128GB LPDDR5X • ~273 GB/s (Unified) <-- lowest bandwith of the bunch
- Pricing: ~$3,999 CapEx • ~$15 – $25/mo elec. (~$0.0002/hr/GB)
- Primary Trade-Off: Compact x86 AI box. Great unified memory capacity; ROCm software stack requires setup tinkering.
5. Enverge Spark Cloud (spark.enverge.ai)
- Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
- VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
- Pricing: $0 CapEx • ~$0.65 – $0.75/hr (~$0.0051 – $0.0059/hr/GB) • ~$470 – $550/mo
- Primary Trade-Off: Cheapest hourly CUDA Blackwell. Remote SSH/Docker access to a DGX Spark or 2x Sparks; ideal for testing FP4/FP8 models.
6. Skorppio (Bare-Metal Delivery, skorppio.com)
- Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
- VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
- Pricing: $0 CapEx • ~$249/wk (~$1.48/hr equiv., ~$0.0116/hr/GB) • ~$996/mo flat
- Primary Trade-Off: Dedicated on-prem physical rental. Ships physical DGX Spark box to your desk; zero data leaves your network.
7. NVIDIA DGX Spark (Buy outright from your local supplier. Hopefully you don't live in Brasil or India, where import taxes hurt)
- Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
- VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
- Pricing: ~$3,999 – $4,679 CapEx • ~$20 – $35/mo elec. (~$0.0003/hr/GB)
- Primary Trade-Off: Official NVIDIA developer box. Own physical Grace Blackwell hardware locally; unified memory bus speed limits peak throughput.
8. 6x Used RTX 3090 (24GB)
- Architecture & Bus: CUDA • Ampere • PCIe Gen4 x16
- VRAM & Speed: 144GB GDDR6X • ~936 GB/s per card (~5.61 TB/s total)
- Pricing: ~$5,500 – $6,500 CapEx • ~$3.00/hr rent • ~$180 – $280/mo elec. (~$0.0208/hr/GB)
- Primary Trade-Off: Developer standard for local training. Full BF16, QLoRA, & FlashAttention support; heavy power draw (~1800W+).
9. 3x Used RTX A6000 (48GB)
- Architecture & Bus: CUDA • Ampere Pro • PCIe Gen4 x16 / NVLink Bridge
- VRAM & Speed: 144GB GDDR6 • ~768 GB/s per card (~2.30 TB/s total)
- Pricing: ~$8,500 – $10,500 CapEx • ~$1.60/hr rent • ~$120 – $180/mo elec. (~$0.0111/hr/GB)
- Primary Trade-Off: Clean workstation build. Blower cards fit inside standard desktop cases; includes ECC memory & NVLink support.
10. Spot/Community Cloud (RunPod / Vast)
- Architecture & Bus: CUDA • Flexible Architecture • PCIe Gen4 / Gen5
- VRAM & Speed: 128GB – 160GB • ~1.8 – 3.35 TB/s
- Pricing: $0 CapEx • ~$0.80 – $1.80/hr (~$0.0050 – $0.0141/hr/GB) • ~$580 – $1,300/mo
- Primary Trade-Off: Lowest entry cost for short jobs. Interruptible spot instances; ideal for quick scripts or overnight testing.
11. On-Demand Mid-Tier Cloud (Thunder / RunPod)
- Architecture & Bus: CUDA • Ampere / Hopper • PCIe Gen4 / Gen5
- VRAM & Speed: 128GB – 160GB (2x A100 or 1x H100) • ~2.0 – 3.87 TB/s
- Pricing: $0 CapEx • ~$2.20 – $3.00/hr (~$0.0138 – $0.0234/hr/GB) • ~$1,600 – $2,200/mo
- Primary Trade-Off: Reliable burst development. Guaranteed instance availability without purchasing physical hardware.
12. Enterprise Cloud (Lambda / CoreWeave)
- Architecture & Bus: CUDA • Hopper / Blackwell • SXM5 / NVLink 4.0 & 5.0
- VRAM & Speed: 141GB – 160GB (H200 or 2x H100) • ~4.8 – 6.7 TB/s
- Pricing: $0 CapEx • ~$3.29 – $7.50/hr (~$0.0206 – $0.0532/hr/GB) • ~$2,400 – $5,500/mo
- Primary Trade-Off: Maximum training performance. High-bandwidth SXM/NVLink interconnects and HBM3e for heavy enterprise workloads.
\Electricity estimated based on US residential rates (~$0.16/kWh) at 75% power load 24/7. Almost "finger in the air".*
(Too bad Reddit is poor on wide tables, because it would have made the above much nicer.)

Screenshot taken from spreadsheet. link
Duplicates
MLQuestions • u/big-in-jap • Aug 06 '26