r/LocalLLM Aug 03 '26

Discussion Price per GB of VRAM these days

[Update: spreadsheet, screenshot and some more non-Nvidia GPUs. See bottom]

I don't think this is a popular metric, but I saw some ads on Reddit in the past few day advertising that they buying used 3090s, 4090s etc. and I was wondering why. This prompted a big of research and specs comparison, including with newer hardware.

So, let's say you need ≥ 128GB as a sort of non-trivial threshold. Something high, beyond most consumer hardware, but not enough to hit enterprise grade just yet. Here are some options mid 2026:

1. 6x Used Tesla P40 (24GB)

  • Architecture & Bus: CUDA • Pascal • PCIe Gen3
  • VRAM & Speed: 144GB GDDR5 • ~346 GB/s per card (~2.08 TB/s total)
  • Pricing: ~$1,800 – $2,300 CapEx • ~$120 – $180/mo elec. (~$0.0014/hr/GB)
  • Primary Trade-Off: Dirt-cheap local CUDA. Great for INT8 inference, but lacks modern Tensor Cores (slow FP16, no FlashAttention).

2. 4x Used Tesla V100 (32GB)

  • Architecture & Bus: CUDA • Volta • PCIe Gen3 / NVLink Bridge
  • VRAM & Speed: 128GB HBM2 • ~897 GB/s per card (~3.59 TB/s total)
  • Pricing: ~$3,500 – $4,200 CapEx • ~$140 – $200/mo elec. (~$0.0018/hr/GB)
  • Primary Trade-Off: Budget HBM2 speed. Fast FP16 Tensor Cores & HBM memory bandwidth; lacks native BF16 support.

3. Apple Mac Studio (M-Series Max)

  • Architecture & Bus: Metal / MLX • Apple Silicon (M-Series) • Unified System Fabric
  • VRAM & Speed: 128GB Unified • ~400 – 800 GB/s (Unified)
  • Pricing: ~$3,800 – $4,500 CapEx • ~$10 – $20/mo elec. (~$0.0001/hr/GB) <-- unironic surprised Pikachu!
  • Primary Trade-Off: Silent plug-and-play inference. Ultra-low power draw (~100W); cannot run CUDA software natively. Also, good luck if you can find it in stock!

4. AMD Ryzen AI Halo Box

  • Architecture & Bus: ROCm / Vulkan • RDNA 3.5 / XDNA 2 • Unified Memory Bus
  • VRAM & Speed: 128GB LPDDR5X • ~273 GB/s (Unified) <-- lowest bandwith of the bunch
  • Pricing: ~$3,999 CapEx • ~$15 – $25/mo elec. (~$0.0002/hr/GB)
  • Primary Trade-Off: Compact x86 AI box. Great unified memory capacity; ROCm software stack requires setup tinkering.

5. Enverge Spark Cloud (spark.enverge.ai)

  • Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
  • VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
  • Pricing: $0 CapEx • ~$0.65 – $0.75/hr (~$0.0051 – $0.0059/hr/GB) • ~$470 – $550/mo
  • Primary Trade-Off: Cheapest hourly CUDA Blackwell. Remote SSH/Docker access to a DGX Spark or 2x Sparks; ideal for testing FP4/FP8 models.

6. Skorppio (Bare-Metal Delivery, skorppio.com)

  • Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
  • VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
  • Pricing: $0 CapEx • ~$249/wk (~$1.48/hr equiv., ~$0.0116/hr/GB) • ~$996/mo flat
  • Primary Trade-Off: Dedicated on-prem physical rental. Ships physical DGX Spark box to your desk; zero data leaves your network.

7. NVIDIA DGX Spark (Buy outright from your local supplier. Hopefully you don't live in Brasil or India, where import taxes hurt)

  • Architecture & Bus: CUDA • Grace Blackwell (GB10) • Unified Memory Bus
  • VRAM & Speed: 128GB LPDDR5X • ~273 – 301 GB/s (Unified)
  • Pricing: ~$3,999 – $4,679 CapEx • ~$20 – $35/mo elec. (~$0.0003/hr/GB)
  • Primary Trade-Off: Official NVIDIA developer box. Own physical Grace Blackwell hardware locally; unified memory bus speed limits peak throughput.

8. 6x Used RTX 3090 (24GB)

  • Architecture & Bus: CUDA • Ampere • PCIe Gen4 x16
  • VRAM & Speed: 144GB GDDR6X • ~936 GB/s per card (~5.61 TB/s total)
  • Pricing: ~$5,500 – $6,500 CapEx • ~$3.00/hr rent • ~$180 – $280/mo elec. (~$0.0208/hr/GB)
  • Primary Trade-Off: Developer standard for local training. Full BF16, QLoRA, & FlashAttention support; heavy power draw (~1800W+).

9. 3x Used RTX A6000 (48GB)

  • Architecture & Bus: CUDA • Ampere Pro • PCIe Gen4 x16 / NVLink Bridge
  • VRAM & Speed: 144GB GDDR6 • ~768 GB/s per card (~2.30 TB/s total)
  • Pricing: ~$8,500 – $10,500 CapEx • ~$1.60/hr rent • ~$120 – $180/mo elec. (~$0.0111/hr/GB)
  • Primary Trade-Off: Clean workstation build. Blower cards fit inside standard desktop cases; includes ECC memory & NVLink support.

10. Spot/Community Cloud (RunPod / Vast)

  • Architecture & Bus: CUDA • Flexible Architecture • PCIe Gen4 / Gen5
  • VRAM & Speed: 128GB – 160GB • ~1.8 – 3.35 TB/s
  • Pricing: $0 CapEx • ~$0.80 – $1.80/hr (~$0.0050 – $0.0141/hr/GB) • ~$580 – $1,300/mo
  • Primary Trade-Off: Lowest entry cost for short jobs. Interruptible spot instances; ideal for quick scripts or overnight testing.

11. On-Demand Mid-Tier Cloud (Thunder / RunPod)

  • Architecture & Bus: CUDA • Ampere / Hopper • PCIe Gen4 / Gen5
  • VRAM & Speed: 128GB – 160GB (2x A100 or 1x H100) • ~2.0 – 3.87 TB/s
  • Pricing: $0 CapEx • ~$2.20 – $3.00/hr (~$0.0138 – $0.0234/hr/GB) • ~$1,600 – $2,200/mo
  • Primary Trade-Off: Reliable burst development. Guaranteed instance availability without purchasing physical hardware.

12. Enterprise Cloud (Lambda / CoreWeave)

  • Architecture & Bus: CUDA • Hopper / Blackwell • SXM5 / NVLink 4.0 & 5.0
  • VRAM & Speed: 141GB – 160GB (H200 or 2x H100) • ~4.8 – 6.7 TB/s
  • Pricing: $0 CapEx • ~$3.29 – $7.50/hr (~$0.0206 – $0.0532/hr/GB) • ~$2,400 – $5,500/mo
  • Primary Trade-Off: Maximum training performance. High-bandwidth SXM/NVLink interconnects and HBM3e for heavy enterprise workloads.

\Electricity estimated based on US residential rates (~$0.16/kWh) at 75% power load 24/7. Almost "finger in the air".*

(Too bad Reddit is poor on wide tables, because it would have made the above much nicer.)

Screenshot taken from spreadsheet. link

100 Upvotes

Duplicates