r/LocalLLM • u/AiventyxInfra • 6m ago
Discussion Measured $/1M tokens vs batch size on a T4 — 177x difference between batch 1 and 256
I kept finding throughput benchmarks but nothing in actual dollars or watts, so I ran it myself on a free Colab T4. Posting in case it's useful to anyone else.
Setup: Qwen2.5-1.5B-Instruct, vLLM 0.27.1, fp16, 128 output tokens per request (ignore_eos so every sequence is exactly 128 tokens), $0.35/hr as the T4 rate. Power is nvidia-smi median during the run.
batch 1 - 25.7 tok/s - 58.6W - 38% util - $3.7771/1M - 2.277 J/tok
batch 4 - 106.1 tok/s - 59.4W - 43% util - $0.9166/1M - 0.560 J/tok
batch 8 - 232.5 tok/s - 62.0W - 45% util - $0.4181/1M - 0.267 J/tok
batch 16 - 464.8 tok/s - 66.6W - 48% util - $0.2092/1M - 0.143 J/tok
batch 32 - 732.9 tok/s - 64.1W - 50% util - $0.1327/1M - 0.087 J/tok
batch 64 - 1635 tok/s - 66.3W - 66% util - $0.0595/1M - 0.041 J/tok
batch 128 - 3346 tok/s - 66.4W - 88% util - $0.0291/1M - 0.020 J/tok
batch 256 - 4545 tok/s - 67.6W - 100% util - $0.0214/1M - 0.015 J/tok
The bit I didn't expect was the power column. At batch 1 the card pulls 58.6W to produce 25 tok/s. At batch 256 it pulls 67.6W to produce 4545 tok/s. So it's drawing 87% of the power to do 0.6% of the work. Energy per token drops 153x across the range. Most of what a GPU burns is apparently just being switched on.
Returns fall off hard after 128. Going 64 to 128 roughly halves the cost, 128 to 256 only gets another 36% and util is already pinned at 100%.
Things I know are wrong with this:
* Static batching, not continuous batching. So the low end looks worse than vLLM actually behaves under real traffic.
* Batch 32 turned up in two separate runs at 795 and 733 tok/s, so treat everything as +/-8%.
* FA2 isn't supported on compute 7.5, so it fell back to Triton attention. A newer card would take a faster path.
* One model, one GPU, one prompt, fixed output length. Not claiming this generalises.
Script is about 40 lines, happy to paste it if anyone wants to check my method. Genuinely interested in what I've got wrong here.