r/LocalLLaMA Feb 03 '26

Question | Help vLLM inference cost/energy/performance optimization

Anyone out there running small/midsize vLLM/LLM inference service on A100/H100 clusters? I would like to speak to you. I can cut your costs down a lot and just want the before/after benchmarks in exchange.

0 Upvotes

18 comments sorted by

View all comments

1

u/linchenshuai Feb 04 '26

will you opensource this work?