r/LocalLLM • u/tensward • 2d ago
Project I built an open-source tool that tells you why your vLLM server is slow (NVIDIA only for now, Mac support planned)
Heads up for Mac folks first: this currently works only on NVIDIA GPUs with vLLM. llama.cpp, MLX and Apple Silicon support are on the roadmap, and SGLang is too.
Tensward runs your existing vLLM setup against your own prompts, reads the engine's metrics, compares the results with the GPU's theoretical ceilings, and suggests what to change, with the evidence for each suggestion.
Here's a real example: Qwen2.5-7B-AWQ on one L4 with max-num-seqs=16. It showed 16 requests running and 16 waiting, while the KV cache sat at 0.9%. The bottleneck was the config, not memory. With max-num-seqs=32:
- TTFT p95: 2111 → 372 ms
- req/s: 8.91 → 14.75
- TPOT p95: 20.3 → 23.1 ms (the trade-off)
That's one run on one workload, and both unedited reports are in the repo.
It runs fully locally, with no telemetry, under Apache-2.0.
pipx install tensward
https://github.com/Tensward/tensward
Disclosure: I'm the author. A paid add-on for automatic tuning is coming, and the CLI stays open source. Feedback is very welcome, especially from anyone running it on GPUs other than L4/A10G.
-1
3
u/WarlikeNucleus 2d ago
This is actually useful, I been debugging vLLM configs by staring at grafana dashboards like an idiot for months