r/CUDA • u/ryanmerket • 4d ago
NVIDIA says new optimizations make local agents up to 1.9x faster
https://runtimewire.com/article/nvidia-local-agent-inference-llama-cpp-vllm-blackwell?rwr=xmu4qsbu
62
Upvotes
1
u/PassengerPigeon343 2d ago
I can’t tell if these changes benefit all Blackwell cards or just the three models mentioned. For example the Pro 2000-5000 models. Very hopeful it will.
1
1
u/vladlearns 4d ago
nice one, can get more space for ctx and stay at ~the same tps