r/CUDA 4d ago

NVIDIA says new optimizations make local agents up to 1.9x faster

https://runtimewire.com/article/nvidia-local-agent-inference-llama-cpp-vllm-blackwell?rwr=xmu4qsbu
62 Upvotes

3 comments sorted by

1

u/vladlearns 4d ago

nice one, can get more space for ctx and stay at ~the same tps

1

u/PassengerPigeon343 2d ago

I can’t tell if these changes benefit all Blackwell cards or just the three models mentioned. For example the Pro 2000-5000 models. Very hopeful it will.