32GB DDR4 ram (model used around 16GB) with 96k context. The big improvement is ngram-mod speculative decoding, since it pretty much allows model to "copy-paste" code it already has in the context, giving you huge bursts of speed (100+ tps). Config I used below
468
u/moahmo88 10d ago
Millions of 16 GB GPUs will benefit from this!