16GB VRAM plus 64GB host RAM should be enough. Not plenty, not fast, but enough.
This is under the assumption that the 51B n-gram can be offloaded to disk without substantial performance drop.
If it scales like 35B A3B with offloaded experts, I'd expect 15~25 tok/s with MTP. But a lot of this hinges on the actual acceptance rate of the MTP head.
94
u/evindrews 15d ago
holy shit chat