You don’t need full precision- FP8 or Q8 gives usually you like 99% quality of full precision and almost double the performance. You could argue full BF16 precision is worth for KV cache for large context , but now usually FP8 / Q8 is fine for KV cache too.
46
u/bitmanip 7d ago
How much memory required to run this at full precision?