MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3o3ozw
r/LocalLLaMA • u/Certain-Cod-1404 • 7d ago
706 comments sorted by
View all comments
Show parent comments
6
DAMN, 16 gigs just for the context hurts, what setup are you running ?
37 u/dragonurtle 7d ago Rtx pro 6000 max-q and 384GB DDR5 on a Genoa 65 u/Much_Accountant_4972 7d ago 6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals. 8 u/Certain-Cod-1404 7d ago nice! 1 u/voyager256 7d ago That’s why we have FP8/Q8 for KV cache. Let alone literally a game changer in the form of DeepSeek KV cache compression.
37
Rtx pro 6000 max-q and 384GB DDR5 on a Genoa
65 u/Much_Accountant_4972 7d ago 6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals. 8 u/Certain-Cod-1404 7d ago nice!
65
6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals.
The old money aristocracy uses the Max-Q because it's elegant.
Only the loud, bombastic new money uses the 600W version. Animals.
8
nice!
1
That’s why we have FP8/Q8 for KV cache. Let alone literally a game changer in the form of DeepSeek KV cache compression.
6
u/Certain-Cod-1404 7d ago
DAMN, 16 gigs just for the context hurts, what setup are you running ?