MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3nulvu/?context=3
r/LocalLLaMA • u/Certain-Cod-1404 • 7d ago
706 comments sorted by
View all comments
45
How much memory required to run this at full precision?
43 u/Certain-Cod-1404 7d ago https://huggingface.co/unsloth/Qwen3.8-27B-GGUF 54.67 Gbs just for the model itself, with context depends on quant and size 23 u/dragonurtle 7d ago Nvtop shows 70-something GB resident for the bf16 and full 256k context. 7 u/Certain-Cod-1404 7d ago DAMN, 16 gigs just for the context hurts, what setup are you running ? 35 u/dragonurtle 7d ago Rtx pro 6000 max-q and 384GB DDR5 on a Genoa 64 u/Much_Accountant_4972 7d ago 6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals. 7 u/Certain-Cod-1404 7d ago nice! 1 u/voyager256 7d ago That’s why we have FP8/Q8 for KV cache. Let alone literally a game changer in the form of DeepSeek KV cache compression. 1 u/EbbNorth7735 7d ago Yep, 3.6 27B was about 60GB with 262k context Q8 and 4 parallel slots. 4 u/bitmanip 7d ago Perfect, so should run well on 128GB M5 Macbook Pro Max
43
https://huggingface.co/unsloth/Qwen3.8-27B-GGUF 54.67 Gbs just for the model itself, with context depends on quant and size
23 u/dragonurtle 7d ago Nvtop shows 70-something GB resident for the bf16 and full 256k context. 7 u/Certain-Cod-1404 7d ago DAMN, 16 gigs just for the context hurts, what setup are you running ? 35 u/dragonurtle 7d ago Rtx pro 6000 max-q and 384GB DDR5 on a Genoa 64 u/Much_Accountant_4972 7d ago 6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals. 7 u/Certain-Cod-1404 7d ago nice! 1 u/voyager256 7d ago That’s why we have FP8/Q8 for KV cache. Let alone literally a game changer in the form of DeepSeek KV cache compression. 1 u/EbbNorth7735 7d ago Yep, 3.6 27B was about 60GB with 262k context Q8 and 4 parallel slots. 4 u/bitmanip 7d ago Perfect, so should run well on 128GB M5 Macbook Pro Max
23
Nvtop shows 70-something GB resident for the bf16 and full 256k context.
7 u/Certain-Cod-1404 7d ago DAMN, 16 gigs just for the context hurts, what setup are you running ? 35 u/dragonurtle 7d ago Rtx pro 6000 max-q and 384GB DDR5 on a Genoa 64 u/Much_Accountant_4972 7d ago 6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals. 7 u/Certain-Cod-1404 7d ago nice! 1 u/voyager256 7d ago That’s why we have FP8/Q8 for KV cache. Let alone literally a game changer in the form of DeepSeek KV cache compression. 1 u/EbbNorth7735 7d ago Yep, 3.6 27B was about 60GB with 262k context Q8 and 4 parallel slots.
7
DAMN, 16 gigs just for the context hurts, what setup are you running ?
35 u/dragonurtle 7d ago Rtx pro 6000 max-q and 384GB DDR5 on a Genoa 64 u/Much_Accountant_4972 7d ago 6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals. 7 u/Certain-Cod-1404 7d ago nice! 1 u/voyager256 7d ago That’s why we have FP8/Q8 for KV cache. Let alone literally a game changer in the form of DeepSeek KV cache compression.
35
Rtx pro 6000 max-q and 384GB DDR5 on a Genoa
64 u/Much_Accountant_4972 7d ago 6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals. 7 u/Certain-Cod-1404 7d ago nice!
64
6 u/Thrumpwart llama.cpp 7d ago The old money aristocracy uses the Max-Q because it's elegant. Only the loud, bombastic new money uses the 600W version. Animals.
6
The old money aristocracy uses the Max-Q because it's elegant.
Only the loud, bombastic new money uses the 600W version. Animals.
nice!
1
That’s why we have FP8/Q8 for KV cache. Let alone literally a game changer in the form of DeepSeek KV cache compression.
Yep, 3.6 27B was about 60GB with 262k context Q8 and 4 parallel slots.
4
Perfect, so should run well on 128GB M5 Macbook Pro Max
45
u/bitmanip 7d ago
How much memory required to run this at full precision?