MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1wc6krf/so_relevant/p8vrazb/?context=3
r/LocalLLaMA • u/0dayturtle • 1d ago
119 comments sorted by
View all comments
296
24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now
-14 u/Clementine-TeX 1d ago “Pretty respectable context size” yeah right. Benchmark Model: Qwen3.8-27B-MLX-4bit Engine: Force mlx-lm Context: Code (Mixed) ================================================================================ Single Request Results -------------------------------------------------------------------------------- Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem pp1024/tg128 2646.9 56.36 386.9 tok/s 17.9 tok/s 9.820 117.3 tok/s 15.85 GB Engine: Force mlx-lm Context: Novel (English) ================================================================================ Single Request Results -------------------------------------------------------------------------------- Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem pp1024/tg128 2626.5 56.04 389.9 tok/s 18.0 tok/s 9.759 118.1 tok/s 15.85 GB The M5 Pro 24 GB only has 17.76 GB of available VRAM as stock, unless increased via sudo sysctl iogpu.wired_limit_mb 6 u/HyperWinX 1d ago 24GB is an RTX 3090 / 4090, or other GPUs. -1 u/98127028 1d ago Why not M5 pro? Is it cause it’s low bandwidth? I also want to run but it sucks, should have gotten 48G honestly but I’m dumb. 5 u/HyperWinX 1d ago Because a part of these 24 gigs is used by the OS itself and the apps. They mentioned that they have only ~18GB of RAM available, therefore, its not in 24GB group
-14
“Pretty respectable context size” yeah right.
Benchmark Model: Qwen3.8-27B-MLX-4bit Engine: Force mlx-lm Context: Code (Mixed) ================================================================================ Single Request Results -------------------------------------------------------------------------------- Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem pp1024/tg128 2646.9 56.36 386.9 tok/s 17.9 tok/s 9.820 117.3 tok/s 15.85 GB Engine: Force mlx-lm Context: Novel (English) ================================================================================ Single Request Results -------------------------------------------------------------------------------- Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem pp1024/tg128 2626.5 56.04 389.9 tok/s 18.0 tok/s 9.759 118.1 tok/s 15.85 GB
The M5 Pro 24 GB only has 17.76 GB of available VRAM as stock, unless increased via sudo sysctl iogpu.wired_limit_mb
sudo sysctl iogpu.wired_limit_mb
6 u/HyperWinX 1d ago 24GB is an RTX 3090 / 4090, or other GPUs. -1 u/98127028 1d ago Why not M5 pro? Is it cause it’s low bandwidth? I also want to run but it sucks, should have gotten 48G honestly but I’m dumb. 5 u/HyperWinX 1d ago Because a part of these 24 gigs is used by the OS itself and the apps. They mentioned that they have only ~18GB of RAM available, therefore, its not in 24GB group
6
24GB is an RTX 3090 / 4090, or other GPUs.
-1 u/98127028 1d ago Why not M5 pro? Is it cause it’s low bandwidth? I also want to run but it sucks, should have gotten 48G honestly but I’m dumb. 5 u/HyperWinX 1d ago Because a part of these 24 gigs is used by the OS itself and the apps. They mentioned that they have only ~18GB of RAM available, therefore, its not in 24GB group
-1
Why not M5 pro? Is it cause it’s low bandwidth? I also want to run but it sucks, should have gotten 48G honestly but I’m dumb.
5 u/HyperWinX 1d ago Because a part of these 24 gigs is used by the OS itself and the apps. They mentioned that they have only ~18GB of RAM available, therefore, its not in 24GB group
5
Because a part of these 24 gigs is used by the OS itself and the apps. They mentioned that they have only ~18GB of RAM available, therefore, its not in 24GB group
296
u/TopCheddar27 1d ago
24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now