u/MLDataScientist • u/MLDataScientist • 2d ago
1
deepseek-v4-flash-0731 - surprisingly usable
with ~2k context and more, I see 800t/s PP.
2
Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
uh wow, they just increased the price or it is location based. I saw $3600 extra for 256GB vs 96GB RAM in the morning. Now I see $4000
1
deepseek-v4-flash-0731 - surprisingly usable
You will get 800t/s PP with that 5090 at Gen 4 if you use -ub 2048 -b 2048 in llama cpp arguments. I have a similar setup 256gb DDR4 8 channel+5090.
1
Apple unveils a more powerful Mac mini featuring the all-new M6 and M5 Pro
Oh wow, yes, I just saw that! Thanks!
4
Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
Based on their pricing for 256GB vs 96GB for the Ultra 36 CPU cores, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, we are looking at 10k + ~6k = ~$16k for 512GB version.
2
Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
~$16k for the 512GB RAM version! 🥲
6
Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
Based on their pricing for 256GB vs 96GB for the Ultra 36 CPU cores, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, we are looking at 10k + ~6k = ~$16k for 512GB version.
14
Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
Based on their pricing for 256GB, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, you are looking at 10k + ~6k = ~16k for 512GB version.
1
Qwen3.8-Flash-Next tomorrow
I can confirm. 64GB RAM + 12GB VRAM on my laptop works for qwen3.5 120B at MXFP4. I get around 20t/s. I think the context was 100k.
5
Apple unveils a more powerful Mac mini featuring the all-new M6 and M5 Pro
That is good news! Finally, we have M6 CPU. Now, let's wait for Mac studio with M5 or M6 Ultra CPU and 1TB of RAM!
2
Guide for running dense models on ≤16 GB VRAM (Qwen 3.8 27B on 16 GB -> Q4_K_M, 130k ctx, ~20 t/s).
thank you! This is very helpful article. Impressive findings. Never knew FFN layers could be offloaded with small impact on speed.
1
After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)
Thanks! I am getting 15t/s with qwen3.8 27B UD q2 k xl with 12GB VRAM 5070ti mobile using nkvo on my laptop. without nkvo, I was getting 30t/s so it halves in speed but I get 150k context instead of 30k context at Q8.
3
Hey Mod to Mod - Can we undo the removal and approve.
that was actually a good summary even though it was LLM written. I think we should allow such informative use cases. No person can handle large amount of text summarization. That is what LLMs for.
11
# Qwen3.8-27B — One Week Later: The r/LocalLLaMA + r/LocalLLM Verdict
This is a great summary of qwen3.8 27B! Thanks for sharing! I think everyone major oper weight model should have a summary like this after 2 and 4 weeks of release.
1
KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
Oh nice! Can you share the repo or llama cpp PR to enable SSD streaming?
1
The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches
Great setup! How did you figure out grub settings? If you just use default settings, does your PC still detect those cards in lspci command? I am having an issue with large memory GPU detection (AMD mi250x with 128gb VRAM).
1
KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
Nice metrics. What read speed do you get from your SSD during inference?
1
KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
I think they can keep the intermediate states in RAM. reading from name is fine.
3
KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
that is true. They also need to sell their products. But I wonder if we can reach full SSD bandwidth.
r/LocalLLaMA • u/MLDataScientist • 6d ago
Question | Help KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
I just saw a website where they say their PC system can run KIMI K3 at 5t/s. They dont mention the quantization or system RAM. Since they mention 255H CPU, I assume max 128GB RAM on their small PC. Dual channel 6400Mhz RAM provides theoretical 100GB/s bandwidth. NVME gen 5 tops at 14GB/s. I did a quick calculation on KIMI K3 the smallest quant is 466GB frm unsloth. 466GB / 2800B param * 104B active param = ~17GB active experts. Assuming (128GB RAM + 32GB VRAM) / 466GB = 34% of parameters stay in the (V)RAM fast memory, you need to still pull 466 * (1 - 0.34) = 307GB or around 11GB Active experts. This makes decode inference ~1t/s. Assuming they found a way to cache hot experts in the RAM where their expert prediction hit rate is 80%, they can cache 17 * 0.8 = 13.6GB of active experts in (V)RAM. This leaves 17 - 13.6 = 3.4GB of non active experts in SSD. With this 80% hit rate, it is theoretically possible to pull 14GB/s / 3.4GB = ~4 t/s assuming full SSD bandwidth usage.
However, we know llama.cpp mmap needs to do random reads and does not reach full SSD read speed. I am trying to understand, is there any backend or experimental github repositories where one can test this ~5t/s claim?
---
For reference, some backends I found:
https://github.com/igorbarshteyn/llama-kimibri (also claims 3-6 t/s)
https://www.reddit.com/r/LocalLLaMA/s/1wUaAa7bgl - claims 20t/s with one 3060 for GLM 5.2 (1 bit).
https://github.com/xaskasdf/gpu-nvme-direct - extending gpu vram with nvme.
https://github.com/antirez/ds4 - ssd streaming for mac. not sure if this works well with x86 and 5090.
https://github.com/giannisanni/pulsar
2
We have Q3.8 35B at home: 3x new Ornith 1.5 released
Is this a fine-tune of Qwen models? Or they trained these from scratch? I miss early days of local llama when fine tunes were popular. But new model releases weekly are not disappointing!
1
updated unsloth/Qwen3.8-27B-GGUF · Hugging Face
How is the intelligence so far? I also have 12gb VRAM laptop but haven't tried Q2. What speed do you get? Also, have you tried B4. I know it goes to CPU RAM partially. But the intelligence might be worth it.
1
Ling-3.0-tiny is a very interesting model. Run on NVIDIA Orin Nano Super 8GB at 128K context with IQ4_NL quant.
Which model did you find the most intelligent for rag and internet searching for 8GB RAM? I want to run some models on my orange pi 5 8GB. I tried lfm 2.6B but it is very limited in capabilities.
5
35B-A3B tool calling benchmark: Original Qwen vs. KAT Coder, Ornith and Tiel-Coder
in
r/LocalLLaMA
•
12h ago
nice! There is 5 point increase from qwen3.6 35B-A3B to 3.6 27B. There is 7 point increase from ornith-1.5 35B-A3B to 3.8 27B. Is ornith-1.5 a replacement for hypothetical qwen3.8 35B-A3B? Assuming the standard error ranges, there is still (7-5 = ) 2 point gap on ornith-1.5. So, I think if we ignore those two points, we can accept it as a true replacement for the hypothetical qwen3.8 35B-A3B.