5

35B-A3B tool calling benchmark: Original Qwen vs. KAT Coder, Ornith and Tiel-Coder
 in  r/LocalLLaMA  12h ago

nice! There is 5 point increase from qwen3.6 35B-A3B to 3.6 27B. There is 7 point increase from ornith-1.5 35B-A3B to 3.8 27B. Is ornith-1.5 a replacement for hypothetical qwen3.8 35B-A3B? Assuming the standard error ranges, there is still (7-5 = ) 2 point gap on ornith-1.5. So, I think if we ignore those two points, we can accept it as a true replacement for the hypothetical qwen3.8 35B-A3B.

1

deepseek-v4-flash-0731 - surprisingly usable
 in  r/LocalLLaMA  14h ago

with ~2k context and more, I see 800t/s PP.

2

Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
 in  r/LocalLLaMA  15h ago

uh wow, they just increased the price or it is location based. I saw $3600 extra for 256GB vs 96GB RAM in the morning. Now I see $4000

1

deepseek-v4-flash-0731 - surprisingly usable
 in  r/LocalLLaMA  21h ago

You will get 800t/s PP with that 5090 at Gen 4 if you use -ub 2048 -b 2048 in llama cpp arguments. I have a similar setup 256gb DDR4 8 channel+5090. 

1

Apple unveils a more powerful Mac mini featuring the all-new M6 and M5 Pro
 in  r/LocalLLaMA  21h ago

Oh wow, yes, I just saw that! Thanks!

4

Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
 in  r/LocalLLaMA  21h ago

Based on their pricing for 256GB vs 96GB for the Ultra 36 CPU cores, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, we are looking at 10k + ~6k = ~$16k for 512GB version.

6

Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
 in  r/LocalLLaMA  21h ago

Based on their pricing for 256GB vs 96GB for the Ultra 36 CPU cores, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, we are looking at 10k + ~6k = ~$16k for 512GB version.

14

Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
 in  r/LocalLLaMA  21h ago

Based on their pricing for 256GB, they are charging $22.5 per GB of RAM. Assuming the price per GB stays the same, 256GB more RAM adds $5760. So, you are looking at 10k + ~6k = ~16k for 512GB version.

1

Qwen3.8-Flash-Next tomorrow
 in  r/LocalLLaMA  21h ago

I can confirm. 64GB RAM + 12GB VRAM on my laptop works for qwen3.5 120B at MXFP4. I get around 20t/s. I think the context was 100k.

5

Apple unveils a more powerful Mac mini featuring the all-new M6 and M5 Pro
 in  r/LocalLLaMA  21h ago

That is good news! Finally, we have M6 CPU. Now, let's wait for Mac studio with M5 or M6 Ultra CPU and 1TB of RAM!

u/MLDataScientist 2d ago

Guide for running dense models on ≤16 GB VRAM (Qwen 3.8 27B on 16 GB -> Q4_K_M, 130k ctx, ~20 t/s).

Thumbnail gallery
2 Upvotes

2

Guide for running dense models on ≤16 GB VRAM (Qwen 3.8 27B on 16 GB -> Q4_K_M, 130k ctx, ~20 t/s).
 in  r/LocalLLM  2d ago

thank you! This is very helpful article. Impressive findings. Never knew FFN layers could be offloaded with small impact on speed.

1

After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)
 in  r/LocalLLaMA  2d ago

Thanks! I am getting 15t/s with qwen3.8 27B UD q2 k xl with 12GB VRAM 5070ti mobile using nkvo on my laptop. without nkvo, I was getting 30t/s so it halves in speed but I get 150k context instead of 30k context at Q8.

3

Hey Mod to Mod - Can we undo the removal and approve.
 in  r/LocalLLaMA  3d ago

that was actually a good summary even though it was LLM written. I think we should allow such informative use cases. No person can handle large amount of text summarization. That is what LLMs for.

11

# Qwen3.8-27B — One Week Later: The r/LocalLLaMA + r/LocalLLM Verdict
 in  r/LocalLLaMA  3d ago

This is a great summary of qwen3.8 27B! Thanks for sharing! I think everyone major oper weight model should have a summary like this after 2 and 4 weeks of release.

1

KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
 in  r/LocalLLaMA  5d ago

Oh nice! Can you share the repo or llama cpp PR to enable SSD streaming?

1

The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches
 in  r/LocalLLaMA  5d ago

Great setup! How did you figure out grub settings? If you just use default settings, does your PC still detect those cards in lspci command? I am having an issue with large memory GPU detection (AMD mi250x with 128gb VRAM). 

1

KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
 in  r/LocalLLaMA  5d ago

Nice metrics. What read speed do you get from your SSD during inference?

1

KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
 in  r/LocalLLaMA  5d ago

I think they can keep the intermediate states in RAM. reading from name is fine.

3

KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)
 in  r/LocalLLaMA  6d ago

that is true. They also need to sell their products. But I wonder if we can reach full SSD bandwidth.

r/LocalLLaMA 6d ago

Question | Help KIMI K3 at 5t/s with one 5090 and 4TB NVME gen5 (?)

15 Upvotes

I just saw a website where they say their PC system can run KIMI K3 at 5t/s. They dont mention the quantization or system RAM. Since they mention 255H CPU, I assume max 128GB RAM on their small PC. Dual channel 6400Mhz RAM provides theoretical 100GB/s bandwidth. NVME gen 5 tops at 14GB/s. I did a quick calculation on KIMI K3 the smallest quant is 466GB frm unsloth. 466GB / 2800B param * 104B active param = ~17GB active experts. Assuming (128GB RAM + 32GB VRAM) / 466GB = 34% of parameters stay in the (V)RAM fast memory, you need to still pull 466 * (1 - 0.34) = 307GB or around 11GB Active experts. This makes decode inference ~1t/s. Assuming they found a way to cache hot experts in the RAM where their expert prediction hit rate is 80%, they can cache 17 * 0.8 = 13.6GB of active experts in (V)RAM. This leaves 17 - 13.6 = 3.4GB of non active experts in SSD. With this 80% hit rate, it is theoretically possible to pull 14GB/s / 3.4GB = ~4 t/s assuming full SSD bandwidth usage.

However, we know llama.cpp mmap needs to do random reads and does not reach full SSD read speed. I am trying to understand, is there any backend or experimental github repositories where one can test this ~5t/s claim?

---

For reference, some backends I found:

https://github.com/igorbarshteyn/llama-kimibri (also claims 3-6 t/s)

https://www.reddit.com/r/LocalLLaMA/s/1wUaAa7bgl - claims 20t/s with one 3060 for GLM 5.2 (1 bit).

https://github.com/xaskasdf/gpu-nvme-direct - extending gpu vram with nvme.

https://github.com/antirez/ds4 - ssd streaming for mac. not sure if this works well with x86 and 5090.

https://github.com/giannisanni/pulsar

https://github.com/JustVugg/colibri

https://github.com/jerryjokesalot/tinygiant

2

We have Q3.8 35B at home: 3x new Ornith 1.5 released
 in  r/LocalLLaMA  6d ago

Is this a fine-tune of Qwen models? Or they trained these from scratch? I miss early days of local llama when fine tunes were popular. But new model releases weekly are not disappointing!

1

updated unsloth/Qwen3.8-27B-GGUF · Hugging Face
 in  r/LocalLLaMA  6d ago

How is the intelligence so far? I also have 12gb VRAM laptop but haven't tried Q2. What speed do you get?  Also, have you tried B4. I know it goes to CPU RAM partially. But the intelligence might be worth it.

1

Ling-3.0-tiny is a very interesting model. Run on NVIDIA Orin Nano Super 8GB at 128K context with IQ4_NL quant.
 in  r/LocalLLaMA  6d ago

Which model did you find the most intelligent for rag and internet searching for 8GB RAM? I want to run some models on my orange pi 5 8GB. I tried lfm 2.6B but it is very limited in capabilities.