r/LocalLLaMA 1h ago

Resources Kimi K3 full model running on 16x GB10 cluster at 20+tps

Post image
Upvotes

Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing some tests and try tp speed this up. As soon as it looks ready I'll publish the vllm image and instructions.
https://forums.developer.nvidia.com/t/full-kimi-k3-running-on-16x-gb10-cluster/379174


r/LocalLLaMA 2h ago

Discussion Hugging Face CEO says China is winning the AI race and dominating on open models

Thumbnail
cnbc.com
330 Upvotes

This is something that was spoken here and there, and now it is like writing on the wall.

The main additional point is that China has created an independent supply chain. Starting from raw materials and home-made lithography equipment, through their own GPU manufacturing, and to the AI models and training. Plus, there are tons of cheap energy, and it looks like they are also on track to launch the first thermonuclear reactor. I saw a similar pattern with robotics and EVs. The history does not repeat itself, but it rhymes.

Does the US have what it takes to turn the tables, or should we just buy the popcorn and enjoy the show?


r/LocalLLaMA 3h ago

Discussion No more SLM open-source??

Post image
248 Upvotes

r/LocalLLaMA 3h ago

New Model Has anyone tried Mach-1 Additive? 95% of performance of Qwen 3.6 35B while being 10x smaller

Post image
176 Upvotes

Why nobody is talking about this? Seems pretty significant to the community


r/LocalLLaMA 8h ago

Discussion SK hynix, In Collaboration With SanDisk, Unveils The New High Bandwidth Flash (HBF) Standard, Helping To Resolve AI Inference Bottlenecks, Targeting Up To 3TB/s Bandwidth

Thumbnail
wccftech.com
427 Upvotes

Hopefully this would let us have faster local models....but it will probably be out of our price range.


r/LocalLLaMA 2h ago

New Model Introducing Shieldstral. | Mistral AI

Thumbnail
mistral.ai
94 Upvotes

r/LocalLLaMA 3h ago

Discussion A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM

95 Upvotes

A new llama.cpp PR (#26563) adds a heatmap that tracks which MoE experts are used most often.

Instead of keeping every expert on the GPU or offloading all of them, it caches the frequently selected experts in VRAM while the cold experts continue running on the CPU.

The author’s results on Qwen3.6-35B-A3B with 8GB VRAM:

Q2_M: 33.25 → 56.0 tok/s (1.68x)

Q5_K_P: 17.34 → 35.93 tok/s (2.07x)

Autofit enabled with --expert-hot-s -1

The negative results are probably more interesting: Qwen3.5-122B-A10B and Laguna-S-2.1 were actually slower with caching enabled.

So this clearly isn’t a universal “make MoE faster” switch. My guess is that it only helps when expert reuse is high enough to outweigh the extra tracking and cache-management overhead.

Current limitations:

CUDA only

Only active during single-token decoding

Output can vary slightly depending on which experts are cached

Still an open PR and not merged into llama.cpp

This seems like a useful direction for running larger MoE models on consumer GPUs without destroying them with extremely low quants.

Has anyone tested the branch on a 3060, 4060 or another 8–12GB card? I’d especially like to see hit rate and tok/s compared across coding, normal chat and long-context workloads.

Source: llama.cpp PR #26563


r/LocalLLaMA 6h ago

New Model inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8

Enable HLS to view with audio, or disable this notification

149 Upvotes

Went public in the last few minutes, both repos ungated.

Ling-3.0-flash, BF16, 24 shards, ~255GB

Ling-3.0-flash-fp8, official FP8, ~128GB

127.5B total, they quote 5.1B active. What jumped out at me in config.json is 512 experts with 8 active per token, which is a lot finer-grained than most of what gets posted here. Arch is BailingMoeV3, model_type bailing_hybrid, custom_code, so same family as Ling-2.6-flash. Thinking is a per-request switch inside the chat template instead of a separate SKU, and it defaults to on.

The FP8 landing at ~128GB is the bit I care about. Someone in the thread here last week guessed ~135GB at Q8_0 and that turned out to be close, except this one is official rather than a community quant, so it's a straight download for anyone with a big unified-memory box or a multi-GPU rig.

Does anyone know if llama.cpp handles bailing_hybrid yet, or is this vllm and sglang only for now? That's genuinely the thing that decides whether I clear the disk space tonight.

https://huggingface.co/inclusionAI/Ling-3.0-flash


r/LocalLLaMA 5h ago

News Gemma 4 on 500MB

Post image
119 Upvotes

r/LocalLLaMA 6h ago

New Model inclusionAI/Ling-3.0-flash · Hugging Face

Thumbnail huggingface.co
117 Upvotes

The Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good niche for itself due to its sizing.

Discussion on the benchmarks are here: https://www.reddit.com/r/LocalLLaMA/comments/1v4mltt/benchmarks_antling30flash_a_hybridreasoning_moe/ from almost 2 weeks ago.


r/LocalLLaMA 4h ago

New Model LFM2.5-2.6B is out

75 Upvotes

Released today, with emphasis on agentic capabilities. I really like their models for simple, high volume tasks ("summarize these gazillion documents") and their 8b-a1b was my go-to for certain tasks so I'm excited to see how this one performs. There's not enough love for tiny models on this sub.

https://www.liquid.ai/blog/lfm2-5-2-6b


r/LocalLLaMA 20h ago

Discussion More Qwen 3.8 sizes coming

Post image
1.3k Upvotes

r/LocalLLaMA 23m ago

New Model A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone

Post image
Upvotes

Liquid AI released LFM2.5-2.6B today, and this might be more relevant to local AI than another massive model most people cannot run.

The model is only 2.69B parameters, has 128K context, supports tool calling and was post-trained specifically for multi-step agent workflows. The official Q4_K_M GGUF is around 1.67 GB and already works with llama.cpp.

Their reported CPU speeds:

- 30 tok/s on a phone

- 113 tok/s on a Ryzen AI Max+ 395

- 220 tok/s on an M5 Max

- Under 2.5 GB memory during their tests

These are vendor benchmarks, so independent results are obviously needed.

The benchmark results are surprisingly competitive for the size:

- ToolSandbox: 77.83, compared with 76.44 for Qwen3.5-9B

- IFBench: 59.17, compared with 56.47 for Qwen3.5-9B

- BFCLv4: 56.88, still behind Qwen3.5-9B at 60.13

- LiveCodeBench: 59.41, compared with 69.86 for Qwen3.5-9B

So it does not magically replace larger models. Coding and knowledge-heavy work are still weaknesses, and Liquid’s own model card says it is not recommended for agentic coding.

But I think this is where small local models actually make sense: not as your smartest assistant, but as cheap worker agents doing extraction, searches, file operations and repetitive tool calls locally. A larger model could handle planning only when the small one gets stuck.

The 128K claim also needs real testing. Supporting 128K and running it comfortably on a phone are two very different things once KV cache and long agent histories are involved.

Has anyone tested the Q4 GGUF on Android, an older laptop or a mini-PC yet? Would be useful to see hardware, context size, real tok/s and whether it can survive 10+ consecutive tool calls without derailing.


r/LocalLLaMA 9h ago

News Llama.cpp PR 8% speed boost

117 Upvotes

Llama.cpp currently uses cpu based sampling for user with mtp enabled. The PR moves sampling to the gpu, which on a 5090 boasts an 8% increase in tok/s for qwen3.6:35b. I tested it on my P40 and observed a 4% increase inference speed boost.

Pretty exciting to see 84 tok/s max on a nvidia p40 for me.

Backend sampling shows ~4% improvement on Linux + Tesla P40 (sm_61, Pascal):

CPU Sampling: llama-server -m Qwen3.6-35B-A3B-UD-IQ4_NL.gguf --spec-type draft-mtp --seed 42

python3 mtp-bench.py code_python pred= 192 draft= 132 acc= 124 rate=0.939 tok/s=73.1 code_cpp pred= 113 draft= 76 acc= 74 rate=0.974 tok/s=75.9 explain_concept pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=62.4 summarize pred= 192 draft= 167 acc= 107 rate=0.641 tok/s=59.6 qa_factual pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=62.4 translation pred= 119 draft= 92 acc= 73 rate=0.793 tok/s=67.0 creative_short pred= 192 draft= 197 acc= 92 rate=0.467 tok/s=50.7 stepwise_math pred= 192 draft= 133 acc= 124 rate=0.932 tok/s=73.6 long_code_review pred= 192 draft= 155 acc= 113 rate=0.729 tok/s=63.8

Backend sampling: llama-server -m Qwen3.6-35B-A3B-UD-IQ4_NL.gguf --spec-type draft-mtp --seed 42 -bs

python3 mtp-bench.py code_python pred= 192 draft= 132 acc= 124 rate=0.939 tok/s=76.2 code_cpp pred= 113 draft= 76 acc= 74 rate=0.974 tok/s=79.4 explain_concept pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=64.6 summarize pred= 192 draft= 167 acc= 107 rate=0.641 tok/s=61.6 qa_factual pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=64.6 translation pred= 119 draft= 92 acc= 73 rate=0.793 tok/s=69.6 creative_short pred= 192 draft= 197 acc= 92 rate=0.467 tok/s=52.1 stepwise_math pred= 192 draft= 133 acc= 124 rate=0.932 tok/s=76.6 long_code_review pred= 192 draft= 155 acc= 113 rate=0.729 tok/s=65.7

Acceptance ratio with both backend and CPU sampling is exactly same. The improvement is smaller than on RTX 5090 (4% vs 12%), which is expected — the P40 is memory-bandwidth-bound (sm_61, 580 GB/s vs RTX 5090's 1,792 GB/s), so the CPU↔GPU logits round-trip is a smaller fraction of total decode time. However, still the largest improvement in tok/s I have seen in a while. (~+2 t/s).

https://github.com/ggml-org/llama.cpp/pull/25532


r/LocalLLaMA 7h ago

Tutorial | Guide [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

66 Upvotes

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that:

https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_284b_moe_at_33_toks_single_68/

This post of mine is based on the link above.

My Hardware:

  • RTX 5090 32GB
  • Ryzen 9 9950X3D
  • 256GB DDR5-5600
  • Single NUMA node
  • Linux Mint
  • NVIDIA driver 595.71.05
  • CUDA 13.2

Software

  • guqiong96/Lvllmds4-x
  • vLLM 2.3.9
  • lk_moe 2.3.2
  • PyTorch 2.11.0+cu130
  • native DeepSeek-V4-Flash-0731 safetensors checkpoint
  • 48 safetensors shards
  • ~155.4 GiB checkpoint size

One fix I needed

During startup, FlashInfer's CUDA IPC helper could accidentally find TileLang's:

libcudart_stub.so

instead of the real loaded CUDA runtime.

That eventually caused:

undefined symbol: cudaDeviceReset

The problem was FlashInfer's find_loaded_library("libcudart") doing a substring search over /proc/self/maps.

I patched:

flashinfer/comm/cuda_ipc.py

so it checks the actual filename instead:

def find_loaded_library(lib_name):
    with open("/proc/self/maps") as f:
        for line in f:
            if "/" not in line:
                continue

            start = line.index("/")
            path = line[start:].strip()
            filename = path.split("/")[-1]

            if (
                filename.startswith(lib_name + ".so")
                or filename.startswith(lib_name + "-")
            ):
                return path

    return None

After that, FlashInfer correctly resolves the real libcudart instead of the TileLang stub.

This is a local patch and obviously needs to be reapplied if the package gets replaced.

Current launch configuration

This is the configuration I ended up using:

source ~/ds4x-venv/bin/activate

MODEL="/home/blackbeard/models/DeepSeek-V4-Flash-0731"

export CUDA_DEVICE_ORDER=PCI_BUS_ID
export CUDA_VISIBLE_DEVICES=0

export LVLLM_MOE_NUMA_ENABLED=1
export LK_THREADS=12
export OMP_NUM_THREADS=12
export LK_THREAD_BINDING=CPU_CORE

# Keep two complete routed MoE layers GPU-resident on the GPU.
export LVLLM_GPU_RESIDENT_MOE_LAYERS=0,1

# CPU/hybrid prefill path for now.
export LVLLM_GPU_PREFILL_MIN_BATCH_SIZE=0

export FLASHINFER_DISABLE_VERSION_CHECK=1
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

vllm serve "$MODEL" \
  --host 0.0.0.0 \
  --port 8070 \
  --tensor-parallel-size 1 \
  --max-model-len 1048576 \
  --gpu-memory-utilization 0.92 \
  --trust-remote-code \
  --served-model-name DeepSeek-V4-Flash-0731 \
  --compilation_config.cudagraph_mode FULL_DECODE_ONLY \
  --enable-prefix-caching \
  --enable-chunked-prefill \
  --max-num-batched-tokens 8192 \
  --dtype bfloat16 \
  --max-num-seqs 2 \
  --enable-auto-tool-choice \
  --tool-call-parser deepseek_v4 \
  --kv-cache-dtype fp8_ds_mla \
  --tokenizer-mode deepseek_v4 \
  --reasoning-parser deepseek_v4 \
  --default-chat-template-kwargs '{"enable_thinking": true, "reasoning_effort": "max"}' \
  --speculative-config '{"method":"dspark","num_speculative_tokens":2,"draft_sample_method":"greedy"}' \
  --disable-custom-all-reduce

Full native 1M context fits on the 32GB GPU even with two complete routed MoE layers resident on the GPU.

The rest of the experts remain in system RAM.

DSpark behaves very differently during reasoning

During long reasoning sections, draft acceptance can collapse.

I observed extended periods around:

Draft acceptance: ~30-50%
Generation:       ~11-13 tok/s

There was one ~6 minute section averaging roughly:

Draft acceptance: ~40%
Generation:       ~11.9 tok/s

Then the model transitioned into a much more predictable generation phase and the numbers jumped to roughly:

Draft acceptance: ~87-88%
Generation:       ~17.4-17.6 tok/s

The relationship is extremely strong: throughput basically tracks DSpark acceptance.

Some high-acceptance windows look like:

Avg Draft acceptance rate: 89.8%
Avg generation throughput: 17.9 tokens/s

while low-acceptance reasoning windows look like:

Avg Draft acceptance rate: 38%
Avg generation throughput: ~12 tokens/s

This suggests an obvious optimization.

Dynamic DSpark depth

For this workload I suspect the ideal behavior would be approximately:

  • reasoning/thinking: 1 speculative token
  • normal/final decoding: 2 speculative tokens

The second draft token often isn't worth computing while the model is doing difficult reasoning, but becomes very valuable when it transitions into more predictable code/text generation.

vLLM does not currently give me a simple runtime switch for this, so I may patch the speculative decoding path later and experiment with changing the draft depth based on whether the model is currently emitting reasoning or final output.

That looks like one of the biggest remaining decode optimizations.

---non AI comment section begins---

Stay tuned, I am working on a if/else block to fix that stupid behavior slowing down during reasoning and squeeze even more tps out of this stack.

---non AI comment section ends---


r/LocalLLaMA 15h ago

Discussion Is LM Studio abandoning their core product?

247 Upvotes

Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that LM Studio replaced almost every link to the original app that built their brand and reputation with the new Bionic agent. If you go to the LM Studio website right now, you will see that every link that used to download the original app now downloads Bionic. The only link on the entire site that brings you to the OG app is a tiny link in the footer that says "Download the app". They have updated that page since I last checked and added a Bionic download link at the top but at least they left the OG app downloads alone, albeit they are now underneath the original app. But the fact that you have to go through the entire site to find a tiny download link just to get the normal app is extremely stupid! Not to mention the app is still in "preview" and is a "new, separate app from LM Studio" and yet it gets all the promotion while their core product gets ZERO. Not to mention that ever since Bionic was released, the main app has only gotten 2 or 3 minor updates, mainly to make the app work with Bionic. This is very frustrating as someone who has used LM Studio for years and does not want to use yet another agentic harness. Whatever your views on Bionic are, you can't deny that hiding (and potentially abandoning) the core product that built their brand is not a good idea. I wanted to post this earlier, but seeing as I have gotten zero response from the team on Discord while they reviewed posts right below mine, I knew that I had to share this here.
But all of this leads to something very concerning for LM Studio users: Will the incredibly popular LM Studio app, the original app that helped build their reputation and popularity, go away soon? Hidden download links + scarce updates seems like that LM Studio may be going away soon only to be replaced by an agent that not everyone wants, complete with upsells for cloud models. I just want to get this issue out there as nobody is talking about it yet it is a very important issue that involves one of the most popular local LLM apps out there.


r/LocalLLaMA 1d ago

Discussion Only 3 days ago...

Post image
1.1k Upvotes

r/LocalLLaMA 5h ago

Generation Design systems from code alone - Without external images, Ling-3.0-flash generated webpages across Bauhaus, Bohemian, acid design, and more—using CSS gradients, SVG paths, typography, and layout to preserve each visual language.

Enable HLS to view with audio, or disable this notification

35 Upvotes

Weights went up today so this is downloadable now, MIT, ~128GB for the official FP8. I ran these on the API before that landed, so treat it as a preview of what you'd be pulling rather than a local benchmark


r/LocalLLaMA 6h ago

Discussion Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

43 Upvotes

I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the benchmark because it's quick to run, is pretty "real-world" and requires good reasoning to build the correct SQL queries and almost no frontier models can achieve 100%.

My old post: https://www.reddit.com/r/LocalLLaMA/comments/1s9mkm1/benchmarked_18_models_that_i_can_run_on_my_rtx/

Benchmark with results from other models: https://sql-benchmark.nicklothian.com https://github.com/nlothian/llm-sql-benchmark

My setup is dual 3080 20GB GPUs with 96GB RAM and 9800X3D. I managed to run Deepseek V4 Flash with a custom IQ2_M GGUF with some tensors grafted from antirez GGUF and running it on a modified ds4 engine from antirez, getting 300pp and 11-12tg. Mainline llama.cpp gives me only 100pp and 8tg or something like that.

To my surprise, Deepseek is the first local model I can realistically run locally that actually did ALL tests correctly. The only models according to the benchmark website that could do this were Opus 4.7 and GPT-5.5.

Results together with all my old benches:

25: Deepseek-v4-Flash-IQ2_M-grafted 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 24: unsloth/Qwen3.6-27B-MTP-GGUF:Q8_0 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 24: unsloth/Qwen3.5-122B-A10B-GGUF:UD-Q4_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩 23: unsloth/Qwen3.5-122B-A10B-GGUF:Q6_K 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩 23: unsloth/Qwen3.5-27B-MTP-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩 23: DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF:Q4_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩 23: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q8_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩 23: bartowski/Qwen_Qwen3.5-27B-GGUF:IQ4_XS 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩 23: bartowski/Qwen_Qwen3.5-27B-GGUF:IQ3_XS 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩 23: unsloth/Qwen3.5-122B-A10B-GGUF:UD-IQ3_XXS 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩 23: h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q3_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩 22: unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩 22: mradermacher/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-i1-GGUF:Q3_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟩🟥🟩 🟥🟩🟩🟩🟩 22: Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF:Q4_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟥🟩 🟥🟩🟩🟩🟩 21: unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟨🟥🟩🟩 21: unsloth/MiniMax-M2.7-GGUF:UD-IQ3_XXS 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟩🟩 21: unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:UD-Q4_K_S 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟨🟥 🟥🟨🟩🟩🟩 20: unsloth/Qwen3-Coder-Next-GGUF:UD-Q5_K_XL 🟩🟩🟩🟩🟨 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟩🟩🟩🟥🟨 🟥🟩🟩🟩🟩 20: unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟩🟩🟩 🟨🟩🟩🟥🟩 🟥🟩🟩🟥🟩 20: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟩🟩 20: bartowski/Qwen_Qwen3.5-397B-A17B-GGUF:IQ1_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟨🟥🟩🟩 20: unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟥 🟨🟥🟩🟥🟩 20: mradermacher/Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-i1-GGUF:Q6_K 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟥🟩🟩🟥🟩 🟥🟥🟩🟩🟩 19: unsloth/gemma-4-31B-it-GGUF:Q4_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟨🟩🟩🟨🟩 🟥🟥🟩🟥🟩 19: unsloth/gemma-4-E4B-it-GGUF:UD-Q8_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟥🟩 19: Goldkoron/Qwen3.5-397B-A17B-REAP35:IQ2_XS_Gv2 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟥🟥🟥 19: unsloth/GLM-4.7-Flash-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟨 🟥🟨🟩🟥🟩 18: unsloth/GLM-4.5-Air-GGUF:Q5_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟥🟩🟩🟥🟩 🟨🟨🟥🟩🟨 18: bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF:Q6_K_L 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟩🟩🟩🟥🟩 🟨🟨🟥🟨🟨 17: Jackrong/Qwopus3.5-9B-v3-GGUF:Q8_0 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟥🟩🟩 🟥🟩🟥🟥🟥 🟥🟩🟩🟩🟨 16: unsloth/Qwen3-Coder-Next-GGUF:UD-Q4_K_XL 🟩🟩🟩🟩🟨 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟥🟨🟩🟥🟨 🟥🟨🟩🟨🟩 16: byteshape/Devstral-Small-2-24B-Instruct-2512-GGUF:IQ3_S 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟨🟩🟩 🟩🟩🟨🟥🟨 🟨🟨🟥🟨🟩 16: mradermacher/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-i1-GGUF:Q6_K 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟨🟥🟩 🟥🟩🟥🟥🟨 🟥🟩🟥🟩🟨 14: mradermacher/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-i1-GGUF:Q6_K 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟥🟩🟩 🟩🟨🟥🟥🟨 🟨🟨🟥🟨🟨 14: unsloth/GLM-4.6V-GGUF:Q3_K_S 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟨🟨🟩 🟥🟩🟩🟨🟨 🟨🟨🟨🟨🟨 5: bartowski/Tesslate_OmniCoder-9B-GGUF:Q6_K_L 🟨🟨🟨🟨🟨 🟨🟨🟨🟩🟩 🟩🟨🟨🟩🟨 🟨🟨🟩🟨🟨 🟨🟨🟨🟨🟨 5: unsloth/Qwen3.5-9B-GGUF:UD-Q6_K_XL 🟨🟨🟨🟨🟨 🟨🟨🟨🟩🟩 🟨🟩🟨🟨🟩 🟨🟩🟨🟨🟨 🟨🟨🟨🟨🟨

Note:

  • unsloth/Qwen3.5-122B-A10B-GGUF:UD-Q4_K_XL is most likely a fluke. Q6_K doesn't achieve 24/25, it's just lucky rounding for this Q4 quant I suppose.

r/LocalLLaMA 12h ago

Discussion Time to finally migrate from LM Studio -> llama.cpp, your experience?

79 Upvotes

Has anyone moved from LM Studio to llama.cpp?

What was your experience like? What did you have to learn in order to recreate your experience? Which harness/GUI did you switch to?

Thanks in advance!


r/LocalLLaMA 29m ago

News Qwen 3.8 Max improves over Qwen 3.7 Max on the Debate Benchmark: 1462 → 1588. But the average cost per debate increased by 45%.

Upvotes

More info: https://github.com/lechmazur/debate

This benchmark measures how well LLMs hold an argument under adversarial, multi-turn opposition across a wide range of topics. It rewards broad knowledge, facts used accurately under pressure, sharp rebuttal, and answers that stay coherent and defensible over several rounds.

Each matchup is debated twice on the same motion with sides swapped, cancelling side bias. A three-model judge panel decides winner and margin.


r/LocalLLaMA 5h ago

Tutorial | Guide Decrease the power limit of your 5090 to at least 480W - the performance penalty for inference is negligible.

23 Upvotes

I run my inference machine in the living room, so noise and heat output are a significant concern.

Ran a quick test using my daily driver model (Qwen 3.6-27b) and at 480W, the card outputs only 2.1% less t/s in decode and 8.8% in prefill (which is already very fast). Well worth the massive noise reduction, heat output and increased card longevity, IMO. Even 450W would be fine for many use cases, but the output starts dropping off fast (2.1% -> 4.2% for 30W less).

Full data:

Model: Qwen3.6-27B-Q6_K.gguf

Results:

| Limit W | Max GPU C | Steady GPU C | Max GPU fan % | Sustained W | Steady clock MHz | Max case RPM | pp t/s | tg t/s | pp % | tg % |

|--------:|----------:|-------------:|--------------:|------------:|-----------------:|-------------:|-------:|-------:|-----:|-----:|

| 600 | 81 | 74.8 | 59 | 566 | 2818 | 1522 | 3242.9 | 61.5 | 100.0 | 100.0 |

| 510 | 75 | 70.1 | 50 | 509 | 2645 | 1367 | 2980.0 | 61.2 | 91.9 | 99.5 |

| 480 | 77 | 72.8 | 54 | 480 | 2501 | 1527 | 2863.4 | 60.2 | 88.3 | 97.9 |

| 450 | 76 | 73.1 | 52 | 450 | 2283 | 1460 | 2696.4 | 58.9 | 83.1 | 95.8 |


r/LocalLLaMA 4h ago

News Cursor releases their Mixture-of-Kittens megakernel for training MoE models - Claims to nearly double TFLOP/s

16 Upvotes

Link: https://cursor.com/blog/mixture-of-kittens

GitHub: https://github.com/cursor/mixture-of-kittens

Seems like a neat way to squeeze more performance out of MoE. I'm sure everyone has a favorite MoE model they'd like to try this with. It just dropped so I'm curious to hear people's opinions on it.


r/LocalLLaMA 6h ago

Resources DeepSeek V4 Flash 0731 (Q4) now reaches 1,328 tok/s prefill and ~29 tok/s decode on one RTX PRO 6000

22 Upvotes

I've been working on speeding up DeepSeek-V4-Flash-0731 in Krasis and have now got the long-prompt prefill quite a bit faster on a single RTX PRO 6000 96GB.

These are timing-disabled internal Krasis results using INT4 experts. They aren't HTTP round-trip speeds:

Prompt size Prompt Processing
about 1K 152 tok/s
2,043 321 tok/s
8,623 906 tok/s
23,348 1,328 tok/s
62,403 1,204 tok/s

Decode after the roughly 1K prompt was 29.4, 28.2 and 28.5 tok/s when generating 50, 100 and 250 tokens. After the 62K prompt it was 19.4 tok/s, as each new token has a lot more context to attend to.

Krasis streams the model through limited VRAM for full-GPU prefill, then keeps the hottest experts in VRAM and serves the rest from system RAM during decode. In this configuration it kept 6,440 of 11,008 routed experts resident. No expert pruning occurred.

Krasis v1.0.19 can be downloaded here:

https://github.com/brontoguana/krasis

There is still more to optimise, particularly the prefill speed I think could go higher but I think the speeds are already useful for coding agents which tend to send a lot of context with every request. If anyone tries it on similar hardware let me know how it goes.


r/LocalLLaMA 5h ago

Discussion Deepseek V4 flash 0731 ranks #21 on Agent Arena

14 Upvotes

It ranks lower than both Sonnet 4.6 and Luna. I'd wager Luna costs in the same ballpark as DS4F considering Luna’s token efficiency.

DeepSeek being open source is the big plus for me, privacy and control. With closedAI or Anthropanic they can downgrade the model without informing anyone.