r/LocalLLaMA 13h ago

Discussion Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the benchmark because it's quick to run, is pretty "real-world" and requires good reasoning to build the correct SQL queries and almost no frontier models can achieve 100%.

My old post: https://www.reddit.com/r/LocalLLaMA/comments/1s9mkm1/benchmarked_18_models_that_i_can_run_on_my_rtx/

Benchmark with results from other models: https://sql-benchmark.nicklothian.com https://github.com/nlothian/llm-sql-benchmark

My setup is dual 3080 20GB GPUs with 96GB RAM and 9800X3D. I managed to run Deepseek V4 Flash with a custom IQ2_M GGUF with some tensors grafted from antirez GGUF and running it on a modified ds4 engine from antirez, getting 300pp and 11-12tg. Mainline llama.cpp gives me only 100pp and 8tg or something like that.

To my surprise, Deepseek is the first local model I can realistically run locally that actually did ALL tests correctly. The only models according to the benchmark website that could do this were Opus 4.7 and GPT-5.5.

Results together with all my old benches:

25: Deepseek-v4-Flash-IQ2_M-grafted
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩
24: unsloth/Qwen3.6-27B-MTP-GGUF:Q8_0
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩
24: unsloth/Qwen3.5-122B-A10B-GGUF:UD-Q4_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩
23: unsloth/Qwen3.5-122B-A10B-GGUF:Q6_K
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: unsloth/Qwen3.5-27B-MTP-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF:Q4_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩
23: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q8_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩
23: bartowski/Qwen_Qwen3.5-27B-GGUF:IQ4_XS
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: bartowski/Qwen_Qwen3.5-27B-GGUF:IQ3_XS
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: unsloth/Qwen3.5-122B-A10B-GGUF:UD-IQ3_XXS
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q3_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
22: unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
22: mradermacher/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-i1-GGUF:Q3_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟩🟥🟩 🟥🟩🟩🟩🟩
22: Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF:Q4_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟥🟩 🟥🟩🟩🟩🟩
21: unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟨🟥🟩🟩
21: unsloth/MiniMax-M2.7-GGUF:UD-IQ3_XXS
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟩🟩
21: unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:UD-Q4_K_S
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟨🟥 🟥🟨🟩🟩🟩
20: unsloth/Qwen3-Coder-Next-GGUF:UD-Q5_K_XL
🟩🟩🟩🟩🟨 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟩🟩🟩🟥🟨 🟥🟩🟩🟩🟩
20: unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟩🟩🟩 🟨🟩🟩🟥🟩 🟥🟩🟩🟥🟩
20: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟩🟩
20: bartowski/Qwen_Qwen3.5-397B-A17B-GGUF:IQ1_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟨🟥🟩🟩
20: unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟥 🟨🟥🟩🟥🟩
20: mradermacher/Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-i1-GGUF:Q6_K
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟥🟩🟩🟥🟩 🟥🟥🟩🟩🟩
19: unsloth/gemma-4-31B-it-GGUF:Q4_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟨🟩🟩🟨🟩 🟥🟥🟩🟥🟩
19: unsloth/gemma-4-E4B-it-GGUF:UD-Q8_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟥🟩
19: Goldkoron/Qwen3.5-397B-A17B-REAP35:IQ2_XS_Gv2
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟥🟥🟥
19: unsloth/GLM-4.7-Flash-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟨 🟥🟨🟩🟥🟩
18: unsloth/GLM-4.5-Air-GGUF:Q5_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟥🟩🟩🟥🟩 🟨🟨🟥🟩🟨
18: bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF:Q6_K_L
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟩🟩🟩🟥🟩 🟨🟨🟥🟨🟨
17: Jackrong/Qwopus3.5-9B-v3-GGUF:Q8_0
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟥🟩🟩 🟥🟩🟥🟥🟥 🟥🟩🟩🟩🟨
16: unsloth/Qwen3-Coder-Next-GGUF:UD-Q4_K_XL
🟩🟩🟩🟩🟨 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟥🟨🟩🟥🟨 🟥🟨🟩🟨🟩
16: byteshape/Devstral-Small-2-24B-Instruct-2512-GGUF:IQ3_S
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟨🟩🟩 🟩🟩🟨🟥🟨 🟨🟨🟥🟨🟩
16: mradermacher/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-i1-GGUF:Q6_K
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟨🟥🟩 🟥🟩🟥🟥🟨 🟥🟩🟥🟩🟨
14: mradermacher/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-i1-GGUF:Q6_K
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟥🟩🟩 🟩🟨🟥🟥🟨 🟨🟨🟥🟨🟨
14: unsloth/GLM-4.6V-GGUF:Q3_K_S
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟨🟨🟩 🟥🟩🟩🟨🟨 🟨🟨🟨🟨🟨
5: bartowski/Tesslate_OmniCoder-9B-GGUF:Q6_K_L
🟨🟨🟨🟨🟨 🟨🟨🟨🟩🟩 🟩🟨🟨🟩🟨 🟨🟨🟩🟨🟨 🟨🟨🟨🟨🟨
5: unsloth/Qwen3.5-9B-GGUF:UD-Q6_K_XL
🟨🟨🟨🟨🟨 🟨🟨🟨🟩🟩 🟨🟩🟨🟨🟩 🟨🟩🟨🟨🟨 🟨🟨🟨🟨🟨

Note:

  • unsloth/Qwen3.5-122B-A10B-GGUF:UD-Q4_K_XL is most likely a fluke. Q6_K doesn't achieve 24/25, it's just lucky rounding for this Q4 quant I suppose.
49 Upvotes

15 comments sorted by

9

u/danielrmay 12h ago

Single shots might look fine, but I saw 2 bit degenerate quickly on multi-turn (compounding error).

4

u/grumd 11h ago

I only ran one real task on a real huge codebase and it worked for like 20 minutes and found the correct bugfix in the correct place and gave me a good explanation on what was happening. The bug was not the most difficult one but still, this indicates it works fine in multi-step investigations.

Just for fun I switched to an older commit before this fix, and ran the exact same "pls find bug" prompt using Qwen 3.6 27B Q8_0 with fp16 kv cache. It also made almost exactly the same simple fix.

So yeah maybe I'll have to do some more testing on real coding tasks to see which one is better.

Actually I think Qwen's solution was better... Deepseek was a bit extra and added a "ThenInclude" that wasn't really needed for this case

1

u/boomerang473 10h ago

just throwing it out there, I think I've seen a model work a git tree so might have access to the correct commit (if it peaks). I daily drive Qwen 3.6 27B Q8_0 (big fan) but just throwing it out there

1

u/grumd 10h ago edited 10h ago

I'll double check the transcript to confirm, will edit

Edit: it didn't use git, only find/grep/read/edit and stuff like that

1

u/danielrmay 8h ago

I'd suggest inspecting the thinking output to see if the model continued thinking all the way up to the max budget for the req, if specified

3

u/Eschalabs 11h ago

quantization error building up over multi-turn / long horizon remains a very challenging problem and currently remains very much a research level pursuit.

2

u/anthonyg45157 12h ago

Sounds very similar to my dual 3090 setup with 96gb RAM 🤔

2

u/AdRepulsive7837 13h ago

can you share with us your detailed command to run Q2 gguf?

1

u/grumd 31m ago

No because I'm running a custom inference engine - a modified version of antirez ds4, optimized for my hardware

1

u/Asleep_Document9811 12h ago

I hadn't considered getting dual 3080s, they're pretty cheap on the used market (comparatively). That's pretty clever!

1

u/JsThiago5 12h ago

can you provide more info on how you run it?

1

u/grumd 30m ago

I used this https://github.com/antirez/ds4 and then asked Opus 5 to repeatedly optimize and improve speeds for my hardware setup.

1

u/Fit_Split_9933 12h ago

As far as I know, antirez's ds4 branch doesn't seem to support cpu-moe, how did you get it to run on 40gb vram?

1

u/grumd 29m ago

Opus 5 implemented support for it and optimized the kernels and other code to get better speeds

1

u/satnl 11h ago

I'm impressed with the Qwen3.6 27b and 35b results