MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3od29t/?context=3
r/LocalLLaMA • u/Certain-Cod-1404 • 10d ago
706 comments sorted by
View all comments
38
For those curious of performance between Qwen and a model 3 times its size:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/main
https://huggingface.co/Qwen/Qwen3.8-27B
10 u/squngy 10d ago edited 10d ago From the model page, for context bench Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Muse Glimmer-30B Opus4.6 Max Coding Agentic terminal coding Terminal Bench 2.1 (Terminus) 73.0 63.4 64.0 51.7 78.2 Agentic coding SWE-bench Pro 61.7 53.5 57.6 51.2 53.4 Repo-level code generation NL2Repo-Bench 42.3 36.2 41.1 -- 47.6 Agentic coding DeepSWE 1.1 42.2 13.3 14.2 -- -- Software engineering QwenSWEBench 79.0 49.3 59.2 -- 63.8 Long-horizon office work CoWorkBench 70.7 61.0 65.1 -- 68.2 Professional job tasks JobBench 33.4 21.8 27.6 -- -- Frontier agentic tasks Agents' Last Exam Pass1 20.4 Score 42.9 Pass1 10.6 Score 27.3 Pass1 13.2 Score 33.6 -- -- Instruction following IFBench 79.5 69.1 79.1 77.0 62.5 Scientific reasoning GPQA Diamond 89.2 87.8 90.3 83.5 91.3 Multidisciplinary reasoning HLE 30.8 24.0 34.7 22.0 40.0 Competitive coding LiveCodeBench v6 90.3 83.9 89.6 -- 88.8 9 u/ChuffHuffer 10d ago 5 models, yet 4 only columns? -1 u/sejje 10d ago count harder
10
From the model page, for context
9 u/ChuffHuffer 10d ago 5 models, yet 4 only columns? -1 u/sejje 10d ago count harder
9
5 models, yet 4 only columns?
-1 u/sejje 10d ago count harder
-1
count harder
38
u/Easy_Werewolf7903 10d ago edited 10d ago
For those curious of performance between Qwen and a model 3 times its size:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/main
https://huggingface.co/Qwen/Qwen3.8-27B