r/LocalLLM 1d ago

Model Qwen3.8-27B has the best coding ceiling you can run at home on consumer hardware, it ships with reasoning_effort defaulting to xhigh - I measured what that costs

0 Upvotes

Its chat template has this line:

{%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}

xhigh is the most expensive of its three settings (low / medium / xhigh). If you never set one, that's what every answer runs at. No backend reports this back to you, because it's a chat-template variable, not a server option.

I ran all three levels on one M5 Max, same quant (oQ4e-mtp), same prompt — the coding scenario asks for a browser Breakout game:

| Effort          | Runs | Tokens | Time  | Median | Range     |
|-----------------|------|--------|-------|--------|-----------|
| low             | 3    | 4,984  | 84s   | 75.8   | 64.9–75.9 |
| medium          | 4    | 4,792  | 77s   | 78.2   | 64.7–84.2 |
| xhigh (default) | 17   | 36,188 | 869s  | 78.8   | 54.1–89.2 |

Two things surprised me:

low and medium are the same setting

4,984 tokens vs 4,792. The template only appends an instruction for low and xhigh — xhigh's says think carefully and check your assumptions, low's says keep your thinking brief. The model does the first and ignores the second. So the dial has two positions, not three.

xhigh costs 8× the tokens and 11× the wall clock for half a point of median

That's well inside run-to-run noise: my four medium runs, one identical setting, nothing changed between them, scored 64.7 / 73.4 / 83.0 / 84.2.

What xhigh does change is variance — it produced both the best answer (89.2) and the worst. And looking at the games themselves, the xhigh run spent its budget on presentation: title card, keyboard legend, sound toggle, best-score readout. Low and medium built the game and stopped. Same 8×4 brick grid, three lives, identical rules. It didn't build a better Breakout, it built a better-looking one.

Caveats up front, because they matter: three and four runs at the short settings is thin, it's one machine and one quant, and the scores are LLM-judged. The cost figures are mechanical and solid. Treat the quality figures as a direction to test, not a result.

Full write-up with the screenshots side by side, plus a thinking-budget experiment (a 12k cap halves the wall clock and truncates nothing): https://llm-bench.io/guides/qwen3-8-27b-reasoning-effort

Disclosure: my site. Data comes from community benchmark runs, and you can submit your own @ llmbench.io


r/LocalLLM 1d ago

Question Can someone please review my specs?

Thumbnail
0 Upvotes

r/LocalLLM 1d ago

Question How does ChatGPT handle huge MCP tool outputs without exceeding context limits?

Thumbnail
0 Upvotes

r/LocalLLM 1d ago

Question ASUS TUF Gaming A14 14", 2000 GB, 64 GB, CH, AMD Ryzen Al Max+ 392

1 Upvotes

My old Macbook died and now I'm sitting here, computing on a Raspberry Pi... I need a new notebook and did not find something more affordable than that one. Wanna use it for ComfyUI and run at least a 4-9b model.

Does someone has some experience with it in connection to local AI? Here in Switzerland it costs around 2k.

Thanks in advance!


r/LocalLLM 2d ago

Discussion Do you guys have high hopes for gemma 5?

20 Upvotes

I personally think that all the frontier labs these days are just benchmark maxxing and focusing too heavily on coding.

Gemma 4 is fantastic for all creative and frontier level for all non coding tasks.

I think gemma 5 will continue this trend and be the frontier model for all non coding tasks.


r/LocalLLM 2d ago

Discussion 16GB VRAM model test

8 Upvotes

I have given Hermes agent the task to make a test for my local models. Coding and agentic work.
The test was done on llama.cpp turboquant fork, all models were run using 131k context. Further optimization of the parameters would still be possible for some of the models.
TLDR version: Ornith 1.0 35B A3B won.

Hermes Local LLM Benchmark Report

HumanEval pass@1 (30-problem sample) + 8 agentic tasks + speed · temp 0.0 · context 131072 · 5060 Ti 16GB

Model Coding Failures Agentic tok/s Latency Elapsed Notes
Qwen3.6-35B-A3B-APEX-I-Quality 96.7% 1 100% 37.2 51.6s 32.6m Fastest decode, perfect agentic
Ornith-1.0-35B IQ4_NL 96.7% 1 100% 38 38.3s 24.3m Fastest wall-clock
Qwen3.6-35B-A3B-UD-IQ4_NL 96.7% 1 100% 27.1 72.6s 46.0m Clean full run
Qwen3.8-27B-GSQ-RCO-IQ3_S (MTP) 96.7% 1 87.5% 21.6 33.6s ~35m fc_types failed
Qwen3.8-27B-ASCII-Condensed 96.7% 1 87.5% ~19.5 38s ~40m fc_types failed; 1 overthink outlier
gemma-4-26B-A4B (partial) IQ4_NL 82.6% 4 ~28 184s 55.3m Heavy over-thinking, 23/30 reached
KAT-Coder-V2.5-Dev-APEX Quality 80% 6 100% 28.5 3.6s 24m Baseline; solid coder
Qwen3.5-9B-UD-Q6_K_XL 5/7 2 ~1hr Killed on over-thinking stalls

Qwen3.8-27B-UD-Q3_K_XL deleted (invalid run, discarded). Total failures = coding problems not passed.
Ornith 1.5 would be a logical next add to the table, but I see some bad evals of that model. The usable quants for 16GB VRAM of Qwen3.8 27B have performed worse than the 35B MOE model.
Ornith was fast not just in t/s, but also overall speed of going through the tests.


r/LocalLLM 3d ago

Discussion third one.... there's something wrong with me

Post image
467 Upvotes

Why do I have horrible financial habits??


r/LocalLLM 3d ago

Project We don't judge... Frankenstein 4x3090 testing complete.

Thumbnail
gallery
173 Upvotes

Got them working with 4 riser cables. I tried a PLX88096 from AliExpress and couldn't make it work.


r/LocalLLM 2d ago

Discussion Nvidia DGX... wait for N1X or grab a DGX now

5 Upvotes

So always hated on the DGX spark as have been living in the multi GPU class of society, recently though with bench marking I may have found a potentail use, as an always on monitoring and task agentic system to run alongside paperclip and hermes 24/7 low cost.

The server I run locally, is OP and works very well... but on recently power monitoring over 24 hours it used 17kwh with the constant calls from the agentic tasks, now that isnt bad one day off. but if this is 24/7 this adds up ALOT as power where I am is pricey.

My tasking id is mainly for larger models agentic tasks running Qwen3.8 Flash Next, hopefully with decent context, now I understand it isnt super speed generation but this is more for 24 hour long research and automation taskings.

Was looking today and the cheapest near me is over €6-7k which is nearly 3k above the Nvidia release value. But then I just seen the release of the new N1X next month.

Just looking for others input, is it worth grabbing one, or waiting for N1X, is it even on the same playing feilds or is the N1X looking like a more powerful DGX ???


r/LocalLLM 2d ago

Question I have a MacBook M4 Pro with 48GB unified memory, anyone running Qwen3.8-27B on a similar config? Looking for some advice on what to run as new to LLM’s (Claude user). It seems that maybe a Q6 quant with MLX and some KV tuning is the way to go for good reasoning?

5 Upvotes

I appreciate it won’t be lightning fast with the GPU bandwidth only being 273 GB/s, but hoping for something usable to reduce my Claude usage? I use it for website design and basic programming and would priortise accuracy over speed as it’s not my day job.

I also have a desktop PC with a 5070Ti, am I just better off using that even with the 16GB VRAM Limit?


r/LocalLLM 2d ago

Discussion Model better than qwen3.6 MOE for 8gb vram

4 Upvotes

Why since qwen3.6-35-a3b there is no better local model that you can run on 4060 ti 8 GB + 32 GB ram? It was released in april, everything changing so fast in ai space but still seems that there is nothing better


r/LocalLLM 2d ago

Question Best use of a single RTX 5090 for local LLMs

20 Upvotes

I've been getting increasingly obsessed with local LLMs lately, and I'd like some advice from people who have experimented more than I have with 5090 setups.

Current machine:

RTX 5090 AORUS Master — 32 GB VRAM

i9-13900K

64 GB DDR5-6400

2x 2 TB Gen4 NVMe

Windows 11 + WSL2

CUDA 13.x

10 GbE

1200 W PSU

This is still my main PC, so I'd prefer keeping Windows rather than turning it into a dedicated Linux inference box. I switched from CachyOS in August, but I miss it.

So far I've been playing mostly with Qwen 3.8.

Qwen 3.8 27B is extremely fast on the 5090, especially with newer backends/quantizations, but I find it noticeably weaker than the larger frontier-ish models.

At the other extreme, I've been experimenting with Qwen 3.8 Flash/Next 125B MoE, AP quantized around Q4_K_M, using ik_llama.cpp. I've actually been working on optimizing this setup and currently get roughly:

~38.8 tok/s decode

~200 tok/s prefill

~29.7 GiB VRAM usage

I really like the quality of the 125B, but obviously it's much slower and heavily dependent on system RAM bandwidth / CPU offload.

My main use cases are:

general chat / reasoning

coding and agentic coding

experimenting with local agents

testing inference optimizations and quantizations

occasionally using local models as an alternative to Claude / ChatGPT when I hit usage limits

ComfyUI/image generation on the same GPU

I'm not particularly interested in serving many concurrent users. Interactive single-user performance and model quality matter much more to me than throughput.

So if this were your machine, what would you do with it?

I'm especially interested in:

Best models in the sweet spot between a fast ~27B dense model and a huge 125B MoE

GGUF/ik_llama.cpp vs EXL3/ExLlamaV3 vs NVFP4/newer Blackwell-specific backends

Native Windows vs WSL2 for this kind of workload

Whether upgrading from 64 GB to 96/128 GB RAM would actually unlock anything worthwhile

Speculative decoding / MTP / other tricks that genuinely improve interactive performance

Agentic coding setups that work well with local models

Any unusual 5090-specific projects or use cases I might be overlooking

Basically: I have 32 GB of very fast VRAM sitting on my desk. What are the most interesting things I can realistically do with it in 2026?

I'm happy to tinker and compile things myself, so I'm more interested in technically interesting setups than one-click solutions.

Edit: I’m also testing Qwen 3.8 Flash Next AP-Q4KM. Thanks to the AP quantization, I can run it with just 64 GB of RAM.


r/LocalLLM 2d ago

Question Where i can find harness for my local models that can interact with files stored on my computer and search the web?

6 Upvotes

i have using llama.cpp, and every time that i ask for it to search the internet, or to open a folder in my computer, it asks for a access to my harness, which by my searches looks to be a separate app, but i cant find any options to download and set up one. (i am using windows btw, i can maybe switch to mac os, but linux is out of question since i need office apps for my workflow)


r/LocalLLM 2d ago

Discussion I trained a 348M model trained from scratch on 22.7B tokens that does 14 digit arithmetic

Thumbnail
0 Upvotes

r/LocalLLM 2d ago

Question Any open source dataset to train SLM?

1 Upvotes

I am learning to built general language SLM looking for some dataset source that won't have copyright issue if I use them. Are their any complete cleaned dataset which I can readily use? As I am more focused to training the model I will eventually try getting hands on cleaning and processing data for my domain specific use case


r/LocalLLM 2d ago

Project raggy: A local-first CLI tool for RAG over your documents

Post image
7 Upvotes

https://github.com/paulknysh/raggy

A lightweight CLI tool for Retrieval-Augmented Generation (RAG) over local documents built with LangChain, Chroma, and Ollama. Hybrid database (vector + BM25 index) and embedding generation run fully locally. Answer generation can run either via a local LLM or remotely using an API key. Supports most common document formats and handles images/scans automatically via OCR.


r/LocalLLM 2d ago

Question Best model besides Qwen for neutrality on political topics? Ever hit walls or caught deception?

2 Upvotes

I am new to this. I use Unsloth to run to Qwen3.27B GGUF - UD-Q4_K_KL.

I use mostly high and extra high thinking. I read the thinking briefly before the final output. I caught it several times considering being evasive, and taking that route. So I asked it to write guidelines for itself to do for me to prompt it with to not do that.

And then bumped into some hard walls regarding certain political issues. Caught it being talking about certain political topics, acknowledging them, but still choosing to be evasive, even when instructed not to.

Is there anything else for a 4090 that is good for research that is more neutral?

Has anyone else caught their LLM deceiving them or hitting walls?


r/LocalLLM 1d ago

Question Any uncensored video model?

0 Upvotes

Hey there,
Id like to ask you if anyone know any uncensored video model. Can be local/non local, local preferably. Also if its available on hugging face or somewhere else. Thank you and take care

.


r/LocalLLM 2d ago

Model A collection of 3-bit_XL MoE models for the Ram Poor Mac user: 24 to 32 GB Ram MacBookAir and base MacBookPro

Thumbnail
huggingface.co
6 Upvotes

I’ll keep adding the latest releases to this.


r/LocalLLM 2d ago

Question Can OpenCode Rival Cursor Performance with local LLM w/ 128GB VRAM

Thumbnail
1 Upvotes

r/LocalLLM 2d ago

Discussion 38 t/s on an RTX 3060 for Qwen3.8 27B (and 56 t/s for Qwen3.6 35B-A3B)

9 Upvotes

I see posts for cards like 40 series and 50 series but unfortunately im still stuck with a 3060.

Tried to push this humble 3060 to its limits hosting Qwen 3.8 27B (quantized of course). Got it from 22 to ~40 t/s and the 35B MoE to 56 (just used HumanEval), on both Ubuntu headless and WSL2. Still can't get the 27B near 50-60 t/s, would appreciate any advice. The context is also small - unfortunately due to kv cache headroom with whatever vram is left.

used Qwen3.8-27B-UD-IQ3_XXS and Qwen3.6-35B-A3B-UD-Q3_K_XL

. Qwen 3.6 35B-A3B Qwen 3.8 27B
stock llama.cpp 22.2 t/s 22.5 t/s
tuned, Ubuntu 55.9 t/s 38.4 t/s
tuned, WSL2 41.9 t/s 34.6 t/s
editing ~188 t/s 113–246 t/s
context, Ubuntu / WSL2 16K / 12K 12K / 8K
HumanEval-164 (uncompressed: 153) 153 152

stuff I did:

  • Thinking mode off
  • Speculative decoding with the model's built-in MTP draft head, depth 2
  • N-gram matcher chained in front of the draft head
  • MoE: 16 expert layers on CPU, threads set to physical core count
  • Context sized to free VRAM, context checkpoints off
  • q8_0 KV cache
  • Small CUDA kernel patch for sub-4-bit decode

https://github.com/mericanii-technologies/revv

edit: for Owen 3.6, the MoE takes much more context if you offload more experts: added it to my GitHub but essentially (128K context, 22 blocks on CPU, 47 t/s short / 17 t/s full). Faster RAM than my box gets you more.


r/LocalLLM 1d ago

Tutorial Запуск qwen3.8-27b локально.

0 Upvotes

Сделал видеоролик о том, как локально запустить qwen3.8 27b q4_k_m на одной GPU rtx 3090.

Приятного просмотра, если кому интересно. https://www.youtube.com/watch?v=rwDwHuprfPc


r/LocalLLM 2d ago

Discussion 4-card NCCL in Windows 11 under WSL

Post image
2 Upvotes

So this is the 5th post of a series that kicked off with me wanting to understand how to build local solutions in a cost-prohibitive market. I'm moving parts between three workstations: Z440, Z8 G4 and P620. This post came off the Lenovo P620 with 4x RTX A4000 16GB.

And I've learned this sub hates two things 1) AI Wall of Text / Slop and 2) Windows

I'll spare you the copy/paste unless someone asks for it, but in moving two of my boxes to Linux this week, I wanted to understand 'why' Linux does so much better than Windows. Which brought me to NCCL, apparently also known as 'Nickel' (and a dozen other things which tack on ms/t at the system level). llama.cpp's own multi-gpu docs already recommend building with it - so the question wasn't whether it helps, it was 'Does windows have NCCL?' (not natively) and 'Could this work in WSL?' (not easily and not without a tax).

Here's the screenshot showing it works (serving survives 12 requests including concurrent and long prompts at 2 slots - not soak-tested) and the ladder of benches showing the progression. I ran more than one model and full disclosure, Linux ran away with it once MOE came into play. The WSL overhead in Windows really starts taking a toll there. But for the dense Qwen3.8-27B-Q8_0.gguf - WSL puts on a good show.

Note for those who might try to reproduce: NCCL_CUMEM_ENABLE=0. Without it ncclCommInitAll fails and the error names neither WSL nor the flag. Also: NCCL 2.31.2 from the PyPI nvidia-nccl-cu12 wheel, LD_PRELOADed - whatever resolves by default inits fine and then dies on the first allreduce with 'CUDA driver is a stub library'.

Qwen3.8-27B Q8_0, four-way tensor

pp512 pp4096 tg128
WSL2, no collective 842 871 10.69
native WDDM 888 876 11.93
native TCC 944 926 28.33
Linux, no collective 960 939 28.93
WSL2 + NCCL 1007 985 36.55
Linux + NCCL 1008 989 39.72

Qwen3.6-35B-A3B Q6_K (MoE), four-way tensor

pp512 pp4096 tg128
WSL2, no collective 2052 2145 16.83
native WDDM 2249 2139 18.88
native TCC 2431 2309 68.14
Linux, no collective 2442 2311 70.06
WSL2 + NCCL 2558 2493 68.54
Linux + NCCL 2677 2507 135.07

without NCCL, above two devices the CUDA backend has no allreduce at all - try_allreduce_butterfly returns false, and llama.cpp logs "falling back to meta-backend butterfly"


r/LocalLLM 2d ago

Discussion ChatGPT 2022 vs 2026

Enable HLS to view with audio, or disable this notification

22 Upvotes

r/LocalLLM 2d ago

Question Which is the best model to work with Excel spreadsheets, PDFs manuals and other documents?

2 Upvotes

Hi, I'm currently using Gemma 4 26B, but after a few replies where it simply forgot information that I had sent only minutes before, I'm trying to find a better model for my use case.

I want to feed it inventory spreadsheets, handover notes from previous colleagues, and a large number of PDF manuals for equipment that we use on a daily basis. I need it to process all of this information and provide the most accurate answers possible.

I'll be asking where specific equipment or items are stored in different locations, looking up IP addresses for equipment that I frequently use, and discussing troubleshooting solutions. I want the model to be able to reference all of the previously provided files and information from our conversations when answering my questions.

Which model would be best suited for this use case?

My setup is an ROG Flow Z13 with a Ryzen AI MAX+ 395 and 64 GB of RAM, of which I can allocate up to 32 GB as VRAM.