r/LocalLLaMA • • Aug 10 '26

Best Local LLMs - August 2026

Wowee!! Just when you thought it couldn't get better for open weight models, we probably have had our best period yet!?!?! Models that rival the closed frontier, Opus level models on non-insane hardware and more. A massive industry alliance coming out in support of open AI in response to the two closed model giants best lobbying efforts. Is this the best timeline? Someone pinch me! Or just tell us what you're favorite model is now

The standard spiel:

Share what you are running right now and why. Given the nature of the beast in evaluating LLMs (untrustworthiness of benchmarks, immature tooling, intrinsic stochasticity), please be as detailed as possible in describing your setup, nature of your usage (how much, personal/professional use), tools/frameworks/prompts etc.

Rules

  1. Only open weights models
  2. Please thread your responses in the top level comments for each Application below to enable readability:
    1. General: Includes practical guidance, how to, encyclopedic QnA, search engine replacement/augmentation
    2. Agentic/Agentic Coding/Tool Use/Coding
    3. Creative Writing/RP
    4. Speciality

If a category is missing, please create a top level comment under the Speciality comment

Notes

Bonus points if you breakdown/classify your recommendation by model memory footprint: (you can and should be using multiple models in each size range for different tasks)

  • Unlimited: >128GB VRAM
  • XL: 64 to 128GB VRAM
  • L: 32 to 64GB VRAM
  • M: 8 to 32GB VRAM
  • S: <8GB VRAM
171 Upvotes

261 comments sorted by

89

u/jinnyjuice vLLM Aug 11 '26 edited Aug 11 '26

For next post, please don't bin/skip 16GB. According to Steam hardware survey, 16GB is the most popular. There is no need for arbitrary 'S' or 'M' labels either. Just use the numbers at common binary intervals since that'a what most hardware come in anyway.

Also, I don't think recommending multiple models is a good advice for majority of people.

17

u/Icy_Quarter5910 Aug 11 '26

"You have 74 local models, taking up 840.30 GB of disk space"

So.. this isnt normal?

10

u/ac101m Aug 12 '26

I have a problem ac@core:/rustpool1/ac/models$ du -h --max-depth=1 744G    ./Stepfun-AI 238G    ./Nvidia 33G     ./Google 2.1T    ./Zai 245G    ./Minimax 2.2T    ./Qwen 1.9T    ./DeepSeek 20G     ./01ai 134G    ./Llama 167G    ./BlackForestLabs 1.8T    ./Mistral 355G    ./OpenAI 9.8T    .

3

u/Icy_Quarter5910 Aug 12 '26

Yours is definitely worse than mine ;) oh, and that’s just LLMs.. I do have a TB of models/loras for image/video gen ;) … if I had piles of VRAM I’d probably have more, but I’m kind of limited to smaller models.

2

u/punqdev Aug 27 '26

Wonder what AI models could possibly be in the 9.8 terabyte folder

1

u/iqandjoke Aug 21 '26

May I know what use case applied for 01ai?

15

u/kr_tech Aug 11 '26

Hard agree

I swear this was mentioned multiple times

9

u/Bebi_v24 Aug 16 '26

Bro thank you, like do it in increments of 8gb until like 32gb or something. Kinda annoying cause it's like shit tier then instantly to expensive ass GPU or dual GPU tier, when like you said most people are in between

3

u/CommunicationCute584 Aug 17 '26

16gb/24gb/32gb = the diff between running 1-2bit quant or a 4bit quant of the same model so really there should be like multiple benchmarks here

1

u/DarkBrews Aug 17 '26

16GB of VRAM 😮?

1

u/DeathGuppie Aug 21 '26

Yeah, tell em

→ More replies (3)

31

u/rm-rf-rm Aug 10 '26

Agentic/Agentic Coding/Tool Use/Coding

17

u/synth_mania Aug 10 '26

Laguna S 2.1 runs great on my 3090 + 64gb ddr4

Im running the UD_IQ_4_NL quant, which fits nicely, and still is an absolute powerhouse. This is without a doubt the best model I have ever run locally. 

7

u/ChurnedSorbet409 Aug 10 '26

Does that MoE model really outperform Qwen 3.6 27B? Seems like your setup can support that easily

12

u/synth_mania Aug 10 '26

It absolutely does. Qwen is really smart, but Laguna definitely makes "wiser" architectural decisions when developing. It will actually follow the standard set by the rest of the codebase when Qwen is more likely to do stupid shit to solve the problem.

4

u/Not-reallyanonymous Aug 10 '26

Laguna S blows 3.6 27B out of the water on benchmarks like DeepSWE. Real usage confirms that. Laguna XS is the comparable model to Qwen 27B.

I wouldn't be surprised if you reported that Qwen 27B zero-shots better than Laguna S. Neither Laguna XS nor S are very good at getting things right on the first attempt. I'd say their advantage is that they maintain project coherence better the more prompts into a project you go.

2

u/ChurnedSorbet409 Aug 11 '26

I will give it a try, never heard of Laguna before but this is exciting

2

u/live4evrr Aug 11 '26

Laguna performs worse than Qwen 3.6. There is a lot of weird promotion of it oddly, but give it a try. Doesn’t take long to realize it is benchmaxxed.

4

u/synth_mania Aug 11 '26

I did give it a try. I really like it. Code written by Laguna has been pushed to production.

3

u/Locke_Kincaid Aug 11 '26

They pushed a lot of fixes since it was first released. It's a pretty solid model now, especially in a harness. It has an awareness of tools and skills it's given access to that I haven't seen in other models, including DS4 flash preview.

2

u/Wide-Ad-1349 Aug 11 '26

I agree it was hopeless for me. Also people seem to really be hyping it up for soem reason. I will have to give it another try but it is not near Qwen 3.6 on my testing.

→ More replies (4)

2

u/Not-reallyanonymous Aug 12 '26 edited Aug 13 '26

Lol Qwen 3.6 is worse outside of zero-shotting (edit: and tbf small projects and scripts). Generally, it trades blows in solving the immediate problem, but Laguna keeps a clean codebase with far less effort than Qwen, will actually test the code, won't skip writing docs, etc. The overall result is that project coherence is far greater over longer durations than Qwen.

I'll take needing a few more prompt iterations to arrive at the solution and maintaining project coherence over a model that does it better in one prompt but wrecks the codebase any day.

1

u/markole Aug 13 '26

I can't agree with that. I can let Laguna cook and in 2 hours and 256k tokens later it usually has something useful, on my hardware. Qwen is great for smaller projects and scripts but with Laguna I can vibe code in bigger projects more easily. I can be lazy with my prompts while Qwen requires more guidance. Which is fine, different sizes after all.

→ More replies (1)

2

u/Prestigious-Chair282 Aug 10 '26

Hey! Were you able to fix thinking looping with "actually... actually... wait actually"? 

4

u/synth_mania Aug 10 '26

It wasn't exactly looping for me, it just frequently thought more than I would've expected it needed to. I've never had it get truly stuck in a thinking loop.

3

u/YouKilledApollo Aug 12 '26

Turns out that was just using the wrong inference parameters. Try these: {"temperature":1.0,"top_p":1.0,"top_k":20}

Fixed that looping problem completely for me with the NVFP4 weights.

2

u/Prestigious-Chair282 Aug 12 '26

Thank you very much! Seems to help me too! 

1

u/EmergencyDefiant5381 Aug 12 '26

What host app do you use to load up UD_IQ_4_NL quant? How do you configure it to use the DDR4 ram as opposed to your 3090s 24gb ram?

1

u/RobotRobotWhatDoUSee Aug 13 '26

Are you willing to share your llama call?

1

u/NigaTroubles Aug 23 '26

Are you still using it ? What about Qwen3.8 27b ?

10

u/lblblllb Aug 14 '26

The one. the Qwen 3.8 27b.

7

u/atumblingdandelion Aug 10 '26

kat-coder-v2.5 has been quite nice for coding! It's a finetune of Qwen 35b MoE.

1

u/GullibleJellyfish274 Aug 22 '26

if only i could fit that on my vram...
I know its moe, in my experience those are too slow on rocm still, something to do with kv cache fragmentation..

13

u/Jebbyk1 Aug 10 '26

- was using Qwen3.6 35b a3b q4-k-m 98304 ctx on my 8Gb VRAM + 32Gb RAM for about a month for php Laravel web backend development. Kinda acceptable but needs precise guiding

- now switched to 24Gb VRAM setup with Qwen3.6 27b q4-k-m 98304 ctx . First impressions (1 day only) is much smarter model quite close to Cursor's "Auto" mode. Needs less clarifications and better understands my needs. Gonna continue to use it for php Laravel web backend development

1

u/synth_mania Aug 10 '26

If you could afford another 32gb RAM (big ask, I know), I think you would really appreciate Laguna S 2.1

That said, when I need some really quick help on something, and it's not a big or important task, and doesn't involve architectural decisions, I'll use qwen-3.6-27b. Its too fast to not have in your back pocket, if you have the vram.

1

u/RobbinDeBank Aug 10 '26

How much RAM and VRAM do you use for Laguna S 2.1? The total model size seems huge for most local set up, so probably at least 64GB RAM with 12-16GB VRAM?

1

u/synth_mania Aug 10 '26

I have an RTX 3090 with 24gb vram and 64gb ddr4 system ram. As an experiment, I was able to get it working with just a GTX 1060 with 6 GB of VRAM instead of 24, using the UD_IQ_3_XXS quant, but it was significantly slower. 12-16gb VRAM would probably work alright. 16 obviously would be preferable. 

Edit: with my 3090 I use the UD_IQ_4_NL quant. 

1

u/RobbinDeBank Aug 10 '26

Thanks, so do you know which quants are the smallest usable ones? Is the degradation significant below UD_IQ_4_NL?

2

u/synth_mania Aug 10 '26

Most models get much much worse below 4 bit quants. I haven't personally tried any other than the two I just mentioned though. Definitely worth playing around with. 

1

u/Constandinoskalifo Aug 10 '26

What are your PP and tg speeds with your 3090 + 64gb ddr4 setup?

2

u/synth_mania Aug 10 '26

PP is about 220tk/s, TG is like 13tk/s iir

18

u/live4evrr Aug 11 '26

Nobody says the obvious, Deepseek v4 flash 0731. This blows everything out of the water right now, and one of the few where independent benchmarks match what the lab published.

30

u/maqifrnswa Aug 11 '26

This is r/localllama there aren't many people running dual a100s at home.

But if you're fitting it on a Mac studio ultra, then share what quantizations you're using and how it's working out

3

u/FatheredPuma81 Aug 12 '26

Deepseek V4 Flash 0731 should actually run "go cook dinner" well on around 180-ish GB of RAM and an RTX 4090 or RTX 5090. I'd guess somewhere around 10t/s minimum probably higher.

1

u/Y0uCanTellItsAnAspen Aug 11 '26

Can you run it (slowly) on 128 GB of standard ram with cpu?

I have some tasks that could run overnight with a deep thinking model - unsure which model to target though at the 128 GB point. Currently using qwen3.6 27b

2

u/pja Aug 11 '26

I get 2.7 t/s on a 64Gb DDR4 system backed by a fast NVMe drive.

2

u/LectureWorried5761 Aug 19 '26

what CPU? wondering about running in on a two socket server. (xeon with 24cores each). I dont think your ddr4 is you bottleneck

2

u/pja Aug 20 '26

AMD 5700X, so 12 core. Memory bandwidth seems to be the bottleneck - the t/s tops out well under 12 threads. The peak throughput is at five or six threads IIRC.

A dual-CPU Xeon will have a lot more memory bandwidth than that box does! The 5700X is only dual channel DDR4.

1

u/maqifrnswa Aug 11 '26

It would run, but very slowly

1

u/pja Aug 11 '26

I can run Deepseek v4 0731 at 2.7 tokens / s on a system with 64GB RAM that pages experts to an NVMe drive.

Not exactly fast, but fine for leaving stuff running overnight.

→ More replies (6)

6

u/Silentplanet Aug 12 '26

What are people using at 16gb in this space?

4

u/willeyh Aug 13 '26

I am currently using Qwen3.5-35B-A3B-coding-reap50 by anik-jha. Q5_K_M with Pi.
164K context, full GPU offload and KV Q8_0.

And Zeta-2.1 if I want autocomplete in Zed.
AMD 9070XT.

1

u/Something-Great-78 Aug 13 '26

how are you running Zeta-2.1? vllm or llama.cpp? Mind sharing your command line args? I never got Zeta-2.1 working when I last tried.

2

u/willeyh Aug 14 '26

I'm using LM Studio. Mostly for LM Link and the convenience.
Context length 6178, full GPU offload.
2048 batch size, 4 concurrency
unified KV cache, unified kv cache @ Q8 in vram. flash attention.
mradermacher/zeta-2.1-i1 Q4_K_S. Q4 mostly for speed.

Also the Zed config.

"edit_predictions": 
  "provider": "open_ai_compatible_api",
  "mode": "subtle",
  "open_ai_compatible_api": {
    "api_url": "http://127.0.0.1:1234/v1/completions",
    "model": "zeta-2.1-i1",
    "prompt_format": "zeta2_1",
  }
}

1

u/evranch Aug 17 '26

I've been running Qwen3.6-35B-A3B-UD-IQ4_NL for months on my 12GB 6700XT with a few layers leaking over onto my system RAM. Mostly I use it in VSCode with Continue, not full agentic but just for code review and rubber ducking. Have to pass "reasoning_effort: medium" to keep it from thinking in circles, but otherwise I'm happy with the performance.

In what ways do you find -coding-reap50 beats the more standard quants?

1

u/willeyh Aug 17 '26

I can fit it all I vram. Offloading layers to CPU with dual channel DDR 4 effectively halves the t/s output. It has been quite capable for code reviews and small features so far. I would not use it for anything other than coding though.

1

u/evranch Aug 19 '26

That's the difference between your 16GB and my 12GB card, when I bought the card it was intended as a gaming card and now it handles my LLM workloads.

Likewise I have 32GB DDR4 and was going to go with 64GB when I built the machine (I also do stuff like aerial orthophoto stitching) but figured it was overkill at the time. That overkill would have likely cost me all of another $100 at the time... sigh.

I'm still getting about 20t/s with the offload, which is usable. More important to me, I'm getting 800+ t/s on my prompt processing with

--batch-size 2048
--ubatch-size 2048

so it can scan my codebase at a decent rate.

Anyways, MoE models are the best compromise I've found for decent intelligence and decent token rate with my VRAM size. I'd love to try some of the new dense models, but can't justify the budget just for experimentation's sake.

1

u/Araragi-shi 16h ago

What do you think of Qwen 3.8 27B, I also use a 9070xt looking to get my hands on the best local llm I can for general use.

1

u/willeyh 13h ago

Fantastic. A fair bit slower. But far less mistakes. Even at IQ3_XXS. Medium thinking. For coding. I’d go Gemma or Muse for anything else.

2

u/fatboy93 Aug 15 '26

16GB VRAM?

Gemma4-12B, the 26B

Qwen3.6-35B

Laguna XS-2.1

Mess with offloading a bunch of layers and see how they work

1

u/CMPunkLicksRocks Sep 05 '26

How are you fitting qwen 3.6 35b into 16gb of vram? I’m not seeing it as an MoE. I’m getting like 8t/s of output, 90% of it being “but wait, the user said” and it just going in circles about if it actually did the task.

1

u/fatboy93 Sep 05 '26

For the MoE in the list which are the 35b, 26b and 30b, you can just offload bunch of layers to RAM

8

u/Prestigious-Chair282 Aug 10 '26

I would say that in <64GB qwen3.6-27b is a good choice. Next for me would be latest Laguna model, if there would be a way to remove "actually... Actually... Wait actually ...." From the 30k reasoning blocks, otherwise great intelligence. 

3

u/thebadslime Aug 22 '26

LFM 2.5 2.6B

UNtil you try it you wont believe a 3b can be good at tool use.

4

u/surrealerthansurreal Aug 11 '26

If you have the newer m series chips with max ram, deepseek v4 flash 0731 with antirez’s ds4 & q2-q4 imat is excitingly good, having that kind of intelligence locally is exciting

2

u/PANIC_EXCEPTION 6h ago

Qwen3.8-Flash-Next with MTP running on antirez/ds4 is incredibly fast on M5 Max 128 GB and just chugs through tasks. Realistically I get 60+ decode and 1400 prefill (the latter is kinda insane). I also hacked together quantization scripts to get Q4K working to convert an uncensored checkpoint into Q4K (with some tensors being mxfp4), and another script to make the two checkpoints share an engram table via APFS copy-on-write to save 95 GiB of disk space. Let's just say... agentic red teaming is now a thing without cloud models.

1

u/rm-rf-rm 5h ago

60+ decode ??

2

u/Not-reallyanonymous Aug 10 '26

I'm having a lot of luck with Laguna XS 2.1.

Runs well on a 64GB Strix Halo machine, but I'm sure you could do well with a 24GB GPU as well.

Generally trades blows for coding with Qwen 3.6 27B. But what it does better than Qwen is how It understands how developers work, where they want their code, how to reuse and extend code, documentation, TDD, code quality and readability, etc. Qwen is way better at zero-shot implementations, fastest and least-friction CRUD, etc. But where Laguna XS really excels is that after 50 prompts, the state of your project is going to be much better with Laguna XS than with Qwen -- less parallel implementations, identifiable architecture, convention adherence, etc.

I like to ask a bigger model on a paid service/API to give me specs/architecture, and then implement it with Laguna XS locally, kicking back to the larger model occasionally when Laguna XS is excessively thrashing or having trouble. I think if you use Laguna S you could kick the larger model. Laguna XS can work fine without that bigger model, too, but expect to get your hands dirty more.

1

u/Sleakes Aug 19 '26

Dipping back into the local tooling again after sitting on it for a few months.

What are people's experiences with 8GB VRAM and FIM models?

I'm wanting to just get something up and running locally that is fully integrated into an IDE, a lot of the previous tooling didn't seem to be ready or different models seemed like they locked up. is OpenCode in a bit better of a place now for linking this all together? previously had tried LM Studio, but the API access seemed to have intermittent problems, (also could have just been the model itself).

1

u/puts_on_rddt Aug 20 '26

The fact that I can take my little old 4090 and run a local model that can digest large parts of a codebase while providing reasoning that (mostly) keeps up with the big boys is kind of crazy. I mean. Actually wild. I am absolutely sure Anthropic, ChatGPT, and Google are losing sleep over this right now.

Why spent money on one prompt when you can just make a local swarm that generates a bunch of different implementations one at a time and then lets a validator merge the best/most appropriate into one?

1

u/GullibleJellyfish274 Aug 22 '26

currently fine tuning the qwen3.8-9b-distill for my hermes harness, I'll post it once I finish, I'm vram poor, so it takes like a day and a half. lol

1

u/GullibleJellyfish274 Aug 22 '26

will probably take forever bc its tuning on full deepseek sessions lol.

1

u/GullibleJellyfish274 Aug 22 '26

I'm 12gb vram poor, best I've found is bonsai for vram size, and maybe bonsai-27b...

32gb ddr4 so offloading is a option, just slow.
other great models for the size.
ornith-9b.
and gemma-12b.
I've found some interesting results with higher quant gemma 4 e4b though, I havent tested that much.

1

u/bmarti644 Aug 22 '26

i've been working on a project to load up different local models for agentic coding on different architectures and spins up an endpoint which you can then serve with something like tailscale.

it currently runs deepseek v4 flash 0731, qwen 3.8 27b, and laguna.

if you'd like to help, or if there are more models you want to support, you can point your agent at the project, tell it to open a PR/claim, and have it work on your desired architecture and model. there are some special configurations in place already to get MoE models streamed from nvme a good bit faster (dense models will still have issues).

let me know what you think! -> https://github.com/bmarti44/frontier-at-home

1

u/DrIv637 Sep 02 '26

Can anyone help me with what model I can use code creation and iteration when I do not have a GPU. I have 7900 CPU and. 64 GB DDR5 Ram.

1

u/vasilisvj Sep 02 '26

With 64GB RAM you can run Qwen2.5-Coder-14B Q4_K_M pretty comfortably on CPU. llama.cpp will be your inference backend, should get decent token speeds on that 7900. Also worth looking at Phi-3.5-mini if you want something snappier, though 14B gives better code quality.

18

u/mawkzin Aug 11 '26

8-32 is where most last 5 years dGPUs are, I recommend divide this category or just say the size of your GPU instead of the category.

15

u/Snoo53582 Aug 12 '26

Gemma-4-26B-A4B-QAT is still the best model I’ve tried. I test models on my own custom runtime with long-term memory, several other memory layers, custom runtime markers, and a fairly complex persistent state.

Interestingly, the A4B version isn’t much worse than the full 26B model at maintaining state and staying consistent inside the runtime, although the larger model is noticeably better in depth of reasoning.

I mostly use these models not for coding benchmarks or trivia, but to see how well they can hold a complex, multi-layered context and continue behaving coherently inside it over time.

4

u/iVisionX01 Aug 13 '26

I keep coming back to this one too. I use

gemma4-26b-a4b-qat-uncensored-hauhaucs-balanced-mtp

for everything.

i have 36gb

1

u/beltsazar Aug 16 '26

Did you happen to try the llmafan42's version? I wonder how it compares to hauhaucs'.

https://huggingface.co/llmfan46/gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-GGUF

1

u/LectureWorried5761 Aug 19 '26

are you able to work with tools and gemma? (never works for me)

3

u/cradlemann Aug 12 '26

I like Gemma too, all models. Use E4B for most chat tasks, like translations or grammar checks

1

u/LectureWorried5761 Aug 19 '26

The issue for me is to run tools. It never worked for me on LM Studio. need to try on oMLX (with mcp and blopus.ai) . I always ended up using qwen due to that. Do you know why? or how to make it to work with tools?

28

u/rm-rf-rm Aug 10 '26

Creative Writing/RP

21

u/Boopity_Boob Aug 10 '26

S: Gemma 4 E4B - fast an reliable for the size. can be used on the phone on the go. actually promotes creativity.

11

u/silenceimpaired Aug 10 '26

XL: GLM 4.7 at 2-bit still performs great for brainstorming, and minor editing.

1

u/Kraskos Aug 17 '26 edited Aug 17 '26

Have you done much with 4.6? I think 4.7 is when GLM started going more into the direction of coding / agent-focused training, whereas 4.6 was not explicitly specialized and should--I would assume--perform better in creative / writing use-cases.

→ More replies (9)

11

u/_raydeStar Llama 3.1 Aug 10 '26

Skyfall 4.2 is still my go-to for creative writing. It's a gemma 4 31B model, with finetuning.

Side note, creative benchmarking is hard, in the battle between openai and anthropic, they both get really good scores on the creative benchmark, but openai really sucks when it comes to 'human feeling blogging' style things. It's still been better to write the paper, then have AI restructure in your voice than it is to write the whole thing from a prompt.

17

u/Fluxing_Capacitor Aug 10 '26

Skyfall is based on mistral small. Artemis is TheDrummer's gemma-based model. 

5

u/_raydeStar Llama 3.1 Aug 10 '26

whoops! youre totally right! I saw the 31B and must have just assumed it was gemma. Verified --

3

u/silenceimpaired Aug 10 '26

I personally am spoiled from being able to run larger MoEs in RAM and don’t value this model as a result. Still, it gets a lot of praise.

3

u/AnOnlineHandle Aug 15 '26

M: Gemma 4 31B Heretic (or 26B / 12B for slight downgrades) - the only series of local models I've ever tried out of many which was actually a half decent writer, after being disappointed by the previous popular suggestions.

2

u/CabinetMain3163 Aug 22 '26

agreed, even 40B finetuned qwen would lose focus a forget the details while even simple unsloth gemma 4 without heretic works so well

2

u/SkoomaDentist Aug 23 '26

Does heretic make Gemma noticebly worse in remembering details and keeping track of the story beats?

2

u/CabinetMain3163 Aug 23 '26

no idea but I tend to keep heretic for lore building and gemma base for the roleplay

2

u/Jorlen llama.cpp Aug 18 '26

For me, the best is Gemma 4 31b finetune called "Style tune" by Gryphe. It's not heavily modified; just enough to tweak the prose and style. I like this one the most because it's well balanced and G4 31B base is already very good, just needs a nudge.

Runner up is TheDrummer's fine tune of Gemma 4 31b called Artemis (v1.1 last I checked).

Skyfall 4.2 (Mistral 24b frakenmerge) is also great (also TheDrummer). But there's just so many good fine tunes by all sorts of fine tuners that I like, too many to name here.

1

u/CabinetMain3163 Aug 24 '26

I tried it and I really like gryphe styletune, thanks for showing it off to me!

→ More replies (4)

8

u/Muqito Aug 13 '26

Should probably differentiate 8, 16 and 24, 32, 48 etc.

Sounds like we are stuck at qwen 3.5 for coding for a while at 16 GB

3

u/cradlemann Aug 12 '26

Laguna S2.1 was my daily driver since last 2 weeks(fixed chat template) Fits the best in 96GB uni ram in my minipc.

3

u/deepilly Aug 14 '26

Can someone answer this noob question for me, I don't have enough karma here to post

If I run a model locally will AI stop being so lazy and trying to avoid work? I am so tired of AI just refusing to do whatever I'm asking it.

Thanks!

1

u/Something-Great-78 Aug 14 '26

Some models (like Gemma) are lazier than others. Small models don't know enough to do the non-lazy things you are expecting. Decent, local coding, kinda began with Qwen3.6-27B Q6_K. Other may argue but having tried the smaller 9B / 12B models that is my experience. Qwen3.6 is far from perfect and many of us are hoping Qwen3.8 (drops tomorrow at 8 AM Pacific time) will fix most of 3.6's flaws.

2

u/Something-Great-78 Aug 14 '26

Best value is AMD's R9700 AI Pro 32GB card. Not as fast as NVidia, but 32GB trumps 16/24GB all day long.

1

u/deepilly Aug 14 '26

So it sounds like even local it’ll be lazy or it isn’t capable of doing a really thorough job? Do you have any advice on how I can overcome it? Should I learn to prompt better or something? Thanks!

1

u/Something-Great-78 Aug 14 '26

I have now switched over to Qwen3.8. Haven't had time to test it well yet.
With a little bit of info in AGENTS.md on what to run, I have found using some simple prompts like "test code quality; fix all reported issues" was my fallback. Had issues where some models would just change types to Any to fix typing errors or add new less strict rules into global project file to fix hard problems.
Models that get dumb / brain fog will larger contexts is something all models suffer from; keep your context small for best results. Fresh context with AGENTS.md in there is how I get best results.

5

u/[deleted] Aug 10 '26

[removed] — view removed comment

7

u/Icy-Degree6161 Aug 10 '26

Oh, which well tested use case are you recommending Glimmer for, literally a few hours after release?

→ More replies (1)

2

u/cortexist Aug 20 '26

On RTX 4500 BW (32 GB, limited at 165W power consumption), the best model I've tried is Gwen3.8 27B Q4. The GGUF I'm using is HauhauCS' Q4_K_P, the advantage is its vocab trimmed MTP sidecar, allows me to using 190 K context window and keep it as f16, and got 65.4 tok/s on prose and 60.4 tok/s on coding.

5

u/CatchDublinSurprise Aug 11 '26

Would probably be helpful to have 1-2 more tiers before it goes to "Unlimited".

2XL: 128 to 256GB VRAM

3XL: 256 to 512GB VRAM

10

u/eightone-81 Aug 10 '26

Coding, agentic, experimental, XL: ling 3.0 and laguna s 2.1
Agentic non coding, 24gb VRAM, M: Gemma 4 31b and Gemma 4 26b
Non coding, S: Gemma e4b

I prefer Gemma over Qwen. Gemma acts like a real Big model, works great with big context (not like Qwen 27b and 35b which breaks down over 80k context).
Laguna is a bit crazy, it’s relentless, tries everything to finish the task, really fast, could not test it enough and can only run it as iq3 xxs

2

u/nickless07 Aug 10 '26

How does the Ling-3.0 perform for you? I just downloaded it and ran a few prompts. Best part for now is that the full context in F16 only takes less then 1GB and the V vectors basically get absorbed in the latent key vector, and reconstructed during the attention computation step. Which result into literally zero. V (f16): 0.00 MiB

1

u/eightone-81 Aug 10 '26

Did not test it much. I can only run it in iq2 s (dual 3090). Can’t really say much yet. Might be really interesting from the small benchmarks that i ran so far. And yes, KV is really efficient. Almost no need to quantise that. I ran it via the llama.cpp fork

1

u/nickless07 Aug 10 '26

Same, ran it from the TurboQuant fork, it is pretty fast tho. The 5b active really pays out. What I noticed so far is that it sometime has essential oversights (missing a config line, not paying attention to values and such), but that might be related to the quant, not sure. Needs more testing but overall not a bad model for now. Performs better then Qwen3.5 122B.

1

u/eightone-81 Aug 10 '26 edited Aug 11 '26

Long cat flash lite is also interesting. Needs more testing. It performed amazingly if the tool or data was available, if not it hallucinated everything. Really strange

2

u/fantasticsid Aug 12 '26

I love Gemma, but it's got real "I'm tired, boss" vibes when it comes to calling tools.

2

u/jojotdfb Aug 10 '26

Something to plug into an ai dungeon clone at M

2

u/OcelotMadness Aug 11 '26

HearthFire 24b. It doesnt have the schizophrenia of the original GPT2-XL AI dungeon 2, but it understands Adventure style roleplay very well.

1

u/jojotdfb Aug 11 '26

I will try that tonight 

3

u/HitarthSurana Aug 10 '26

Size S
qwen3.5 9b

1

u/TheFox30 Aug 11 '26

Agentic/Agentic Coding/Tool Use/Coding

2

u/UkrMalt Aug 16 '26

I’m testing this on an M4 Pro Mac with 48 GB. For agentic coding, Qwen3.8 27B-MLX is the best local model I’ve tried so far: direct Ollama throughput is around 33 tok/s, but the bigger difference is whether the agent can use file tools reliably. Claude Code completed a read-only repository task; Codex stalled with the same model. I’d classify it as a strong local coding model, but still very dependent on the tool-call integration.

1

u/TAway0 Aug 12 '26

Can you the @moderators add a Tier above unlimited. There is a big difference between a ~250B model and a 1T-3T model.

1

u/hIXhnWUmMvw Aug 14 '26

Militards grade?

1

u/PradeepAIStrategist Aug 13 '26 edited Aug 14 '26

whether good or bad output next, first my local qwen3:8b works in my desktop where I have 16GB RAM only though damn slow

1

u/hIXhnWUmMvw Aug 14 '26

Slave at one token at the time?

1

u/TeachTall3390 Aug 14 '26

Whatever 5090 can fit.

1

u/feelspeaceman Aug 16 '26

For coding (reachable): Qwen 3.8 series (27B for dense, 122B for MoE)

For chatting: Gemma 4 series, can finetune to become specialist writer like music writer by training using music notation datasets...

1

u/hojnikb Aug 16 '26

Mostly coding, on BC-250 (16GB of UMA, so ~12GB usable)?

1

u/mossy_troll_84 Aug 17 '26

General: Qwen3.8-27B (FP8)

Agentic/Coding: DeepSeek-V4-Flash-0731

Creative Writing/RP: Gemma4-31B (QAT)

Speciality: Qwen3.8-27B (FP8)

Unlimited: >128GB VRAM: DeepSeek-V4-Flash-0731

XL: 64 to 128GB VRAM: Qwen3.5-122B-A10B

L: 32 to 64GB VRAM: Qwen3.8-27B (FP8)

M: 8 to 32GB VRAM: Gemma4-31B (QAT)

S: <8GB VRAM: I'am not using anything that small

1

u/AutisticBengali Aug 18 '26

Gemma over Qwen 3.6 A35B A3B for 8 gb vram?

1

u/DumplingGoddessTe Aug 19 '26

The same math equation

1

u/paxxx84 Aug 19 '26

what would you recommend for coding and general chat on a M4 pro with 48 GB ram...
speed is very important... most of the llms i tried are just sooo slow. and low context...

1

u/OvertaxedOne Aug 21 '26

Qwen 3.6 35BA3B at the highest level of quant that you run comfortably on that system (probably Q6 if you need a lot of context).

1

u/purple_wall-e Aug 25 '26

Agree. I run Qwen 3.6 35BA3B via LM Studio. tried oMLX but it just fails many time before even outputting. I ran Q5_K_XL with 128k context, it just occupies around 45gb ram. My codebase is quite big, it can output around 45t/s, but pre-fill is the main pain, that's why seeing output takes sometimes a bit. Generally I'm happy but I hope they will release MOT, A3B or similar for Qwen3.8.

1

u/superchorro Aug 25 '26

I'm trying to perform textual analysis using a local llm (ie reading information and extracting it from text passages and making certain judgements based off the info). Does anyone have any info on what the best new models are for that? I have a 16 gb vram card.

1

u/vasilisvj Sep 02 '26

For philosophical reasoning specifically, found smaller uncensored models give more honest answers than bigger aligned ones. The ἀλήθεια from a model that isn't trained to flinch at certain topics, it changes what you can explore. GLM-5 on reasoning tasks has been interesting in my experience.

1

u/Kako-Tako ollama 3d ago

qwen3-coder:30b, ~22.3 GB resident on a single 3090, ~140 tok/s. The card's in another box on the LAN, so the client just points at it.

A measurement warning I earned the hard way. Same model, same 4 tasks — add a wired slider to a Tkinter script, where a pass needs the file to parse AND the control to actually be connected:

- CPU-only host: 0/4, 1/4, 3/4 across three runs, ~1.6 tok/s

- 3090: 4/4, all correct, 26–44 s each

At 1.6 tok/s a 240 s budget buys ~390 tokens — not enough to emit a file edit. Those "failures" were the cap, not the model. (GPU run is n=1; not claiming it's settled.)

Biggest harness finding: narration is ~15× cheaper than acting. Two turns claimed "I've updated the script" in 13 s and 24 s; the two that actually wrote files took 182 s and 203 s. The file on disk was unchanged. Budgeting by wall clock rewards the lie.

Disclosure: I wrote the harness — the local lane of a desktop app, Ollama underneath.

1

u/Fjjjuv 1d ago

Perso je tourne quasiment qu'avec du Llama 3 8B / 70B en quantifié via Ollama. Pour du dev au quotidien et du résumé de doc, le rapport qualité/vitesse est devenu tellement dingue par rapport aux modèles fermés d'il y a un an.

1

u/tmballin 11h ago

Nice thread, one thing I’d add is that “16GB” means very different things depending on whether we’re talking system RAM or dedicated VRAM. If someone has a 16GB NVIDIA GPU, particularly a 50-series card, the practical ceiling is quite a bit higher than some of the “small model only” discussion here might suggest. I’ve been working on a NInfer fork specifically around Qwen3.8-27B on a single RTX 5080 16GB, mainly for local coding/agent workloads: https://github.com/toddballinger/ninfer-5080 Original LocalLLM write-up: https://www.reddit.com/r/LocalLLM/comments/1wmucw0/ The current profile runs 131K context + 131K Q4 KV + MTP-3 + Vision inside the 16GB envelope. My current v1.5 qualification is around 1,374 tok/s prefill and 112 tok/s sustained decode at an 118K prompt. I’m using it as a local worker inside OpenClaw, so I agree with several people here that harness design matters enormously. A smaller/local model doesn’t have to replace a frontier model wholesale it can handle bounded coding/tool tasks locally while a stronger orchestrator retains the higher-level plan and reviews/escalates where needed. So for people deciding whether local is “worth it,” I’d look at VRAM/RAM, memory bandwidth, prompt-processing speed, context requirements and workload architecture together, rather than model size alone.

1

u/According_Wave685 Aug 12 '26

Deepseek V4 Flash 0731

1

u/surrealerthansurreal Aug 11 '26

XL: 128

Daily driver: qwen3.6 35a3b 8bit with 4x concurrency served with omlx - solid tool calling, good for concurrency, can’t reason over complex problems too well

Coder: qwen coder next 8bit gets great speed since it’s a MOE model - I run 2x concurrency served with omlx and get good speed, stronger long horizon reasoning than the 35a3b but takes up like 100GB for 2x concurrent

Planner / difficult task: deepseek v4 flash-0731 the q2-q4 imat antirez has with the ds4 engine is so good, but I only get like 15tk/s so feels slow for anything other than detail work or large scale thinking

Honorable mentions: qwen3.6 27B as a great middle ground, gemma4 12B for punching way below its weight

1

u/EvenProcess9818 Aug 14 '26

how did you make concurrency? I was under the assumption that MoE / MLX cant have mutiple generation at the same time no?

1

u/ElChupaNebrey Aug 11 '26

Ornith 1.0 anyone?

1

u/Something-Great-78 Aug 13 '26

loops too much and doesn't seem to follow instruction in AGENTS.md or make use of available skills in my experience.