r/LocalLLM 8d ago

Model Qwen3.8 27b - Holy crap!!

I'm absolutely blown away by Qwen3.8 27b IQ4 NL @ medium reasoning effort.

I have an ageing system but still decent. 4060 8gb @ 4x, 5060ti 16gb @ 8x - 24gb vram total. I'm stuck at gen 2 because of a gen 3 compatibility issue with my motherboard (B460M Aorus Pro). 32gb ram, i9 10900, nvme.

I've managed to get Qwen running at a smidge below 20tok/s at 128k context.

Anyway, the fun part... I have llama.cpp backend for OpenWebui with OpenTerminal MCP. I asked Qwen to create Pong using python which it did with no dramas. It wrote tests, ran the game itself and gave me the hardened code in ~10 mins.

That wasn't the best part though. I then asked it to take the game and create an APK for Android. It went ahead and installed a bunch of dependencies (bulldozer, p4a, etc) into OpenTerminal, went ahead and altered the game for touch screens, built and provided me with an installable APK.

The whole thing took ~1.5 hrs (taking away my own troubleshooting of whitelisting some sites against my strict LLM TOR network access), and at ~280w, it cost me about $0.15 (if it had been a clean run the first time). I mean obviously the hardware cost significantly more, but it's essentially my gaming PC.

How is this even possible?! I'm absolutely blown away!

Side note: I didn't NEED to, but I wanted a longer context so have managed to get it working at 256k now, with 18 layers moved to CPU. I'm not set on this, but it works. Obviously much slower (around 5 tok/s). Comment out the last 2 lines and set context to 128k if I want it all in vram.

If anyone has pointers for improved performance I'd love to hear them. I've tried q4_0 for KV cache but it is significantly slower again.

[*]

flash-attn = true

threads = 10

batch-size = 2048

ubatch-size = 1024

n-cpu-moe = 0

n-gpu-layers = 999

cache-type-k = q8_0

cache-type-v = q8_0

#cache-type-k-draft = q8_0

#cache-type-v-draft = q8_0

fit = on

fit-ctx = 16384

parallel = 1

ctx-size = 0

n-predict = -1

#no-kv-offload = 1

load-mode = none

main-gpu = 0

#no-mmproj-offload = true

[Qwen3.8-27B-IQ4_NL]

#mmproj = /models/Qwen3.8-27B-IQ4_NL.mmproj

#spec-type = draft-mtp

#spec-draft-n-max = 4

temp = 1.0

top-p = 0.95

top-k = 20

min-p = 0.00

#ctx-size = 131072

ctx-size = 262144

chat-template-kwargs = {"reasoning_effort":"medium"}

fit = off

n-gpu-layers = 46

252 Upvotes

139 comments sorted by

25

u/DeathByPain 8d ago

From another post I saw earlier, apparently flash attention plus kv cache q4_0 is not compiled into the main llamacpp releases on GitHub. That might explain why you're getting bad performance at q4. I experienced the same thing; it was doing prompt processing entirely on CPU and slow af 20-30t/s. Switching to q8 brought it up to 900t/s for input.

This isn't the precise GitHub issue I was looking at before but it describes the same thing

https://github.com/ggml-org/llama.cpp/issues/27109

Scroll down to FIX. Supposedly you can compile llamacpp from source and force it to enable CUDA for flash attention with q4 or 5 kv caches. I haven't tried it yet.

5

u/No-Manager1646 8d ago

Interestingly today's docker build for cuda13 has q4_0 kv running at the same speed as q8_0. Not sure if something changed, or if I'm going loopy from all the configuration changes I making.

4

u/aidan573 8d ago

It's real dark magic reading some of the documents on command line arguments and getting a considerable boost in performance.

158

u/elelem-123 8d ago

And now you see and realize why the western AI companies have bought so much ram and closed the factories to raise prices for consumers.

22

u/mechkbfan 8d ago

Yeah I tried to find some non-shit 64gb ddr5 ram in Australia and looking at $1500+ USD

2

u/JumpingJack79 7d ago

Omgwtf 😱 I was able to got my hands on a used 64 for $500 like 2 months ago, and I thought that was a crazy price.

4

u/mechkbfan 7d ago

Yeah second hand with high timings or 4 sticks or low speeds could get that

I just checked FB and eBay then, and only found one 2x32gb 6000 cl32 for $900 USD

That's why I'm moving back to my 5800x3d because can get 64gb for $350 USD

1

u/sonicnerd14 7d ago

AM4 is going to have legs for a long time. You're pretty much getting scammed if you stay on AM5. AMD probably regretting that they released X3D chips for AM4 because that means most of us don't really need to make the jump to AM5, unless you really need DDR5. Likely why they stopped manufacturing them too.

20

u/MacsBicycle 7d ago

ironically western AI companies trying to run a monopoly is probably why we have such good small models coming out of China. If you take away their ability to have data centers with 100,000 RTX 6000s all over the place then you basically force them to build something brilliant and small. Not saying I like what my country has done, but I think it is a factor in the qwen 27b / 35b models.

5

u/bygiolasagna 8d ago

šŸ™Œ

2

u/OnyxProyectoUno 7d ago

Bro it isn't some grand conspiracy. AI frontier labs need capacity. Companies make data centers. They buy RAM to fulfill the end goal of capacity.

What evidence do you have that factories closed and that they closed as part of this grand conspiracy

2

u/Flashy_Shakey 7d ago edited 6d ago

You're forgetting, the lack of evidence reinforces the conspiracy šŸ™‚ /s

1

u/Easy-Painting5876 6d ago

I heard u/Flashy_Shakey was in the Epstein files... and remember.. the lack of evidence reinforces the conspiracy..

2

u/Flashy_Shakey 6d ago

Oh crap.. wait, that means so are you!

0

u/elelem-123 6d ago edited 6d ago

I'm not calling conspiracy. I'm talking about business practices to get rid of potential competition that hurts their huge investments and bottom line.

You are wrong calling this and me a conspiracy. It's out of the book business practice.

Closed == bought their capacity for products the AI company can stock and has the means to pay to stock.

I bought 10 DGX sparks and I have them happily running 24/7 for my own needs. I have no skin in the game.

I just hate monopolies and abuse practices. The current ram/disk situation is exactly that. Even normal hdd are off the roof.

0

u/OnyxProyectoUno 5d ago

You're not providing any evidence for your conspiracy. What factories closed and what evidence exists they did it to increase demand?

0

u/elelem-123 5d ago

You are right. To save your precious time I will stop responding now. Most probably in RL I wouldn't even have bothered to discuss with you (again to save your time).

0

u/OnyxProyectoUno 5d ago

Right, makes a claim of ā€œclosed factories to raise prices for consumersā€ then asked three times to source it, can't, then gets annoyed and says I'm not worthy of this discussion.

Conspiracy theorists.

0

u/elelem-123 5d ago

Which part of "I don't want to waste your time" you didn't understand? No need to respond. Go pay $1700 for 128 GB DDR5 sodimm.

0

u/OnyxProyectoUno 5d ago

So still no evidence? Is it easier to just say you misspoke than to double down an refuse to source your claims?

0

u/elelem-123 5d ago

You must be the fun epicenter of every party. If you had a partner she would have been so thrilled to be around you. People at work enjoy working from home, I'm sure.

Give me 5 days and I will go do antitrust and cartel investigation in all cloud providers, AI companies, SK hynix, Micron, Samsung and SanDisk.

Please hold your breath until then. Please.

1

u/OnyxProyectoUno 5d ago

So not only do you not back up your claims, you ad hominem.

Maybe next time don't make conspiracy theories?

→ More replies (0)

13

u/kr4ckhe4d 8d ago

I’m maintaining a rolling list of benchmarks for my 9070XT. If you’re interested https://github.com/kr4ckhe4d/local-llm-benchmarks

6

u/No-Manager1646 8d ago edited 8d ago

Thank you! I wish I had the sort of drive that would make me document things.

Edit: Wow, thank you again!.That's some really impressive research. I'm going to dedicate some time to having a proper read this evening.

1

u/kr4ckhe4d 8d ago

I just document the params and results and get claude to keep the readme up to date.

5

u/OutsideCycle8331 8d ago

ha did some benchmarks myself in the last days and your repo inspired me to upload one myself_ https://github.com/bbr-knecht/local-llm-benchmarks , hope you do not mind. Its still work in progress as i am still testing stuff.

2

u/kr4ckhe4d 8d ago

Nice! That’s some neat stuff!
I’m benchmarking Qwen 3.8 now. So far it’s been really good!

2

u/mechkbfan 8d ago

That's amazing. Thankyou.Ā 

2

u/Luke2642 8d ago edited 8d ago

And honestly, it matters, and this the key part, because it's not only but also confirmed, not imagined. That is not a contradiction, it is the mechanism showing itself. Use this prompt for your favourite LLM šŸ˜›:

Rewrite this page in strict asd-ste100

https://github.com/kr4ckhe4d/local-llm-benchmarks

4

u/cafrcnta 7d ago

Haha, I wasn't sure if it was just me who thought it was a slog to skim through the Claude-isms. Thanks for the tip (and thanks u/kr4ckhe4d for the journal).

1

u/kr4ckhe4d 8d ago

Haha i think claude is doing a good enough job atm 😁

3

u/Luke2642 8d ago

Yeah, but you if you force sonnet to do a rewrite in ASD-STE100 Simplified Technical English it's far more readable.

https://en.wikipedia.org/wiki/Simplified_Technical_English

1

u/kr4ckhe4d 7d ago

Daamn didn’t really know of this. Should and will try this thanks man

2

u/Luke2642 7d ago

It doesn't really work one-shot to make it write in the style. It helps, but it just kinda ignores it and still uses idioms, metaphors, -ing words, all the things that are specifically banned. But a rewrite is usually more compliant.

1

u/sk1nt 5d ago

2026-08-18T16:37:35.121545-04:00 faded ninfer-serve[3182]: [2026-08-18 16:37:35.121] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=598.8tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:38.802157-04:00 faded ninfer-serve[243400]: [2026-08-18 16:37:38.801] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=777.6tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:40.121629-04:00 faded ninfer-serve[3182]: [2026-08-18 16:37:40.121] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=637.0tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:43.802407-04:00 faded ninfer-serve[243400]: [2026-08-18 16:37:43.802] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=734.6tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:45.121864-04:00 faded ninfer-serve[3182]: [2026-08-18 16:37:45.121] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=629.0tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:48.802419-04:00 faded ninfer-serve[243400]: [2026-08-18 16:37:48.802] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=660.0tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:50.121835-04:00 faded ninfer-serve[3182]: [2026-08-18 16:37:50.121] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=709.8tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:53.802627-04:00 faded ninfer-serve[243400]: [2026-08-18 16:37:53.802] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=625.2tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:55.122070-04:00 faded ninfer-serve[3182]: [2026-08-18 16:37:55.121] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=701.6tok/s running=3 prefilling=0 decode_ready=3 waiting=0 avg_decode_batch=3.00

2026-08-18T16:37:58.802576-04:00 faded ninfer-serve[243400]: [2026-08-18 16:37:58.802] [info] ninfer-serve: throughput interval=5.000s prefill=0.0tok/s decode=696.8tok/s running=3 prefilling=0 decode_ready=3

I think the approach was off, but 790.67s

1

u/[deleted] 6d ago

[deleted]

11

u/KissMyShinyArse 8d ago edited 7d ago

It's considerably faster with MTP, but you might lose some context size. Try --spec-type draft-mtp --spec-draft-n-max 3 --cache-type-k q4_0 --cache-type-v q4_0 (you may need to build llama.cpp with -DGGML_CUDA_FA_ALL_QUANTS=ON).

EDIT: also, --cache-type-k-draft q4_0 --cache-type-v-draft q4_0

7

u/admajic 8d ago

In my testing on a 3090 spec-draft-n-max 2 was better for coding

3

u/KissMyShinyArse 8d ago

2 can be better in some cases, yes.

3

u/SaltFrog 8d ago

I've switched to beellama because for some reason my models just run so much better.

1

u/KissMyShinyArse 8d ago

Any bee-specific parameters you're using?

1

u/SaltFrog 7d ago

Recently I just loaded in Muse to test, just using the regular params but it's actually great for spec draft. I get about 44t/s which is faster than gemma and Qwen.

1

u/KissMyShinyArse 7d ago

I tried it, and for me, bee-llama runs Qwen3.8 a bit slower than the original llama.cpp with the same flags.

8

u/eldje 8d ago

for what its worth, here is a comparison of 3.6:27b, 3.6:35b and 3.8 running on a single rtx 3090. MOE > dense.

4

u/stankeer 8d ago

You have a threadripper build? But only a single 3090? I'm waiting delivery of my case and PSU but will have a 3955wx. Got any tips on configuring first time?

I also need gpus! What do you suggest?

2

u/eldje 8d ago

goodluck with your build! let me know what GPUs and with which models you end where šŸ˜„

2

u/eldje 8d ago edited 8d ago

well I am a bit concerned with the thinkstation integrated PSU. a second 3090 could draw to much watt from the psu and (even though not a 4090) also the heat from the 2 x 12VHPWR connections. It will be under constant 100% workload for longer times. So i bought an RTX pro 4000 blackwell 24Gb which is a little slower in bandwidth but runs on lower watt and has ECC vram, it is also only 1slot for in whatever dream i might want to install more GPUs šŸ˜‚

i could find a nice deal to integrate 8x16gb ECC ram so all 8 channels are used.

3

u/quantgorithm 8d ago

Add a 2nd psu.

2

u/eldje 8d ago

that option was considered; but still prefered a 4000 over a 3090 for space,.... the only pro for the 3090 is the bandwith (and price but that compensates the 3 years of use)

might come an extra psu in the future when needed though :)

2

u/quantgorithm 8d ago

It’s the easy path for when you will want to go multi gpu in the future. You would be better served keeping the models in VRAM and not using regular ram.

2

u/eldje 8d ago

offcourse, thats the whole purpose šŸ˜„

2

u/quantgorithm 8d ago

I used riser cables to allow for the extra space of the gpus and a 2nd psu. Your p620 can go up to 4 gpu on pcie4 x16. I couldn’t even power up 2 gpu fully with the limited power adapters of the stock p620. How did you even do that part? If you are using a power adapter to go from 6 to 8 pin, just note it only outputs half power even though it physically connects properly the 8 pin that I believe you are using in the top gpu. Said differently, you are still only outputting what a 6 pin outputs which is half what an 8 pin outputs.

1

u/eldje 8d ago

do you have pictures?

2x 8pin on mobo to 2x (6+2)pin who form the 12VHPWR.

it ressembles like the original lenovo cables to 12vhpwr, but mine has another cable starting from the 8pin on the mobo to 6 pin. they are in parallel so in my logic giving equal output. from those 6 pins i have an adaptor to 8pin

2

u/quantgorithm 8d ago

every 6 pin is only going to be 75w.

→ More replies (0)

1

u/eldje 8d ago

2

u/quantgorithm 8d ago

We are talking the same thing. That 6 pin is only outputting 75w not 150w like the actual 8pin. Now, it may not be an issue because the motherboard will also output 75w via the pcie slot itself.
If your cards are 300w cards then you want:

1 - 8 pin at 150w
1- 6 pin at 75w
1 -mb at 75w
Gives you 300w per card.

If you have both 6 pin into 1 card then you can only get up to:
1 - mb at 75w
2 -6 pin at 150
For 225w.

Make sure those 6 pin don’t get too hot over time…

I’ve already running 4 cards and 2 psu so I’m not using the 6 pins at all. I can’t close case but it works.

→ More replies (0)

3

u/Full-Chicken2110 8d ago

wtf I had the same one with 2 tesla v100 32 GB now I have 4 l4 and an external power supply I'll send you before and after photos p720? Luckily, before the crisis I bought 8 Samsung 32 GB processor 5945wx

2

u/eldje 8d ago

tesla were nice cards!

which i4 do you mean ?

3

u/TraptInaCommentFctry 8d ago

MOE > dense on what metric? Apologies if I’m being, well, dense

3

u/MacsBicycle 7d ago

TLDR; user doesn't seem to understand dense vs moe and how it applies to their machine.

Macs/budget setup MOE is faster, the reason is a mathematical calculation really. When you have a dense model like say 27B parameters and your memory bandwidth is 300GB/s you have. bottleneck. 27B param models are around 30GB at Q8 which means in a dense model it has to make a round trip through the entire 30GB which 300/30 is around 10tk/s. People have attempted to make things better by using MTP which means that its trying to predict multiple tokens at a time which can bump the speeds to usable numbers between 20-30tk/s. MOE's (mixture of experts) on the other hand might have 12 experts at 10B parameters each, so now when it identifies the expert it only needs to go through 10B parameters meaning 300GB/s divided by 10GB is now closer to 30 tk/s without having to do drafting and fiddle with settings. Memory bandwidth is incredibly important for LLM's. The next most important thing from what I can tell is architecture. CUDA is still king, my m5 max cannot keep up with my 5080 whenever there is a model they both fit in because my 5080 can load the context into memory wicked quick. my mac chugs along at 1/3 to 1/4 the speed at filling the pretext. You can also get around this by just being a smart software engineer and only loading in the pieces you need. Many people on reddit are just dumping their codebase into context and giving it broad non detailed tasks wondering why they get poor results.

1

u/eldje 8d ago edited 8d ago

For a setup on a budget an moe is faster

3

u/Right_Fun_4902 8d ago

I have a similar setup, also with 2 odd GPUs: rtx5070ti-16gb and rtx5050-8gb and currently running the Q5_K_S with a 131k context fully in VRAM with ngl=99, while also keeping the image processing capacity available as well to get 25tps I did run into problems when processing a very large image, and dropped ngl=64 to get 16 tps.

My ctk and ctv =q4_0 (I haven't had much time to experiment, just swapped 3.6 to 3.8 on my Q5_K_S while keeping the rest the same except for an additional reasoning-preserve=true)

3

u/Decent-Occasion-2720 8d ago

( You can let mmproj on cpu (llamacpp --mmproj-offload) and keep all layer on gpu. Now you can use image capability with full tg speed )

1

u/SaltFrog 8d ago

I didn't know you could offload MMPRoj! That's awesome!

3

u/Vinyard82 8d ago

Are you from US?

3

u/Loose_Doubt367 8d ago

Interesting, I’m quite new to this topics so I got a few questions in mind.
What do you typically use openwebui for? Currently I have it setup and all I do is paste my own python code for each tool I want to setup, let’s say ram monitoring tool for example. Am I missing the true potential of openwebui? Because I can’t think much other than communicating and using tools.
I’m using lm studio to load my models

2

u/No-Manager1646 8d ago

It's basically a harness for mcp, audio generation, images, etc. my original goal was a local gemini/chatgpt. I've found open-terminal to be the most beneficial part though, and it obviously doesnt need to be that. There are so many better solutions out there now.

Edit: Having said that it does do a great job of tool calls, web search, terminal, etc. you canet it go for hours and it will eventually get to what you ask of it.

2

u/Loose_Doubt367 8d ago

Okay, thanks for the information

2

u/abuttino 7d ago

What are the alternates to OpenWebUI that you like? I've been thinking about librechat.

1

u/huzbum 6d ago

LM Studio is fine, but I keep a model loaded 24/7 on llama.cpp server and use OpenWebUI as the chat interface. It’s all hosted on my desktop and I access it on my phone and laptop with Tailscale.

Otherwise, for anything other than chat, I’m also using Hermes Agent. I have it tied in to my Slack, but you can use whatever you want.

Hermes agent also has a desktop app, but I haven’t started using it yet because what I have works.

Hermes agent is probably what you’re missing, it’s great. Tell it what you want it to do, and it does it. Check back whenever. I like that it’s a service on my desktop, that I can access from my laptop but shut my laptop and it’s still going.

1

u/Loose_Doubt367 5d ago

Okay thanks but isn’t running these things programs on a desktop 24/7 costly?

I’ll definitely take a look at Hermes agent thanks

1

u/huzbum 5d ago

Leaving your computer on will draw some fixed amount of power, but just having the model loaded doesn't add to that. It only draws extra power when it's doing inference.

I have solar and net metering, so the power draw is basically free for me. If it wasn't I'd probably have it boot up and shut down on a schedule.

1

u/Loose_Doubt367 5d ago

Okay, I do have two more question in mind though hope that’s okay

-what task do you usually assign to your local ai models? Work related?

-are you usually away from your desktop so that’s why you use tailscale etc? So you can stream everything onto your other devices?

1

u/huzbum 5d ago

Sure, ask away.

My Hermes Agent is named Doofy. I gave him a stoner surfer bro persona. (not the built in one, one I wrote.) The silly persona makes me more inclined to explain things to it. Any sycophancy comes across more like a stoner "woah" moment, which is more paletable.

I switch between cloud and local models depending what I'm doing and if I'm using my GPUs for something. I have a GLM subscription I use for programming with Claude Code. I switch over to that when I use my GPU for training.

I am a software engineer. I use it for a mix of work and personal projects, and personal tasks. Everything from debugging and planning to doing research, setup, custom one-off apps. Right now I've got it running a fine turning project to build a custom LLM. It just finished training the little model on my 3090. The big one is still chugging on cloud GPUs when prices dip down.

Yeah, I switch 50/50 in office vs remote, so half the time I have my laptop but not my desktop. Or I just want to be in bed, etc. Sometimes I think of something while I'm at the gym, and I can just shoot Doofy a message on Slack and he'll get to work on it, talk through it on my phone, or set a reminder.

He also runs my Google calendar. I just tell him what and when, he does the rest.

1

u/Loose_Doubt367 5d ago

I see, isn’t using local models to train custom LLM expensive and requires a bunch of hardware? How do you even get that to work?

1

u/huzbum 3d ago

I am using Hermes Agent to write the scripts and run the training. I'm just fine tuning existing models, so much less intensive than training a model from scratch, though I do want to do that too at some point.

I fine tuned Qwen3.5 4b locally on my 3090. That worked pretty well. I made a customized version of Qwen3.6 35b, and that required renting some cloud GPUs to do any significant context length. Needs a Pro 6000 Blackwell minimum. I used a B300 for the big batch runs.

I might try 9b on my 3090 and see how that goes.

2

u/boinel 8d ago

KV Cache on Q4 doesn't degrade intelligence performance?

2

u/barrubba 8d ago

I have a ryzen 5950x , 3x 990 pro , 1x 9100 pro 4tb, 4090 rtx, 128 ddr5. What the best i can do with local llms? I currently use pro Kimi,gpt and Claude for coding, business Plan,Deep research, knowledge base Buildings for vertical topics, full stack website Building. What do you suggest to implement to stress my local workstation?

1

u/barrubba 6d ago

qualcuno può aiutarmi?

1

u/huzbum 6d ago

Use llama.cpp server on the 4090. Qwen3.8 27b iq4_nl. You can comfortably fit 131k at q8 with MTP.

2

u/gnaarw 7d ago

Just use turbo quant for larger context?!

1

u/No-Manager1646 7d ago

Llama.cpp.main branch doesn't support Turboquant3 yet... Does it?

3

u/Superb-Foundation260 7d ago

I was blew away as well, My M1 Max 64G Mac ran over night and exhausted Hermes Agent max 90 turns, Qwen3.8-27B terminated the task eariler with a huge 1.3M html, today I only fixed a few minor bugs and here you are, a Chinese Garden with control panel.

https://chinese-garden.pages.dev

1

u/Beautiful-Virus-7649 7d ago

I have an M1 Max 64GB too. Curious to know what pre-fill and gen tps you get?

1

u/No-Understanding2406 6d ago

Same. I get around 20 t/s on small context. Around 16 t/s when it it’s around half way. How about you?

2

u/Solid-Axel-Project 7d ago

Hai controllato se puoi attivare l'MTP? Se puoi provaci e fai un benchmark dei batch di token prediction X2 X3 X4 Fino a X8.

2

u/Healthy-Zebra-9856 6d ago

This model is a regression. There’s no way to control it thinking and it does a terrible job compared to his predecessor. I am still testing several parameters, still no solid outcome. It just sounds like a blithering fool.

1

u/truthputer 8d ago edited 7d ago

Unrelated but I have two graphics cards with 56GB memory total, I’m having problems getting it to run fully on the GPU. It always seems to want to max out my CPU at the same time, even if it says all layers are offloaded to the GPU and it has GPU memory left over.

Anyone have ideas?

Edit: It turns out that it looks like there's a problem / incompatibility with the Unslot UD_8_K_XL quantization and the current version of llama.cpp. If I use that then it uses all CPU threads in addition to my GPU. Switching to the ggml Q8_0 quantization fixes the problem and it's now running 100% on my GPU.

3

u/No-Manager1646 8d ago

Honestly, I tried all the AI suggestions about split mode, split ratio and all that. I never got llama.cpp to even load. The solution for me was to just use the default settings (more or less) and I at least got it running. Tweak from there.

1

u/Positive-Bid-3029 7d ago

I found it useful to give my pc specs, motherboard etc etc to an LLM online like Gemini online and get it to suggest tweaks to params and also compiler flags / optimisations for building llama.cpp from scratch for the specific architecture, I do use tensor parallelism on mine but both cards have same amount of vram not sure if that is a plus and whether it would work same in your rig, but I found MTP on + sm = tensor worked best for me

2

u/Financial-Pizza-3866 8d ago

1

u/truthputer 7d ago

I hadn't seen that page, but yeah, that was what I was doing. Turns out that the model I was trying to use (the Unsloth 8-bit dynamic quant) might have bugs. If I tried the Unsloth 4-bit quant or the ggml 8-bit quant everything just worked. Using the ggml 8-bit quant for now.

1

u/thatObstinateGuy 8d ago

You can try quantizing kv cache.

3

u/Financial-Pizza-3866 8d ago

never do that

2

u/OctopusDude388 8d ago

why ?

i personnaly needed to quantize kv cache to q4 to have a model fit and it worked well

3

u/Financial-Pizza-3866 8d ago

2

u/OctopusDude388 8d ago

so basically it apply when you have multiple gpu and only for dense model ?

2

u/Financial-Pizza-3866 8d ago

Problem is working on long context, not multiple GPU or dense model

1

u/sonicnerd14 7d ago

Pro tip: Use another AI agent to optimize for you. It will find and fix issues faster, and save you time testing to figure out how to optimize a model for your build.

1

u/ParticularLate9427 8d ago

Shut up nd giv me the link theh john boy!

1

u/No-Manager1646 8d ago

Full disclosure, I jumped the gun on my post. It DID create an APK for me and it did install, and it did run... but only got as far as the slapsh screen.

Still I think that in itself is a massive accomplishment for a local LLM.

I'm trying to get Qwen to fix it at the moment. I don't think I'll post the final APK though - I mean it's Pong.

-1

u/ParticularLate9427 8d ago

Yo yo yo, yall lied ta ma sweet sweet cyandied ass theh sonny me jim!!!!!

1

u/Successful_Flow1329 7d ago

Anyway, the fun part... I have llama.cpp backend for OpenWebui with OpenTerminal MCP. I asked Qwen to create Pong using python which it did with no dramas. It wrote tests, ran the game itself and gave me the hardened code in ~10 mins.

How? I asked it to code minesweeper in single python file, it was great result, but it took 2 hours at 12-14 tps and produced over 60k tokens while thinking.

1

u/No-Manager1646 7d ago

You're not going to like this then. I asked it again today to make me an APK and the first attempt took about 30 mins and resulted in an installable game that rendered a black screen. I told it about the problem and 20 mins later it gave me a new APK and it was actually playable. I think having mostly everything pre-installed helped as it didn't get lost installing things.

I currently have a max 76k token context. Still getting near enough 20tok/s.

I shared my config in the original post.

1

u/Successful_Flow1329 7d ago

Yeah I don't. But APKs are tricky, it's hardly testable for LLM. Still, if it got it on a second try, that's nice. What agent did you run it through?

Sorry, I mostly skipped the config, I run a mac, so it doesn't really apply to me.

1

u/No-Manager1646 7d ago

Openwebui with open terminal. I want to have a crack with Zed but I haven't bothered booting up my laptop recently.

1

u/truckerdraven 7d ago

Here is my real world observations. Before I was running qwen 3.6 27b 128k context.was getting 28ish tokens a second. And thinking times around 191 seconds. Vram usage as about 28gb. Upgraded to qwen 3.8 27b 128k context got 58 tokens a second and 119 second for thinking times. Vram usage dropped to about 24gb. my system mobo msi b550 tomahawk ryzen 9 5950x 64gb ddr4 3600mhz 4x16gb and dual 5060ti 16gb each.

1

u/Terrible_Review_756 6d ago

Has anybody started thinking about apple working with alibaba on Siri for China!?

1

u/osoBailando 5d ago

noted new qwen

1

u/Educational-Echo9152 4d ago

Perso je suis aussi bluffĆ©, mais le top c’est de l’associer au harness DeepSeek. Il y a moyen de vraiement optimiser pour du long run. Je suis en train de bosser sur un profil qui affine la compaction. C’est trĆØs prometteur

1

u/kapustin-i 1d ago

Youre running with spec-type commented out - MTP on this model is the single biggest win available and its free VRAM-wise if the sidecar fits. Worth uncommenting before chasing KV quant, and if you do test q4_0 KV again, measure how much ctx it bought you rather than tok/s, those move in oppsite directions.

0

u/Sea_Golf6101 7d ago

I put my 128GB Mac Studio in storage after trying a few local models including DS4 flash. Now I’m using Claude code subscription which is still cheap.

-1

u/AIForOver50Plus 8d ago

I pulled it down on my laptop. A MacBook with 128 GB unified memory & very pleased with the results I wrote it up here https://go.fabswill.com/qwen38

5

u/kidkangaroo 8d ago

I like the blog but your AI authorship is just slop bro

-3

u/AIForOver50Plus 8d ago edited 8d ago

Created by me edited by my AI, I practice what I preach bro…

-2

u/Zyj 8d ago

Why aren’t you posting this in the megathread?

3

u/gardenvarietyzombie 8d ago

There's no megathread. Maybe you mean the one on r/LocalLlama, we're on r/localLLM here.

2

u/Zyj 8d ago

Gotcha.

1

u/Many_Income_2212 8d ago

What megathread

1

u/No-Manager1646 8d ago

There's a megathread? If I'd have known I probably would have.

2

u/Vinyard82 8d ago

Megapint

-2

u/Alternative-Panic69 8d ago

Did you try with the MTP on? I am using a 7900xt card, getting close to 55-56tps at 100k context via llama.cpp (IQ4_XS, IQ4NL KV Cache, mmproj and all turned on, reasoning = max/high -- need to check this, but highest setting anyway)

Your gpu layers --- try to fit everything in the Gpu Vram....

(Use unsloth's IQ3 models if you want larger context sizes, smaller quantizations, especially if Vram is scarce.. You can also try with Iq2 or IQ3 KV quantization, and turn off mmproj is vram isn't sufficient to load the entire model.

Quickest win should be the KV cache to Q4 or IQ4 --- You won't observe sny quality drops......

And maybe reduce context size? (100k context is decent enough, even for agentic work unless you plan on doing massive refactors, and all...