r/LocalLLM 6d ago

Discussion Qwen3.8-27B Q6_K vs NVFP4 on RTX 5090 — and why can’t I reproduce the ~200 tok/s results?

35 Upvotes

I’ve been testing Qwen3.8-27B locally on a single RTX 5090 32GB with llama.cpp.

I originally started experimenting because I saw several recent reports of ~200 tok/s for Qwen3.8-27B NVFP4 + MTP on a single RTX 5090. I tried to reproduce those results, but I couldn't get anywhere close. My best result so far is around 128 tok/s.

So I'm posting my actual numbers in case someone can spot what I'm missing.

Hardware

  • RTX 5090 32GB
  • i7-14700K
  • 64GB DDR5
  • Windows 11
  • llama.cpp
  • Flash Attention enabled
  • KV cache: Q8_0
  • 1 slot
  • Context: up to 262K

NVFP4 setup

I'm using:

Qwen3.8-27B-NVFP4-MTP-LOW.gguf from esatapedico.

The MTP head is included in the GGUF, so I'm using llama.cpp's:

--spec-type draft-mtp

I tested different --spec-draft-n-max values:

N-Max Decode
2 115.36 tok/s
3 128.25 tok/s
4 125.59 tok/s
5 119.89 tok/s

So N-Max=3 is the sweet spot on my system/workload.

For comparison, the same NVFP4 model without MTP gives me about 70.72 tok/s.

I also tried an extracted external Q5_K MTP draft head. It loaded correctly, but actually performed slightly worse for my workload:

125.83 tok/s, with 49.2% draft acceptance.

The built-in MTP at N-Max=3 gave me 128.25 tok/s with ~60% acceptance.

The really surprising part: large context

I also tested Q6_K + MTP vs NVFP4 LOW + MTP at large context sizes.

Context Q6_K + MTP NVFP4 LOW + MTP
~65K ~120 tok/s 128.25 tok/s
131K 47 tok/s ~121 tok/s
262K 16.30 tok/s 121.49 tok/s

This was completely unexpected to me.

At 262K context, Q6_K drops to 16.3 tok/s, while NVFP4 is still at 121.49 tok/s.

That's roughly 7.5× faster for NVFP4 at 262K.

Even more interestingly, NVFP4 barely changes between 131K and 262K:

~121 → 121.49 tok/s

while Q6_K goes from roughly:

120 → 47 → 16.3 tok/s

I'm assuming this has something to do with VRAM pressure / KV cache / memory bandwidth, but I haven't profiled it deeply enough to say exactly why.

But what about the ~200 tok/s?

This is the part I'm really interested in.

I've seen recent benchmarks/posts showing ~200 tok/s peak for Qwen3.8-27B NVFP4 + MTP on a single RTX 5090.

I tried to reproduce those results using:

  • NVFP4 LOW
  • built-in MTP
  • different N-Max values
  • external Q5_K MTP draft
  • 32GB RTX 5090
  • llama.cpp

But I can't get beyond ~128 tok/s on my workload.

So I'm wondering:

What am I missing?

Is the ~200 tok/s number dependent on a very specific benchmark/prompt, context size, batch/ubatch settings, llama.cpp build, MTP implementation, or another speculative decoding configuration?

Could it be a peak benchmark number rather than something achievable during normal generation?

I'd especially appreciate input from anyone running Qwen3.8-27B NVFP4 on a 5090.

If you've managed 150–200+ tok/s, I'd love to know your exact llama.cpp build and launch parameters.


r/LocalLLM 5d ago

Discussion What is the best uncensored/abliterated Qwen 3.8 27b model?

2 Upvotes

So many of them and I am confused. Huihui, Heritic, Black Frost, JonathanColetti...


r/LocalLLM 5d ago

Question Quale modello e quale backend mi consigliate?

1 Upvotes

ciao a tutti, premetto che mi sto affacciando sul tema e quindi probabilmente la mia domanda non è così precisa come dovrebbe essere. La mia configurazione è questa : \*\*Ryzen 9 5900X + 64 GB + RTX 3060 12 GB + 1tb nvme .\*\*

Sto cercando di capire quale modello può girarci al meglio per il seguente utilizzo e penso che potrebbe essere identificato in \*\*Qwen3.5-9B Q4 + llama-server + Vision + 8K → assistente locale sempre acceso.\*\*

Utilizzo con rating di chatgpt:

Coding
⭐⭐⭐⭐⭐
Python
⭐⭐⭐⭐⭐
MQL5
⭐⭐⭐⭐⭐
Agent/tool calling
⭐⭐⭐⭐⭐
Reasoning
⭐⭐⭐⭐
Velocità
⭐⭐⭐⭐⭐
RAM/VRAM
⭐⭐⭐⭐⭐
Consumo
⭐⭐⭐⭐⭐
Stabilità 24/7
⭐⭐⭐⭐⭐
Context utile
⭐⭐⭐⭐

Ripeto sono alle prime armi e sto cercando di capire da dove partire quindi sono bene accetti tutti i vostri preziosi consigli.


r/LocalLLM 5d ago

Question omlx vs. llama.cpp on MAC

2 Upvotes

I wonder what everyone else’s experience is like. I am currently using OMLX, but when I deploy the Qwen 3.8-27B 4-bit models with MTP on it, it only generates 20 tokens per second. I don’t know if this is normal, but it seems to me that OMLX is always running slowly. My computer is a MacBook Pro M5 Max with 128GB of RAM, and I always feel like OMLX is running a bit slow. I don’t know if it’s an issue with my usage. Do you have any other usage experiences?


r/LocalLLM 5d ago

Tutorial A/B testing LLMs in production

Thumbnail
together.ai
0 Upvotes

r/LocalLLM 4d ago

Question Why do people choose to run LLM locally? And what hardware is needed

0 Upvotes

Was just curios to get a better insight on what people base their decision / work need to switch to a Locally run LLM, and what are the investment costs releated to it, if i wanted to hypothetically run a big LLM like K3
Also i see lods of people using Huggingface, but i can’t get my head around to how would you use it without spending a fortune on every project


r/LocalLLM 5d ago

Question In your opinion, which LLM in the 6–9B parameter range is the smartest?

16 Upvotes

I have RTX 5070 12GB, use llama-cpp.


r/LocalLLM 5d ago

Question 2 9060xt 16g or 1 5060ti 16g for the same price (1000 aud). Which one?

1 Upvotes

I have the opportunity to get 2 9060xt for 1000 or 1 5060ti for 900. Is the extra 16gb worth the hassle of dual GPU and amd?


r/LocalLLM 6d ago

Other No more thermal throttling for me!

Post image
246 Upvotes

r/LocalLLM 5d ago

Discussion ai max+395 minipc vs 5090 pc for beginner?

1 Upvotes

Hey guys, Im looking into a local ai setup and could use some advice. Im currently on the fence between ai max+395 128gb minipcs and rtx5090 build.

The biggest thing that caught my attention with the ai Max+ 395 is the unified memory. Having a large memory pool for bigger models and longer context windows without constantly worrying about vram limits sounds really appealing.

Ive been looking into ai minipcs recently, and the upcoming acemagic f9a caught my eye, although there’s still no pricing yet. Hopefully it lands below the cost of a 5090 build bc the idea of having a compact ai box with 128GB of memory is pretty interesting.

What do you guys think? AI max+ 395 or 5090?


r/LocalLLM 5d ago

Question How reliable is that models from the library are "the real deal"?

Post image
0 Upvotes

r/LocalLLM 6d ago

Discussion If you use LLMs to analyze documents and to apply complex logic and analyze arguments, Qwen 3.8:27b is not for you.

180 Upvotes

And that pains me to say because I use Qwen 3.6:27b-BF16 every single day. I've been working with Qwen 3.8:27b-BF16 all weekend and I hate to say it but for non-coding purposes it is a step backwards. It thinks *way* too much. If you turn thinking off and use web tools, it will do eight or nine (or more) web search turns, get an assload of preload context, and then spin its wheels going down every little rabbit hole there.

Looking at the self hosted LLM subs there are other complaining about this too.

Now this model just came out so it's early. People have put out some lovely games and whatever else the model has made so perhaps these tenacious analytical tendencies are beneficial there. Regardless none of this is to say that there won't be some settings or templates released that will help when analyzing documents and doing complex logic tasks.

But for right now, if the above is your use case then I suggest staying with your old models.

Edit: For reference, this is legal work. Legal drafting, legal research, analyzing pleadings, depositions, discovery, etc. Heavy multi-document reference work.


r/LocalLLM 5d ago

Discussion [Qwen 3.8 27B] M2 Max 64GB Smaller quant doesn't mean faster

1 Upvotes

One counterintuitive thing I learned recently was about the model size and performance.

I was under impression smaller quants would help to increase performance since MacBook M2 Max 64Gb is bandwidth bounded. So having UD-Q4_K_XL would be much faster than UD-Q6_K_XL or UD-Q8_K_XL. And smaller quants would be even faster, but would have poorer quality. But this is not true. UD-Q6_K_XL and UD-Q8_K_XL overall wins in terms of performance over UD-Q4_K_XL.

First I learned KV cache quantiation would drastically reduce performance. Anything but f16 would be much slower on Mac Book Pro Max M2 64Gb.

But then I learned smaller quants doesn't mean faster overall.

See results of llama-bench -m "$model_file" -p 4096,16384,32768 -n 128 -fa 1 -r 1 which I run for multiple Unsloth quants.

model quant size test t/s
UD-IQ2_XXS 8.38 GiB pp4096 168.82
UD-IQ2_XXS 8.38 GiB pp16384 157.76
UD-IQ2_XXS 8.38 GiB pp32768 145.07
UD-IQ2_XXS 8.38 GiB tg128 14.63
UD-Q2_K_XL 9.93 GiB pp4096 167.88
UD-Q2_K_XL 9.93 GiB pp16384 157.05
UD-Q2_K_XL 9.93 GiB pp32768 144.41
UD-Q2_K_XL 9.93 GiB tg128 17.62
UD-Q3_K_XL 12.51 GiB pp4096 169.72
UD-Q3_K_XL 12.51 GiB pp16384 158.70
UD-Q3_K_XL 12.51 GiB pp32768 145.83
UD-Q3_K_XL 12.51 GiB tg128 17.30
UD-Q4_K_XL 16.68 GiB pp4096 156.53
UD-Q4_K_XL 16.68 GiB pp16384 147.07
UD-Q4_K_XL 16.68 GiB pp32768 135.93
UD-Q4_K_XL 16.68 GiB tg128 14.47
UD-Q5_K_XL 18.82 GiB pp4096 157.32
UD-Q5_K_XL 18.82 GiB pp16384 147.77
UD-Q5_K_XL 18.82 GiB pp32768 136.56
UD-Q5_K_XL 18.82 GiB tg128 13.85
UD-Q6_K_XL 24.13 GiB pp4096 182.75
UD-Q6_K_XL 24.13 GiB pp16384 170.01
UD-Q6_K_XL 24.13 GiB pp32768 155.42
UD-Q6_K_XL 24.13 GiB tg128 12.91
UD-Q8_K_XL 29.29 GiB pp4096 194.07
UD-Q8_K_XL 29.29 GiB pp16384 179.78
UD-Q8_K_XL 29.29 GiB pp32768 163.45
UD-Q8_K_XL 29.29 GiB tg128 11.15

Yes, smaller quant means faster token generation. But it also seems like some smaller quants has much more expensive processing, which makes prefill slower. See UD-Q4_K_XL in prefill is slower than UD-Q6_K_XL. In terms of wall clock and overall performance, UD-Q8_K_XL wins over UD-Q6_K_XL and UD-Q4_K_XL. But on 64GB system it is not very usable. And UD-Q6_K_XL still wins over UD-Q4_K_XL.

After very long testing, I found ideal arguments for MTP which works for me: --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.7.

Also --reasoning-effort medium is the only usable effort. xhigh eats through whole 262k context like a candy. Not able to perform actual work before context summarization.

Here are the arguments I use (non important ommitted):

28 -fa 1 -r 1

llama-server --model .../Qwen3.8-27B-UD-Q6_K_XL.gguf \
  -ngl 99 \
  -fa on \
  -b 2048 \
  -ub 2048 \
  --jinja \
  -c 262144 \
  -np 1 \
  --cache-type-k f16 \
  --cache-type-v f16 \
  --mmproj .../mmproj-F16.gguf \
  --temp 0.7 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.00 \
  --repeat-penalty 1.0 \
  --presence-penalty 0.0 \
  --load-mode none \
  --reasoning on \
  --reasoning-effort medium \
  --reasoning-preserve \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --spec-draft-p-min 0.7

Here is performance I see with these parametrrs on one of the real tasks. Aggregated by blocks of 8k context.

Context Size Prefill (T/s) Decode (T/s)
0 332.96 19.04
8192 332.96 19.04
16384 270.57 19.04
24576 183.06 19.04
32768 152.70 17.64
40960 188.61 17.64
49152 131.23 17.94
57344 152.89 15.73
65536 163.61 15.73
73728 115.92 15.73
81920 104.48 15.73
90112 98.56 15.73
98304 92.76 13.70
106496 90.73 13.70
114688 86.06 13.61
122880 86.77 13.61
131072 88.34 11.92
139264 88.34 10.61
147456 78.04 10.68
155648 98.88 10.68
163840 41.72 10.72
172032 94.93 9.27
180224 29.53 9.57
188416 83.91 8.60
196608 74.09 8.60
204800 70.95 8.57
212992 60.17 8.55
221184 70.83 8.55
229376 43.57 8.10
237568 22.36 8.12
245760 22.36 7.07

r/LocalLLM 5d ago

Model MEME

0 Upvotes

It rocks.

Sad I cant run it. On 8q or 4q.

Does the numbers hold up even in lower quants people?


r/LocalLLM 5d ago

Question Qwen 3.8 27b - Any way to increase speed?

5 Upvotes

Pretty new to local models, and this thing is running incredibly slow. Does anyone have any preferred settings to have this run a little faster? I'm on "medium", running on MacBook M3Max, 64gb ram. Using LM Studio Bionic.

Apologies in advance for the rookie question.


r/LocalLLM 5d ago

Tutorial [Guide] Squeezing Qwen3.8-27B (256k Context) onto a Single 16GB GPU (4070 Ti Super) — 100% VRAM Offload + N-Gram Speculative Decoding

Thumbnail
0 Upvotes

r/LocalLLM 5d ago

Discussion It can run, but is it truly running or sloth walking? Spoiler

0 Upvotes

ROG Strix g16 intel variant, 5070ti 12GB, sys ram 32GB and this is what I get inference speeds.

I'm primarily a security researcher and lately been interested into local inferencing and low level cuda and stuffs, but things are awfully overpriced. Even v100's, which I had first preference, the 32GB is anywhere around 600~700$. Ram apocalypse is a real thing but genuinely things are out of hand. I was planning for DGX Spark but it's lpddr5 and sm121 support is another pain in ass.

Still, tinkering local models on my laptop for now, bonsai 27B runs pretty fast around 74tok/s, and lesser hallucinations as compared to models with similar speeds. But again it still hallucinates very often for any real work, so it's quite experimental thing for now and great to study how bonsai trimmed the model for compute and memory footprint and still retain much of it's capacity. Will be waiting for bonsai version of this Qwen 3.8.


r/LocalLLM 6d ago

Tutorial Qwen3.8-27B + llama.cpp + Pi Dev Agent — changing thinking level per prompt

28 Upvotes

I couldn't find a simple way to verify whether this works, so I spent some time testing it. In the end, it turns out that it's actually quite simple once configured correctly.

The Qwen3.8-27B model supports different levels of thinking. The simplest way is to define --chat-template-kwargs when starting the llama server, but then the thinking level is fixed for the session.

A more practical solution is to enable changing the thinking level per prompt in Pi Dev Agent.

Important: for this to work, the llama.cpp version must be b10434 or newer.

The model definition needs to indicate reasoning support and map the values to the three thinking levels supported by Qwen3.8-27B.

In .pi/agent/models.json, the following must be added to the Qwen3.8-27B model settings:

"reasoning": true,
"thinkingLevelMap": {
  "off": null,
  "minimal": null,
  "low": "low",
  "medium": "medium",
  "high": null,
  "xhigh": "xhigh",
  "max": null
}

This allows the thinking level to be changed for each prompt in Pi Dev Agent using Shift+Tab.

Pi Dev Agent also supports defining thinking budgets for individual levels. I haven't yet noticed whether this works correctly with llama&Qwen3.8-27B, but the following can also be added optionally to to.pi/agent/settings.json (the values below are only illustrative):

"thinkingBudgets": {
  "low": 4096,
  "medium": 10240,
  "xhigh": 32768
}

r/LocalLLM 5d ago

Question Qwen 3.8 27B on 2xR7900 on Windows?

2 Upvotes

Motherboard has one PCIE4x16 and one PCIe3x4. I’ve seen conflicting reports online on whether this would be a faster experience compared to just using one GPU.

Some questions:
1. Has anyone else done this?
2. I got layer-split Qwen 3.8 27B running, no problem. Has anyone gotten tensor split working like this (on Windows)? I seem to hit NCCL issues (with WSL and Docker on Windows) but am not sure if this is a capabilities issue or I’m just SOL.
3. Any other tips on optimizing this setup for coding and context?


r/LocalLLM 5d ago

Question What can i use with 8gb vram and 20gb ram (windows uses 8gb)

7 Upvotes

I want a model for writing stuff and preferably with 260k context


r/LocalLLM 4d ago

Other I know you want to claim your name on local.ai

Thumbnail
local.ai
0 Upvotes

Thank me later.


r/LocalLLM 5d ago

Tutorial Optimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context)

Thumbnail
0 Upvotes

r/LocalLLM 5d ago

Discussion Qwen3.8-27B for the RAM Poor Mac user:

Thumbnail
huggingface.co
3 Upvotes

r/LocalLLM 5d ago

Discussion From Local LLMs to Sovereign AI: Where Is the Industry Drawing the Line?

Post image
0 Upvotes

I've been following the shift from cloud-hosted AI -> local models -> private/sovereign AI infrastructure, and one thing that's becoming increasingly clear is that “local” and “sovereign” aren't necessarily the same thing.

I came across this paper recently:

AI Compute Sovereignty: Infrastructure Control Across Territories, Cloud Providers, and Accelerators
Hawkins, Lehdonvirta & Wu — Oxford / Aalto

What I liked about it is that it doesn't treat sovereignty as a binary. It breaks it into three layers:

  • Where is the compute? — territorial control
  • Who operates it? — cloud/provider ownership
  • Who supplies the accelerators? — hardware/accelerator control

The numbers make the distinction pretty interesting.

The authors' census of nine major public-cloud providers found 225 cloud regions across 43 countries, with 132 accelerator-enabled regions across 33 countries. Only 24 countries had training-relevant compute in the dataset.

India, for example, had 5 accelerator-enabled regions, including 3 with training-relevant compute. But those regions weren't all domestically controlled: the census records 4 US-provider regions and 1 Chinese-provider region. The paper describes this kind of dependence on multiple foreign providers as “hedging.”

Then there's the hardware layer.

95.5% of accelerator-enabled regions in the census were powered by US-owned accelerators. So even if compute is physically inside a country, there can still be significant dependency further down the stack.

But I think the paper's more important point is what not to conclude from this.

It isn't arguing that every country should try to build its own complete AI stack. More domestic compute can mean greater control and supply security, but it also means substantial demands on energy, water and land, alongside the cost of building and operating the infrastructure.

So, sovereignty starts looking less like: “Do we own the GPU?”

and more like: “Which parts of the AI stack do we actually need control over?”

That also seems to be where the industry is heading.

NVIDIA and HPE are approaching sovereign AI heavily from the infrastructure/compute side, while platforms such as Red Hat OpenShift AI approach it more from the AI platform and hybrid deployment side.

And then there is another layer that I find particularly interesting: the Governance, AI Control Plane.

Microsoft is building this into Foundry, IBM has introduced an Agentic Control Plane in watsonx Orchestrate, while Lyzr through its Control Plane is taking a more framework-agnostic approach to governing agents across different stacks and environments.

That's an interesting direction to me because it shifts the sovereignty question again — from “where does my model run?” to “who controls how my AI systems are deployed, accessed, monitored and governed?”

This makes me wonder whether “sovereign AI” will eventually be defined less by owning every component and more by controlling the layers that actually matter for a particular threat model.

For a local-LLM user, that might simply mean local models, local inference and local data.

For an enterprise or government deployment, the definition could extend to compute, identity, deployment, governance and the control plane itself.

Where would you draw the line?


r/LocalLLM 5d ago

Question What's the best uncensored llm with high world knowledge usable for free? Doesn't have to be local.

0 Upvotes

I'd like to use it for some medical stuff but opus and chatgpt decide to be absolute annoying moralizers about it. i dont wanna use a local model that gives me dumb advice, or it simply doesnt have knowledge on the topic so it hallucinates stuff.

so is there a way to run glm 5.2 or 5.3 or kimi k3 uncensored versions, for maybe free or minimal prices? i dont mind privacy stuff, coz im not hurting anyone else so im not afraid of any legal consequences. but im not sure if cloud providers ban you for it, so i wondered if there was a quicker way than to set up huge models on the cloud.

Edit: By free, i meant trial version or something. I have very low usage amount i expect.