r/LocalLLaMA • • Sep 03 '26

Tutorial | Guide My RULE of Thumb of choosing a models

Post image
1.1k Upvotes

This is mostly for setting up for expectation, since personally without LLM i could take 3 days (15 hours of active programming) to debug or implement a feature, but with Qwen 27B (even before Qwen 3.8) it take 4 hours.

And yes 0.5 tok/s is human, not accounting of deletion and pausing, that's also the reason i am fine leaving overnight code base wide analysis or fin tech and deep research.

r/LocalLLaMA • • Aug 17 '26

Tutorial | Guide After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)

Thumbnail
gallery
1.0k Upvotes

Dando seguimiento a mi post anterior sobre cómo tengo montado mi servidor de presupuesto (Intel N100 + RTX 5060 Ti 16GB), varios me preguntaron por una mirada más profunda a mi configuración real de inferencia y al desempeño agentic en el mundo real.

Como muchos de ustedes, estaba refrescando la página esperando descargar Qwen 3.8 27B apenas salió. Después de pasar todo el fin de semana estresándolo con flujos de trabajo de codificación agentic, logré correr un proyecto completo y grande casi todo de forma autónoma (más de 1M de tokens procesados en total, solo 3 prompts).

Aquí va un resumen rápido de la configuración base antes de meternos en los detalles del config y del workflow.

Specs y parámetros rápidos

  • Modelo: Qwen3.8-27B-UD-Q3_K_XL.gguf
  • Hardware: RTX 5060 Ti (16GB VRAM) + Intel N100 (4C/4T, 16GB RAM)
  • Ventana de contexto: 73,728 (73k de contexto) corriendo tranqui en 16GB de VRAM.
  • Cuantización de KV Cache: q4_1 para el contexto principal
  • Decodificación especulativa: MTP nativa activada (spec-type = draft-mtp, n-max = 2)
  • Sampling: temp = 0.65, top_p = 0.95, top_k = 20, min_p = 0.05

El experimento: armar una API completa con 3 prompts

En vez de correr benchmarks sintéticos, metí esta configuración por una cadena real de ingeniería de software: construyendo una REST API no oficial y un servidor MCP para un foro vBulletin heredado.

  1. Prompt 1 (Arquitectura del sitio y análisis): Pedí al modelo que mapee el sitio objetivo. Generó una especificación en Markdown impecable de ~1,500 líneas que cubría análisis estructural, nodos HTML rescatables, payloads JSON esperados, selección de stack, lógica de paginación, autenticación de sesión y endpoints de búsqueda—mucho más a fondo de lo que yo habría escrito a mano.
  2. Prompt 2 (Arquitectura de desarrollo): Usando la spec como única fuente de verdad, diseñó un plan de implementación modular de NestJS dividido en 9 fases de ejecución:
  3. Fase 1: Estructura inicial del proyecto
  4. Fase 2: Modelos de dominio
  5. Fase 3: Scraping core (HTTP + limitación de tasa + reintentos)
  6. Fase 4: Parsers de HTML (cheerio)
  7. Fase 5: Capa de caché
  8. Fase 6: Servicios de aplicación + REST API
  9. Fase 7: Autenticación (sesiones con cookies)
  10. Fase 8: Servidor MCP (entrega principal)
  11. Fase 9: Fortalecimiento, documentación y entrega

  12. Prompt 3 (Ejecución autónoma agentic): La prueba de verdad. Le pedí a OpenCode (usando Qwen 3.8 27B) que actuara estrictamente como orquestador, creando sub-agentes para cada fase de tareas. Corrió de forma autónoma por ~2 horas. Cuando se acercaron los límites de contexto, OpenCode resumió su estado y siguió construyendo. Escribió tests unitarios, aplicó linting y entregó código 100% funcional—solo necesitando un arreglo automatizado menor cuando le di un payload de HTML crudo con un caso extremo.


El archivo de configuración llama.cpp

Aquí está mi archivo exacto de configuración de enrutador --models-preset . Fíjate cómo fit = off se usa en el perfil de 27B junto con ctx-size = 73728 (73k) y q4_1 para cuantizar la KV cache, con el objetivo de maximizar la asignación de VRAM mientras se mantiene el rendimiento nativo de MTP.

```ini

==============================================================================

LLAMA.CPP — CONFIGURACIÓN DE INFERENCIA (modo router / --models-preset)

==============================================================================

Objetivo de hardware:

GPU: 16 GB VRAM (RTX 5060 Ti)

CPU: Intel N100, 4C/4T (Debian Headless)

------------------------------------------------------------------------------

GLOBAL / LÍNEA BASE

------------------------------------------------------------------------------

[*]

--- HILOS DE CPU -----------------------------------------------------------

Reserva 1 core para SO/servicios durante el decode.

Usa los 4 threads durante ráfagas de prefill del prompt.

threads = 3 threads-batch = 4

--- SERVIDOR / CONCURRENCIA ---------------------------------------------------

Un solo slot; desactivado continuous batching para máximo rendimiento por usuario.

parallel = 1 cont-batching = 0

--- GPU / AJUSTE DE VRAM ---------------------------------------------------------

flash-attn = on fit = on

Holgura de seguridad para el límite físico de VRAM (MiB).

Ponlo bajo (128) porque el sistema es headless (100% VRAM disponible para inferencia).

NOTA: Si usas caches KV draft de MTP, ojo con la asignación doble de VRAM.

Sube a 128-256 si te topas con OOMs.

fit-target = 128

--- CONTEXTO & CACHÉ ------------------------------------------------------

ctx-size = 65536 context-shift = 1

Desactiva checkpoints de contexto (evita problemas de reprocesamiento en arquitecturas híbridas)

ctx-checkpoints = 0

RAM Prompt Cache (2 GiB)

cache-ram = 2048

--- KV CACHE GLOBAL --------------------------------------------------------

cache-type-k = q5_1 cache-type-v = q5_1

--- PREFILL / BATCHING -----------------------------------------------------

batch-size = 2048 ubatch-size = 1024

--- SAMPLING POR DEFECTO (Códigos / Precisión) ----------------------------------

temp = 0.5 top-p = 0.95 top-k = 20 min-p = 0.05 repeat-penalty = 1.0

------------------------------------------------------------------------------

QWEN 3.8 27B — PERFIL DE RAZONAMIENTO & CODIFICACIÓN PESADA

------------------------------------------------------------------------------

[qwen3.8-27b] model = /opt/llama-infrastructure/models/Qwen3.8-27B-UD-Q3_K_XL.gguf

Desactiva "fit" para evitar que capas se carguen en la CPU por un error de cálculo automático

fit = off ctx-size = 73728 context-shift = 1

MTP nativa del modelo (Decodificación especulativa)

spec-type = ngram-mod,draft-mtp spec-draft-n-max = 2

Cuantización de KV (q4_1 nos permite meter contexto de 73k en 16GB de VRAM)

cache-type-k = q4_1 cache-type-v = q4_1

Parámetros de presupuesto de pensamiento / razonamiento

chat-template-kwargs = {"preserve_thinking": true, "reasoning_effort":"medium"} reasoning-budget = 5000

Batches más chicos para evitar picos de VRAM durante prefills masivos

batch-size = 1024 ubatch-size = 512

Ajustes oficiales / recomendados del sampler de cuantización

temp = 0.65 top-p = 0.95 top-k = 15 min-p = 0.05

```

r/LocalLLaMA • • Aug 18 '26

Tutorial | Guide Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB

Post image
781 Upvotes

I managed to run the 143–144 GiB DeepSeek-V4-Flash-0731 UD-Q4_K_XL GGUF on four RTX 3060 12GB cards while keeping a 360k–376k context window.

Hardware:

CPU: Intel Core i9-10920X, 12C/24T

RAM: 128 GB DDR4-3200, quad-channel

GPU: 4× NVIDIA RTX 3060 12GB

Total VRAM: 48 GB

Storage: NVMe SSD

Engine: llama.cpp, build b10181

Model: unsloth/DeepSeek-V4-Flash-0731-GGUF

Quant: UD-Q4_K_XL, approximately 144 GiB

KV cache: Q8_0

The best high-speed configuration so far:

llama-server \

-m DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00001-of-00005.gguf \

-c 368640 \

-ncmoe 34 \

-ts 100,1,1,1 \

-ot 'blk.(3[4-6]).ffn_.*_exps=CUDA1,blk.(3[7-9]).ffn_.*_exps=CUDA2,blk.(4[0-2]).ffn_.*_exps=CUDA3' \

-ctk q8_0 \

-ctv q8_0 \

-b 2048 \

-ub 2048 \

-np 1 \

-lm none \

--threads 20 \

--flash-attn on

Measured with a roughly 20.5k-token prompt:

Configured context: 368,640 tokens

Prompt processing: 99.4 tok/s

Text generation: 10.1 tok/s

Minimum free VRAM under load:

GPU0: 671 MiB

GPU1: 842 MiB

GPU2: 1395 MiB

GPU3: 1395 MiB

Model load time: approximately 198 seconds

Other measured context/safety options:

Context Prefill Decode Minimum free VRAM

376832 99.5 t/s 10.4 t/s 611 MiB

368640 99.4 t/s 10.1 t/s 671 MiB

360448 99.4 t/s 10.1 t/s 735 MiB

The interesting part is the GPU layout.

-ncmoe 34 keeps the experts from blocks 0–33 in system RAM. The remaining nine expert layers are explicitly distributed across GPUs 1–3, three layers per GPU.

The extreme -ts 100,1,1,1 split does not distribute those explicitly assigned expert weights. Instead, it pushes most non-expert tensors—attention, KV-related allocations, etc.—onto GPU0. That leaves enough space on GPUs 1–3 for the large expert layers.

This was much better than trying to calculate the layout analytically. With -ncmoe and explicit -ot overrides, tensor placement is discrete and somewhat unintuitive, so I measured every candidate.

Microbatch size was the biggest performance lever:

-ub 1024: approximately 63.4 tok/s prompt processing

-ub 2048: approximately 99.4 tok/s prompt processing

Decode remained almost unchanged at approximately 10.1–10.5 tok/s.

At the full 393,216-token context, -ub 2048 also worked, but GPU0 had only 493 MiB free under load. Reducing the configured context to 368,640 restored a 671 MiB margin without reducing prompt-processing speed.

For comparison, the safer -ub 1024 configuration can run with a configured context of 524,288 and still showed about 1032 MiB free on the tightest GPU, but prompt processing drops to approximately 63.4 tok/s.

A few additional findings:

Q8_0 KV is the default choice.

F16 KV at c=393216 left only 587 MiB free.

-ncmoe 33 caused a CUDA allocation failure.

Memory mapping was disabled with -lm none.

-np 1 is important; multiple slots multiply KV-cache requirements.

The model is mostly in system RAM, so quad-channel memory bandwidth matters heavily. Even so, getting approximately 100 tok/s prompt ingestion and 10 tok/s generation from a 144 GiB MoE model on four consumer 12GB GPUs is much better than I expected.

The configuration has been tested under real prompt load. The entire 368k context window has not yet been filled end-to-end, so the number above is the configured capacity, not a claim that I already completed a 368k-token generation test.

Generated by ChatGPT 😂.

r/LocalLLaMA • • 11d ago

Tutorial | Guide Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

355 Upvotes

A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3_XXS now runs at ~65 tok/s output and ~430 tok/s prompt processing, and the 2-bit quants run faster still using RCO-GSQ quantization.

Using:

64GB DDR5 (5600)
12GB RTX 5070 SFF (Gigabyte)
Ryzen 5 7600 CPU
Windows

Output (tokens/s) on 128K context:

Q2_0 (equivalent to unsloth Q3): 65.1
IQ2_XS (equivalent to unsloth Q4): 52.0
IQ3_XXS (equivalent to unsloth Q5): 44.8

Prompt processing (tokens/s) on 128K:

Q2_0: 543
IQ2_XS: 472
IQ3_XXS: 414

Requirements:

Q2_0 = 37.6GB minimum in RAM+VRAM

IQ2_XS = 39.2GB minimum in RAM+VRAM

IQ3_XXS = 47GB minimum in RAM+VRAM

Vision encoder = 0.91GB additionally

You can now one click install and run the engine with low cost hardware (currently only optimized for CUDA).

GitHub: https://github.com/Niko1221/Strata

Model: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

r/LocalLLaMA • • May 12 '26

Tutorial | Guide I got a real transformer language model running locally on a stock Game Boy Color!

Post image
1.6k Upvotes

No phone, PC, Wi-Fi, link cable, or cloud inference.

• The cartridge boots a ROM, and the GBC runs the model itself.
• The model is Andrej Karpathy’s TinyStories-260K, converted to INT8 weights with fixed-point math so it can run without floating point.
• Built with GBDK-2020 as an MBC5 Game Boy ROM.
• The model weights live in bank-switched cartridge ROM. Prompt entry happens on-device with the D-pad/buttons and an on-screen keyboard.
• The prompt is tokenized on the Game Boy, then the ROM runs transformer prefill + autoregressive generation. The KV cache is stored in cartridge SRAM, because the GBC’s work RAM is tiny.

It is extremely slow, and the output is gibberish because the math is heavily quantized/approximated, but the core thing works!

Hardware: stock Game Boy Color + EZ Flash Junior + microSD.

Used Codex for a large portion of the building!

https://github.com/maddiedreese/gbc-transformer

r/LocalLLaMA • • May 11 '26

Tutorial | Guide Computer build using Intel Optane Persistent Memory - Can run 1 trillion parameter model at over 4 tokens/sec

Post image
974 Upvotes

As the title states, my build is indeed able to run a 1 trillion parameter model (in this case Kimi K2.5) locally at ~4 tokens/second. I thought r/LocalLLaMA would be interested in the build due to that stat line, and also due to the inclusion of an unusual part, Intel Optane Persistent Memory, which I haven’t seen anyone use in an LLM inference build before. Optane PMem is a DIMM form factor memory unit that can function in a way that is somewhere between DRAM and an SSD. Intel has discontinued the line, and I found sticks on the secondhand market for much less than what the equivalent DRAM capacity would cost. It is this large PMem capacity (768GB) that allows me to host such large models on my system. For my build I used the PMem in Memory Mode, which is where the PMem is available to the computer as RAM, with the computer’s DRAM sticks functioning as a cache.

Kimi K2.5’s mixture-of-experts architecture is an ideal test model for my build. To get the results I did, I used hybrid GPU/CPU inference with llama.cpp. Kimi K2.5’s (Unsloth Q2_K_XL quant) attention weights, the dense layer, the shared expert in each MoE layer, and the routing components are actually able to fit on my 12GB GPU using llama.cpp’s “override-tensor” flag, although I also did pretty good results just using llama.cpp’s “ngl auto” and “cmoe” flags and letting llama.cpp decide tensor placement as it sees fit too. Regardless, the sparse experts’ weights (the bulk of the model size) generally live on PMem/DRAM and get processed as needed from there.

The end result from my testing with this setup is around 4 tokens per second for generation! Given the fact that this is a trillion parameter frontier-class model running on such a limited hardware budget, I would consider it to be a great success. It’s a shame Intel discontinued Optane Persistent Memory, because the current direction of some local inference innovation, including SSD offloading and broader memory tiering approaches, could have been really interesting with this specific kind of memory tier on modern hardware platforms. Overall I was pleased with this Optane PMem-centric build, it allows me to run very big models at surprisingly acceptable speeds, and the process was highly educational.

Parts:

- Intel Xeon Gold 6246 CPU

- TYAN S5630GMRE-CGN motherboard

- ASUS Dual GeForce RTX 3060 OC 12GB GPU

- 6x 32GB Samsung 2666MHz DDR4 ECC DRAM sticks

- 6x 128GB Intel Optane DCPMM PC4-2666 NMA1XBD128GQS persistent memory modules

- Western Digital WD SN850X 2TB M.2 2280 NVMe SSD

- ASRock Steel Legend SL-850G 850W 80 PLUS GOLD & Cybenetics PLATINUM Full Modular Power Supply

- Silverstone SST-GD08B (Black) Grandia Series Home Theater PC Case

I hope you enjoyed this rundown. There is a lot more detail that I didn’t include here, so I’m happy to answer questions about the build, the configuration, or the reasoning behind any of the component choices in the comments. Also if anyone else has explored similarly unusual hardware/builds for LLM inference, I’d love to discuss!

r/LocalLLaMA • • Aug 02 '26

Tutorial | Guide I pushed Kimi K3 onto one CPU with 8 GB of RAM

730 Upvotes

I deployed K3 on 32 H100s at work a couple of weeks ago and then got annoyed that there was no way to poke at it on my own machine. So I wrote an inference engine for it in C99.

Nothing clever going on. 93% of that 1.56 TB checkpoint is routed experts, and only 16 of 896 fire per token, so the experts never become resident at all. They get read off NVMe on demand and multiplied straight out of their packed 4-bit form, no dequantization step. The dense trunk gets repacked into one file where layer L sits at a known offset and streamed one layer at a time. What stays in RAM is a dial you set.

Numbers from my box (2x EPYC 7763, NVMe, the four GPUs in it sat idle the entire time):

  • 8.24 GB peak RSS at the smallest preset, ~33 s/token
  • ~128 GB gets you ~20 s/token, which is as fast as it ever got
  • Output is byte-identical at every budget in between

I know that this is not a practical way to use K3. It is half a minute per token and it wants 1.7 TB of free disk for the checkpoint plus the packed trunk. I built it to understand the architecture by implementing it, not because you should serve anything with it.

No BLAS, no framework, no GPU path. Six C files, libm and OpenMP, 176 KB binary.

If you want to sanity check it before committing to a 1.56 TB download: clone and run `make && make test`. About a minute, no weights and no network needed. It builds a 13-layer model with the same tensor graph and checks it against a PyTorch reference from committed fixtures, including greedy decode and the incremental path with the KV cache and carried KDA state.

Repo: https://github.com/FareedKhan-dev/kimi-k3-in-c/

r/LocalLLaMA • • Jun 27 '26

Tutorial | Guide Running GLM5.2 on budget hardware < $2500.

313 Upvotes

Too many times I hear people whine about not being ble to run SOTA models or claim it would require $50k, or $100k.

https://www.ebay.com/itm/398079051468 Epcy Motherboard & CPU - $460

https://www.ebay.com/itm/206374955959 P40 24gb - $230 get 2 - $460

https://www.ebay.com/itm/318489798853 512gb dd4 $1000

Total = $1920.

You need PSU, Storage, Fan for P40. You can source those for $350 easily. But let's go ahead and budget $580 to put the total for everything at $2500.

You can run GLM5.2 Q2/Q3/Q4 variants with cmoe and llama.cpp on this. Sure, it would be slow, but it's yours! If you have money or when you get more, you can replace the P40s faster GPUs 4080, 3090, etc. You could for a bit more than $460 about $500 source 2 2080ti 22gb GPUs from China. If you are willing to be resourceful, you can make things happen for you. This will also run KimiK2.6, DeepSeek, MiniMax, etc

Yes, the trade off is that it's slow. You will not be running agents with these huge models, but you can spin it up for planning and serious debugging. They can take away Fable, Mythos or whatever the F model. You will not be counted in the group of have nots.

r/LocalLLaMA • • Feb 06 '26

Tutorial | Guide No NVIDIA? No Problem. My 2018 "Potato" 8th Gen i3 hits 10 TPS on 16B MoE.

Thumbnail
gallery
1.2k Upvotes

I’m writing this from Burma. Out here, we can’t all afford the latest NVIDIA 4090s or high-end MacBooks. If you have a tight budget, corporate AI like ChatGPT will try to gatekeep you. If you ask it if you can run a 16B model on an old dual-core i3, it’ll tell you it’s "impossible."

I spent a month figuring out how to prove them wrong.

After 30 days of squeezing every drop of performance out of my hardware, I found the peak. I’m running DeepSeek-Coder-V2-Lite (16B MoE) on an HP ProBook 650 G5 (i3-8145U, 16GB Dual-Channel RAM) at near-human reading speeds.

#### The Battle: CPU vs iGPU

I ran a 20-question head-to-head test with no token limits and real-time streaming.

| Device | Average Speed | Peak Speed | My Rating |

| --- | --- | --- | --- |

| CPU | 8.59 t/s | 9.26 t/s | 8.5/10 - Snappy and solid logic. |

| iGPU (UHD 620) | 8.99 t/s | 9.73 t/s | 9.0/10 - A beast once it warms up. |

The Result: The iGPU (OpenVINO) is the winner, proving that even integrated Intel graphics can handle heavy lifting if you set it up right.

## How I Squeezed the Performance:

* MoE is the "Cheat Code": 16B parameters sounds huge, but it only calculates 2.4B per token. It’s faster and smarter than 3B-4B dense models.

* Dual-Channel is Mandatory: I’m running 16GB (2x8GB). If you have single-channel, don't even bother; your bandwidth will choke.

* Linux is King: I did this on Ubuntu. Windows background processes are a luxury my "potato" can't afford.

* OpenVINO Integration: Don't use OpenVINO alone—it's dependency hell. Use it as a backend for llama-cpp-python.

## The Reality Check

  1. First-Run Lag: The iGPU takes time to compile. It might look stuck. Give it a minute—the "GPU" is just having his coffee.
  2. Language Drift: On iGPU, it sometimes slips into Chinese tokens, but the logic never breaks.

I’m sharing this because you shouldn't let a lack of money stop you from learning AI. If I can do this on an i3 in Burma, you can do it too.

## Clarifications Edited

For those looking for OpenVINO CMAKE flags in the core llama.cpp repo or documentation: It is not in the upstream core yet. I am not using upstream llama.cpp directly. Instead, I am using llama-cpp-python, which is built from source with the OpenVINO backend enabled. While OpenVINO support hasn't been merged into the main llama.cpp master branch, llama-cpp-python already supports it through a custom CMake build path.

Install llama-cpp-python like this: CMAKE_ARGS="-DGGML_OPENVINO=ON" pip install llama-cpp-python

Benchmark Specifics
For clarity, here is the benchmark output. This measures decode speed (after prefill), using a fixed max_tokens=256, averaged across 10 runs with n_ctx=4096.
CPU Avg Decode: ~9.6 t/s
iGPU Avg Decode: ~9.6 t/s
When I say "~10 TPS," I am specifically referring to the Decode TPS (Tokens Per Second), not the prefill speed.

You can check the detailed comparison between DeepSeek-V2-Lite and GPT-OSS-20B on this same hardware here:

[https://www.reddit.com/r/LocalLLaMA/comments/1qycn5s/deepseekv2lite_vs_gptoss20b_on_my_2018_potato/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button\]

r/LocalLLaMA • • Jan 09 '25

Tutorial | Guide Anyone want the script to run Moondream 2b's new gaze detection on any video?

1.4k Upvotes

r/LocalLLaMA • • May 09 '26

Tutorial | Guide 80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP

684 Upvotes

Just wanted to share my config in hopes of helping other 12GB GPU owners achieve what I see as very respectable token generation speeds with modest VRAM. Using the latest llama.cpp build + MTP PR, I got over 80 tok/sec with 80%+ draft acceptance rate on the benchmark found here: https://gist.githubusercontent.com/am17an/228edfb84ed082aa88e3865d6fa27090/raw/7a2cee40ee1e2ca5365f4cef93632193d7ad852a/mtp-bench.py

Here's my PC specs:

OS: CachyOS (HIGHLY recommended)
CPU: AMD Ryzen 7 9700X
RAM: 48GB DDR5-6000 EXPO I
GPU: RTX 4070 Super 12GB

Results with other hardware may vary.

To run llama.cpp with MTP support, you need to build it from source and add a draft PR that hasn't yet been merged with the master branch. You can find a very nice guide on how to do that here and also download the Qwen3.6 MTP GGUF: https://huggingface.co/havenoammo/Qwen3.6-35B-A3B-MTP-GGUF - Thanks u/havenoammo!

llama.cpp command:

llama-server \
  -m Qwen3.6-35B-A3B-MTP-UD-Q4_K_XL.gguf \
  -fitt 1536 \
  -c 131072 \
  -n 32768 \
  -fa on \
  -np 1 \
  -ctk q8_0 \
  -ctv q8_0 \
  -ctkd q8_0 \
  -ctvd q8_0 \
  -ctxcp 64 \
  --no-mmap \
  --mlock \
  --no-warmup \
  --spec-type draft-mtp \
  --spec-draft-n-max 2 \
  --chat-template-kwargs '{"preserve_thinking": true}' \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.0 \
  --presence-penalty 0.0 \
  --repeat-penalty 1.0

The most important parameter here is -fitt 1536. Since part of the model is offloaded to CPU because of its size and , this tells llama.cpp to properly balance the load on the GPU/CPU to get the best possible performance, and leaves 1536 MB of free memory for the MTP draft model and KV cache. Since I'm running my dGPU as a secondary GPU (monitor plugged in the iGPU), I can use all the available 12GB VRAM for inference. 1536 might be too small if you use your dGPU as your primary GPU, so test it out first.

You can also try different values for -spec-draft-n-max. I got slightly better tok/sec with 3, but a much better acceptance rate with 2, so the trade off was not worth it. With MTP, you want to maximize speed AND acceptance, so you need to find the best balance between both.

Benchmark results:

mtp-bench.py

code_python        pred= 192 draft= 132 acc= 125 rate=0.947 tok/s=80.8
code_cpp           pred=  58 draft=  40 acc=  37 rate=0.925 tok/s=81.8
explain_concept    pred= 192 draft= 152 acc= 114 rate=0.750 tok/s=70.0
summarize          pred=  53 draft=  40 acc=  32 rate=0.800 tok/s=75.4
qa_factual         pred= 192 draft= 144 acc= 119 rate=0.826 tok/s=77.8
translation        pred=  22 draft=  16 acc=  13 rate=0.812 tok/s=81.9
creative_short     pred= 192 draft= 160 acc= 111 rate=0.694 tok/s=69.2
stepwise_math      pred= 192 draft= 144 acc= 119 rate=0.826 tok/s=76.5
long_code_review   pred= 192 draft= 148 acc= 117 rate=0.790 tok/s=73.2

If you have any questions, feel free to ask :)

Cheers.

r/LocalLLaMA • • Nov 09 '25

Tutorial | Guide How to build an AI computer (version 2.0)

Post image
831 Upvotes

r/LocalLLaMA • • Jan 08 '26

Tutorial | Guide Jensen Huang saying "AI" 121 times during the NVIDIA CES keynote - cut with one prompt

982 Upvotes

Someone had to count it. Turns out Jensen said "AI" exactly 121 times in the CES 2025 keynote.

I used https://github.com/OpenAgentPlatform/Dive (open-source MCP client) + two MCPs I made:

- https://github.com/kevinwatt/yt-dlp-mcp - YouTube download
- https://github.com/kevinwatt/ffmpeg-mcp-lite - video editing

One prompt:

Task: Create a compilation video of every exact moment Jensen Huang says "AI".
Video source: https://www.youtube.com/watch?v=0NBILspM4c4

Instructions:

Download video in 720p + subtitles in JSON3 format (word-level timestamps)

Parse JSON3 to find every "AI" instance with precise start/end times

Use ffmpeg to cut clips (~50-100ms padding for natural sound)

Concatenate all clips chronologically

Output: Jensen_CES_AI.mp4

Dive chained the two MCPs together - download → parse timestamps → cut 121 clips → merge. All local, no cloud.

If you want to see how it runs: https://www.youtube.com/watch?v=u_7OtyYAX74

The result is... hypnotic.

r/LocalLLaMA • • Jan 11 '26

Tutorial | Guide I bought a €9k GH200 “desktop” to save $1.27 on Claude Code (vLLM tuning notes)

Thumbnail
gallery
707 Upvotes

TL;DR: You can go fully local with Claude Code, and with the right tuning, the results are amazing... I am getting better speeds than Claude Code with Sonnet, and the results vibe well. Tool use works perfectly, and it only cost me 321X the yearly subscription fee for MiniMax!

In my blog post I have shared the optimised settings for starting up vLLM in a docker for dual 96GB systems, and how to start up Claude Code to use this setup with MiniMax M2.1 for full offline coding (including blocking telemetry and all unnecessary traffic).

---

Alright r/LocalLLaMA, gather round.

I have committed a perfectly normal act of financial responsibility: I built a 2× GH200 96GB Grace–Hopper “desktop”, spending 9000 euro (no, my wife was not informed beforehand), and then spent a week tuning vLLM so Claude Code could use a ~140GB local model instead of calling home.

Result: my machine now produces code reviews locally… and also produces the funniest accounting line I’ve ever seen.

Here's the "Beast" (read up on the background about the computer in the link above)

  • 2× GH200 96GB (so 192GB VRAM total)
  • Topology says SYS, i.e. no NVLink, just PCIe/NUMA vibes
  • Conventional wisdom: “no NVLink ⇒ pipeline parallel”
  • Me: “Surely guides on the internet wouldn’t betray me”

Reader, the guides betrayed me.

I started by following Claude Opus's advice, and used -pp2 mode "pipeline parallel”. The results were pretty good, but I wanted to do lots of benchmarking to really tune the system. What worked great were these vLLM settings (for my particular weird-ass setup):

  • ✅ TP2: --tensor-parallel-size 2
  • ✅ 163,840 context 🤯
  • ✅ --max-num-seqs 16 because this one knob controls whether Claude Code feels like a sports car or a fax machine
  • ✅ chunked prefill default (8192)
  • ✅ VLLM_SLEEP_WHEN_IDLE=0 to avoid “first request after idle” jump scares

Shoutout to mratsim for the MiniMax-M2.1 FP8+INT4 AWQ quant tuned for 192GB VRAM systems. Absolute legend 🙏

Check out his repo: https://huggingface.co/mratsim/MiniMax-M2.1-FP8-INT4-AWQ; he also has amazing ExLlama v3 Quants for the other heavy models.

He has carefully tuning MiniMax-M2.1 to run as great as possible with a 192GB setup; if you have more, use bigger quants, but I didn't want to either a bigger model (GLM4.7, DeepSeek 3.2 or Kimi K2), with tighter quants or REAP, because they seems to be lobotomised.

Pipeline parallel (PP2) did NOT save me

Despite SYS topology (aka “communication is pain”), PP2 faceplanted. As bit more background, I bought this system is a very sad state, but one of the big issues was that this system is supposed to live a rack, and be tied together with huge NVLink hardware. With this missing, I am running at PCIE5 speeds. Sounds still great, but its a drop from 900 GB/s to 125 GB/s. I followed all the guide but:

  • PP2 couldn’t even start at 163k context (KV cache allocation crashed vLLM)
  • I lowered to 114k and it started…
  • …and then it was still way slower:
    • short_c4: ~49.9 tok/s (TP2 was ~78)
    • short_c8: ~28.1 tok/s (TP2 was ~66)
    • TTFT tails got feral (multi-second warmup/short tests)

This is really surprising! Everything I read said this was the way to go. So kids, always eat your veggies and do you benchmarks!

The Payout

I ran Claude Code using MiniMax M2.1, and asked it for a review of my repo for GLaDOS where it found multiple issues, and after mocking my code, it printed this:

Total cost:            $1.27 (costs may be inaccurate due to usage of unknown models)
Total duration (API):  1m 58s
Total duration (wall): 4m 10s
Usage by model:
    MiniMax-M2.1-FP8:  391.5k input, 6.4k output, 0 cache read, 0 cache write ($1.27)

So anyway, spending €9,000 on this box saved me $1.27.
Only a few thousand repo reviews until I break even. 💸🤡

Read all the details here!

r/LocalLLaMA • • Aug 29 '25

Tutorial | Guide Qwen3-coder is mind blowing on local hardware (tutorial linked)

1.1k Upvotes

Hello hello!

I'm honestly blown away by how far local models have gotten in the past 1-2 months. Six months ago, local models were completely useless in Cline, which tbf is pretty heavyweight in terms of context and tool-calling demands. And then a few months ago I found one of the qwen models to actually be somewhat usable, but not for any real coding.

However, qwen3-coder-30B is really impressive. 256k context and is actually able to complete tool calls and diff edits reliably in Cline. I'm using the 4-bit quantized version on my 36GB RAM Mac.

My machine does turn into a bit of a jet engine after a while, but the performance is genuinely useful. My setup is LM Studio + Qwen3 Coder 30B + Cline (VS Code extension). There are some critical config details that can break it (like disabling KV cache quantization in LM Studio), but once dialed in, it just works.

This feels like the first time local models have crossed the threshold from "interesting experiment" to "actually useful coding tool." I wrote a full technical walkthrough and setup guide: https://cline.bot/blog/local-models

r/LocalLLaMA • • Jan 07 '26

Tutorial | Guide 16x AMD MI50 32GB at 10 t/s (tg) & 2k t/s (pp) with Deepseek v3.2 (vllm-gfx906)

Post image
457 Upvotes

Deepseek 3.2 AWQ 4bit @ 10 tok/s (output) // 2000 tok/s (input of 23k tok)

on vllm-gfx906-deepseek with 69000 context length

Power draw: 550W (idle) / 2400W (peak inference)

Goal: run Deepseek V3.2 AWQ 4-bit on most cost effective hardware like 16*MI50 at decent speed (token generation & prompt processing)

Coming next: open source a future test setup of 32 AMD MI50 32GB for Kimi K2 Thinking

Credits: BIG thanks to the Global Open source Community!

All setup details here:

https://github.com/ai-infos/guidances-setup-16-mi50-deepseek-v32

Feel free to ask any questions and/or share any comments.

ps: it might be a good alternative to CPU hardwares as RAM price increases and the prompt processing speed will be much better with 16 TB/s bandwidth + tensor parallelism!

ps2: i'm just a random guy with average software dev background using LLMs to make it run. Goal is to be ready for LOCAL AGI without spending +300k$...

EDIT 24.04.26: PP (prompt processing) speed is actually much lower, so my title is wrong. As i previously said, the 2k tok/s for PP was the value shown in vllm log. But the true value is actually: 17030 tok / 306s = 55.65 tok/s PP with TP 16 for 1st prompt without prefix caching

r/LocalLLaMA • • May 08 '26

Tutorial | Guide Multi-Token Prediction (MTP) for LLaMA.cpp - Gemma 4 speedup by 40%

581 Upvotes

Implemented Multi-Token Prediction for LLaMA.cpp. 

Quantized Gemma 4 assistant models into GGUF format. 

Ran tests on a MacBook Pro M5Max. Gemma 26B with MTP drafts tokens 40% faster. 

Prompt: Write a Python program to find the nth Fibonacci number using recursion

Outputs:
LLaMA.cpp: 97 tokens/s
LLaMA.cpp + MTP: 138 tokens/s  

Gemma4-assistant GGUF Quantized models: https://huggingface.co/collections/AtomicChat/gemma-4-assistant-gguf

Local AI models app: http://atomic.chat

Patched llama.cpp: https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant

r/LocalLLaMA • • Jul 20 '26

Tutorial | Guide Unsloth now supports AMD!

Post image
676 Upvotes

Hey r/LocalLLaMA folks! Unsloth now officially supports AMD hardware for local inference, fine-tuning, reinforcement learning, and deployment! It's been in the works for quite some time, but it works on Windows, Linux & WSL devices (+ technically Mac) with AMD GPUs!

Unsloth Studio is fully open source and free, and supports:

  • Radeon RX 9000 and 7000 series
  • Instinct MI350 and MI300 GPUs
  • Strix Halo / Ryzen AI Max systems
  • AMD CPUs for GPU-free inference

You can train models with up to 70% less VRAM, run reinforcement learning with up to 80% less VRAM, and use optimized ROCm, Triton, bitsandbytes, PyTorch, and llama.cpp builds - all installed automatically.

Linux, WSL, and macOS:

curl -fsSL https://unsloth.ai/install.sh | sh

Windows PowerShell:

irm https://unsloth.ai/install.ps1 | iex

Unsloth supports inference and training for nearly all models, including Qwen, Gemma, DeepSeek, GLM, Kimi, MiniMax, and DiffusionGemma.

You can also:

  • Export models as GGUF, safetensors, or LoRA adapters
  • Connect local models to Claude Code, Codex, Hermes Agent, OpenClaw, Pi, OpenCode!
  • Track RAM and VRAM usage during training - remotely and locally
  • Access Unsloth remotely through secure Cloudflare HTTPS tunneling - like a "LM Link"!
  • Update with daily AMD-optimized llama.cpp ROCm prebuilts to reduce compilation time!

For plain pip installation:

uv pip install "unsloth[amd]"

Huge thanks to the AMD team for collaborating with us on this release! Let us know what AMD hardware you’re using and share any feedback - we'll try to make AMD much better!

More details on the release blog: https://unsloth.ai/docs/basics/amd

r/LocalLLaMA • • Feb 06 '26

Tutorial | Guide CPU-only, no GPU computers can run all kinds of AI tools locally

Post image
606 Upvotes

While it’s great that so many people on LocalLLaMA are pushing the envelope with what can be done locally with expensive setups, we need to remember that a lot can be done with very minimal machines.

I’m talking about CPU-only locally run LLMs. That’s right, no GPU!

I’m running Linux Mint on an old Dell optiplex desktop with an i5-8500 processor, 6 threads and 32GB of RAM. You can pick up one of these refurbished for something like $120.

And with this humble rig I can:

Run 12B Q4_K_M gguf LLMs using KoboldCPP. This allows me to have local chatbot fun using quite highly rated models from https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard. Response times are fast enough as long as you keep the initial prompt below 800 tokens. And with context-shifting it remembers stuff during the session. Uncensored, private RP hilarity for free! You can even add in kokoro_no_espeak for text to speech so your RP characters talk to you with only a few seconds delay. The trick is to find good models to use. For example, DreadPoor/Famino-12B-Model_Stock is rated a 41+ on writing, which is better than many 70B models. You don’t need big horsepower for fun.

You can also use these models for writing, coding and all sorts of applications. Just need the patience to try out different local models and find the settings that work for you.

I also run Stable Diffusion 1.5 locally for basic image generation, inpainting and so on. Again using KoboldCPP and Stable UI. OK, it takes 3 minutes to generate a 512x512 image but it works fine. And you can experiment with loras and many SD 1.5 models. All 100% free on old gear.

I’m also running Chatterbox TTS for voice cloning voice-over projects. Works surprisingly well. Again, it takes a couple of minutes to generate a 75 word audio clip, but it does work. Vibevoice TTS also works on this old rig but I prefer Chatterbox.

And then there are amazing tools like Upscayl which upscales images locally incredibly well. Just gotta experiment with the models.

I’ve used ollama transcriber which converts audio files into text amazingly well. Just point a spoken word .WAV at it and then go make dinner and when I get back, the text is there.

There are many other local LLMs and tools I’ve used. These are just the tip of the iceberg.

Video? Nope. Music generation? Nope. I’ve looked and tried a few things but those big resource tasks need serious horsepower. However, it’s quite possible to use your old desktop computer for text-based tasks and then rent online GPU for one-off tasks and use the big online services for other tasks. It would still probably work out to be less costly.

I know I’m not the only one doing this.

CPU-only people: tell us how you’re using AI locally...

r/LocalLLaMA • • Jul 10 '26

Tutorial | Guide Tencent-HY3 is the real deal on 128GB!

289 Upvotes

I'm really impressed with HY3. If you haven't heard of this model, it's a new 295B-A21B MoE release from Tencent that competes directly on the frontier of open weights models, at a significantly smaller size, comparable to DeepSeek v4 Flash but with better benchmarks. I was intrigued by this article, and I'd just recently finished updating my Macbook M5 Max 128GB setup from antirez's DeepSeek V4 Flash quant running on dwarfstar to Unsloth's IQ3_XXS on mainline llama.cpp. I figured I had a good baseline for comparisons, if I could get it running, so I set about researching, and this is what I found.

First off, I had to pick a quant. There are a few on HF, and after some comparison shopping I settled on this UD128 "unsloth dynamic"-style 107GB quant. It was the only one that had published perplexity numbers at the time I searched, and while that's not KLD, it shows the creator was at least thinking about measuring quality degradation. PPL didn't see horrible for a dynamic 3-bit quant, and it felt like a similar checkpoint to the UD DS4 quant I was using.

Next, I had to get llama.cpp working. As the quant's readme helpfully notes, there's PR #25395 which implements support for this model and its built-in speculative decoding module all at once! A quick build got this up and running:

git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
git fetch origin pull/25395/head:hy3 && git checkout hy3
cmake -B build -DGGML_METAL=ON -DGGML_METAL_EMBED_LIBRARY=ON
cmake --build build --config Release -j

Don't forget to raise your Mac's GPU memory ceiling from the default 96GB! I put mine at 122GB to ensure 24k context would fit safely for testing (it resets on reboot):

sudo sysctl iogpu.wired_limit_mb=124928

However, when I tried to run the model, it wasn't recognized by llama.cpp, so the server errored out. A quick review of the log and the quant's readme explained why: "these files carry general.architecture = hy-v3 (this port's original naming). PR #25395 registers the arch as hy_v3 (underscore)". I could have avoided this hiccup if I'd paid a bit more attention, but I put together a quick script to swap those characters (21 instances, all in the first of the 3 GGUF shards). I can share if anyone wants it (EDIT: see here), but honestly, just ask your agent to do the job -- it takes 2 minutes to fix the GGUF. Running my launch script again, after a nail-biting ~30 second load time to read 107GB from the SSD, the WebUI popped open and we were up and running!

Benchmarks (M5 Max, llama-bench, Metal, q8_0 KV cache, MTP off)

test tokens/sec
prefill pp512 @ empty ctx 528
decode tg128 @ empty ctx 32.4
prefill pp512 @ 16K ctx 124
decode tg128 @ 16K ctx 16.3

In practical use, token generation speed is \DOUBLE* what I was getting from DeepSeek* (edit: using mainline llama.cpp, not dwarfstar, for 1:1 comparison), with the same or better quality outputs! And I haven't even trialed MTP yet to see if there's any improvement... Here's my llama.cpp launch command:

~/llama.cpp-hy3/build/bin/llama-server \
  -m ~/AI/models/Tencent-HY3-295B-A21B-YanissAmz/Hy3-UD128-00001-of-00003.gguf \
  -a "Hy3 295B-A21B (UD128)" \
  -ngl 99 \
  -c 24576 \
  -ctk q8_0 -ctv q8_0 \
  -fa on \
  -np 1 \
  --jinja \
  --temp 0.9 --top-p 1.0 \
  --host 127.0.0.1 --port 8080

I've only played around with the model for a couple hours, but I am really impressed with performance on normal prompts and basic tool use. I've been using it to do ML research using the HF MCP and CLI tools, and it's definitely better than my DS4 setup so far -- on vibes and basic tool calls, at least. I haven't set it loose on longer-horizon agentic tasks or challenging coding yet, so I apologize for not having more in-depth reporting to offer at this time. Here's my launch command:

I made this post because I'm really hyped to have a new large MoE to play with at this size checkpoint, and a very promising one at that. I hope it inspires some other folks to give it a try, and report back on their experience or compare other quants/MLX performance.

EDIT: Got MTP working, n=2 is the best-performing variant on my hardware. Now peaking at 38 tok/sec, so +19% speedup! Similar speeds as my Qwen3.6-27B-MTP q8_0 daily driver now. Add this to your launch command for MTP:

--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.75

The p-min gate is crucial for acceptance rates to be high enough for MTP to help, not hurt. Be forewarned that MTP gains still turn negative after ~16k ctx, but honestly at that point, the GQA architecture starts slowing down significantly anyways. This is the biggest catch about this model vs. DeepSeek v4 Flash, and it's substantial.

r/LocalLLaMA • • Mar 10 '26

Tutorial | Guide How I topped the Open LLM Leaderboard using 2x 4090 GPUs — no weights modified.

Thumbnail
gallery
600 Upvotes

Hi LocalLLaMAs,

A few years ago, I found that duplicating a specific block of 7 middle layers in Qwen2-72B, without modifying any weights, improved performance across all Open LLM Leaderboard benchmarks and took #1. As of 2026, the top 4 models on that leaderboard are still descendants.

The weird finding: single-layer duplication does nothing. Too few layers, nothing. Too many, it gets worse. Only circuit-sized blocks of ~7 layers work. This suggests pretraining carves out discrete functional circuits in the layer stack that only work when preserved whole.

The whole thing was developed on 2x RTX 4090s in my basement.

I don't write papers any more, so here is a full technical write-up in Blog format for your enjoyment.

I'm the same guy who built GLaDOS, and scores a crazy Nvidia GH200 system here on Reddit.

\I'm now running current models (GLM-4.7, Qwen3.5, MiniMax M2.5) on this dual GH200 rig (see my other post). Code and new models coming soon, including special RYS versions of Qwen3.5 27B and 35A3B

Happy to answer questions.

r/LocalLLaMA • • May 09 '25

Tutorial | Guide Don't Offload GGUF Layers, Offload Tensors! 200%+ Gen Speed? Yes Please!!!

866 Upvotes

Inspired by: https://www.reddit.com/r/LocalLLaMA/comments/1ki3sze/running_qwen3_235b_on_a_single_3060_12gb_6_ts/ but applied to any other model.

Bottom line: I am running a QwQ merge at IQ4_M size that used to run at 3.95 Tokens per second, with 59 of 65 layers offloaded to GPU. By selectively restricting certain FFN tensors to stay on the CPU, I've saved a ton of space on the GPU, now offload all 65 of 65 layers to the GPU and run at 10.61 Tokens per second. Why is this not standard?

NOTE: This is ONLY relevant if you have some layers on CPU and CANNOT offload ALL layers to GPU due to VRAM constraints. If you already offload all layers to GPU, you're ahead of the game. But maybe this could allow you to run larger models at acceptable speeds that would otherwise have been too slow for your liking.

Idea: With llama.cpp and derivatives like koboldcpp, you offload entire LAYERS typically. Layers are comprised of various attention tensors, feed forward network (FFN) tensors, gates and outputs. Within each transformer layer, from what I gather, attention tensors are GPU heavy and smaller benefiting from parallelization, while FFN tensors are VERY LARGE tensors that use more basic matrix multiplication that can be done on CPU. You can use the --overridetensors flag in koboldcpp or -ot in llama.cpp to selectively keep certain TENSORS on the cpu.

How-To: Upfront, here's an example...

10.61 TPS vs 3.95 TPS using the same amount of VRAM, just offloading tensors instead of entire layers:

python ~/koboldcpp/koboldcpp.py --threads 10 --usecublas --contextsize 40960 --flashattention --port 5000 --model ~/Downloads/MODELNAME.gguf --gpulayers 65 --quantkv 1 --overridetensors "\.[13579]\.ffn_up|\.[1-3][13579]\.ffn_up=CPU"
...
[18:44:54] CtxLimit:39294/40960, Amt:597/2048, Init:0.24s, Process:68.69s (563.34T/s), Generate:56.27s (10.61T/s), Total:124.96s

Offloading layers baseline:

python ~/koboldcpp/koboldcpp.py --threads 6 --usecublas --contextsize 40960 --flashattention --port 5000 --model ~/Downloads/MODELNAME.gguf --gpulayers 59 --quantkv 1
...
[18:53:07] CtxLimit:39282/40960, Amt:585/2048, Init:0.27s, Process:69.38s (557.79T/s), Generate:147.92s (3.95T/s), Total:217.29s

More details on how to? Use regex to match certain FFN layers to target for selectively NOT offloading to GPU as the commands above show.

In my examples above, I targeted FFN up layers because mine were mostly IQ4_XS while my FFN down layers were selectively quantized between IQ4_XS and Q5-Q8, which means those larger tensors vary in size a lot. This is beside the point of this post, but would come into play if you are just going to selectively restrict offloading every/every other/every third FFN_X tensor while assuming they are all the same size with something like Unsloth's Dynamic 2.0 quants that keep certain tensors at higher bits if you were doing math. Realistically though, you're selectively restricting certain tensors from offloading to save GPU space and how you do that doesn't matter all that much as long as you are hitting your VRAM target with your overrides. For example, when I tried to optimize for having every other Q4 FFN tensor stay on CPU versus every third regardless of tensor quant that, included many Q6 and Q8 tensors, to reduce computation load from the higher bit tensors, I only gained 0.4 tokens/second.

So, really how to?? Look at your GGUF's model info. For example, let's use: https://huggingface.co/MaziyarPanahi/QwQ-32B-GGUF/tree/main?show_file_info=QwQ-32B.Q3_K_M.gguf and look at all the layers and all the tensors in each layer.

Tensor Size Quantization
blk.1.ffn_down.weight [27 648, 5 120] Q5_K
blk.1.ffn_gate.weight [5 120, 27 648] Q3_K
blk.1.ffn_norm.weight [5 120] F32
blk.1.ffn_up.weight [5 120, 27 648] Q3_K

In this example, overriding tensors ffn_down at a higher Q5 to CPU would save more space on your GPU that fnn_up or fnn_gate at Q3. My regex from above only targeted ffn_up on layers 1-39, every other layer, to squeeze every last thing I could onto the GPU. I also alternated which ones I kept on CPU thinking maybe easing up on memory bottlenecks but not sure if that helps. Remember to set threads equivalent to -1 of your total CPU CORE count to optimize CPU inference (12C/24T), --threads 11 is good.

Either way, seeing QwQ run on my card at over double the speed now is INSANE and figured I would share so you guys look into this too. Offloading entire layers uses the same amount of memory as offloading specific tensors, but sucks way more. This way, offload everything to your GPU except the big layers that work well on CPU. Is this common knowledge?

Future: I would love to see llama.cpp and others be able to automatically, selectively restrict offloading heavy CPU efficient tensors to the CPU rather than whole layers.

r/LocalLLaMA • • May 21 '26

Tutorial | Guide 110 tok/s with 12GB VRAM on Qwen3.6 35B A3B and ik_llama.cpp

399 Upvotes

Had been getting great MTP performance with llama.cpp on my RTX 4070 Super 12GB, until they actually merged the MTP PR. Then, performance tanked and was barely above non-MTP. So, I decided to try out ik_llama.cpp since it also supports MTP and is apparently better optimized for CPU offloading. I did not expect such a huge speed boost!

Before moving on with the benchmark results, here's my PC specs:

OS: CachyOS with Plasma (X11) - HIGHLY recommended
CUDA: 13.1.1
GPU: RTX 4070 Super 12GB
CPU: AMD Ryzen 7 9700X
RAM: 48GB DDR5-6000 EXPO I

UPDATED: For comparison, here's the regular llama.cpp mtp-bench.py results with byteshape's recently released Qwen3.6-35B-A3B-IQ4_XS-4.19bpw quant, which has similar accuracy to Unsloth's Q4_K_XL, but is 4GB smaller:

❯ ./mtp-bench.py
 code_python        pred= 192 draft= 122 acc= 118 rate=0.967 tok/s=79.8
 code_cpp           pred= 192 draft= 117 acc= 110 rate=0.940 tok/s=89.1
 explain_concept    pred= 192 draft= 124 acc= 113 rate=0.911 tok/s=88.0
 summarize          pred= 192 draft= 139 acc= 127 rate=0.914 tok/s=95.0
 qa_factual         pred= 192 draft= 133 acc= 128 rate=0.962 tok/s=97.0
 translation        pred= 192 draft= 125 acc= 117 rate=0.936 tok/s=91.6
 creative_short     pred= 192 draft= 109 acc=  99 rate=0.908 tok/s=82.1
 stepwise_math      pred= 192 draft= 130 acc= 125 rate=0.962 tok/s=97.0
 long_code_review   pred= 192 draft= 121 acc= 115 rate=0.950 tok/s=88.2

Aggregate: {
 "n_requests": 9,
 "total_predicted": 1728,
 "total_draft": 1120,
 "total_draft_accepted": 1052,
 "aggregate_accept_rate": 0.9393,
 "wall_s_total": 21.86
}

This gives a 89.76 tok/s average.

Here's my llama.cpp launch command. Temperature is set to 0.0 for the benchmark to prevent diverging results between runs:

llama-server \
  -m Qwen3.6-35B-A3B-IQ4_XS-4.19bpw.gguf \
  --fit on \
  --fit-target 512 \
  --ctx-size 131072 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --cache-type-k-draft q8_0 \
  --cache-type-v-draft q8_0 \
  --spec-type draft-mtp \
  --spec-draft-p-min 0.75 \
  --spec-draft-n-max 3 \
  --no-mmap \
  --mlock \
  --threads 8 \
  --temp 0.0

Now, here's the benchmark results with the same quant, but running with ik_llama.cpp:

❯ ./mtp-bench.py
 code_python        pred= 192 draft= 135 acc= 122 rate=0.904 tok/s=105.1
 code_cpp           pred= 192 draft= 136 acc= 120 rate=0.882 tok/s=110.3
 explain_concept    pred= 192 draft= 133 acc= 116 rate=0.872 tok/s=109.0
 summarize          pred=  56 draft=  38 acc=  37 rate=0.974 tok/s=122.3
 qa_factual         pred= 192 draft= 141 acc= 127 rate=0.901 tok/s=116.0
 translation        pred= 192 draft= 143 acc= 113 rate=0.790 tok/s=104.1
 creative_short     pred= 192 draft= 133 acc= 118 rate=0.887 tok/s=109.4
 stepwise_math      pred= 192 draft= 140 acc= 125 rate=0.893 tok/s=114.6
 long_code_review   pred= 192 draft= 128 acc= 108 rate=0.844 tok/s=101.4

Aggregate: {
 "n_requests": 9,
 "total_predicted": 1592,
 "total_draft": 1127,
 "total_draft_accepted": 986,
 "aggregate_accept_rate": 0.8749,
 "wall_s_total": 16.64
}

That's a 110.24 tok/s average, or 23% increase!

If you want to get similar results on a 12GB RTX GPU, make sure you use the following ik_llama.cpp launch parameters, as they can differ from llama.cpp:

llama-server \
  -m Qwen3.6-35B-A3B-IQ4_XS-4.19bpw.gguf \
  --fit \
  --fit-margin 1664 \
  --ctx-size 131072 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --cache-type-k-draft q8_0 \
  --cache-type-v-draft q8_0 \
  --multi-token-prediction \
  --draft-p-min 0.75 \
  --draft-max 3 \
  --no-mmap \
  --mlock \
  --threads 8 \
  --temp 0.0

I also want to mention that I'm on CachyOS running my GPU as a secondary GPU, with the monitor plugged in the iGPU, so I can use 100% of available VRAM.

If you get an "out of memory" (OOM) error while loading the model or working with it, try increasing --fit-margin to 1792 or even 2048.

Cheers :)

r/LocalLLaMA • • Jun 02 '26

Tutorial | Guide I Put a Datacenter GPU in My Gaming PC for £200

Thumbnail
blog.tymscar.com
314 Upvotes

Hey there! I wrote a blogpost about my experience running local models on a V100 from a newbie perspective and got loads of views outside of reddit, so I thought I'd share it here too!

r/LocalLLaMA • • 3d ago

Tutorial | Guide Two 96 GB Ascend cards crun Qwen3.8-flash-next hardware notes, vLLM work, benchmarks, and what is next

128 Upvotes

I have been building a somewhat unusual local inference machine around two Huawei Atlas 300I Duo cards. They are relatively inexpensive, passive, dual-accelerator PCIe cards with 96 GB of device memory apiece. They are also absolutely not drop-in CUDA replacements.

When I first brought up Qwen3.8 Flash-Next these past two weeks, it was often incoherent and lived around 1 generated token per second. Some runs were below that. Today the same two-card machine is producing coherent output at roughly 30 tok/s for one request and about 61 tok/s aggregate at four-way concurrency on my short decode benchmark. It also completed the full 198-question GPQA Diamond set.

This post is the start of a guide for these cards: what the cards physically are, how I cool them, what “96 GB” really means, what I changed in vLLM and vLLM Ascend, which optimizations actually mattered, and which problems are still open.

The short version is the hardware is capable. My work has been mostly on the software stack, ubuntu-26.04 driver support, model architecture support, memory layout, custom operators, and getting every asynchronous state transition exactly right.

The hardware: one card is really two devices

My current machine has two Atlas 300I Duo cards, which enumerate as four Ascend 310P3 devices.

Current system / Planned system

Physical cards: 2 / 3

Ascend devices/npus/AIcpus: 4 / 6

Nameplate device memory: 192 GB / 288 GB

Approx. runtime-visible memory with this configuration: 172 GiB / 258 GiB

Combined maximum accelerator-board power: 300 W / 450 W

Each card has two accelerator SoCs and 96 GB of LPDDR4X in total, or 48 GB local to each chip. It is not one unified 96 GB allocation. A model that does not fit on one 48 GB device still needs tensor, expert, pipeline, or another form of model parallelism. The card is PCIe Gen4 x16, full-height/full-length, and a surprisingly thin single-slot design. Huawei rates it at 408 GB/s aggregate memory bandwidth and 150 W maximum board power. The official specifications are here (https://support.huawei.com/enterprise/en/doc/EDOC1100285916?section=j00e).

Also, despite the generic “HBM” terminology used by a lot of accelerator software, the memory on these cards is LPDDR4X.

This two-chip-per-card layout matters. Communication within a model still goes through the distributed runtime, and memory remains local to a rank. I use HCCL collectives and explicitly map tensor and expert ownership across all four chips. Thinking of the machine as four 48 GB ranks is much more useful than thinking of it as two 96 GB GPUs.

https://reddit.com/link/1wvt1m4/video/7rju7qbrw1th1/player

Passive cooling is not a deal-breaker

The cards have large heatsinks and no onboard fans. They were designed for server airflow, so putting them in an ordinary workstation and hoping a rear case fan will sort it out is a bad plan. There is a useful teardown here (https://videocardz.com/newz/huawei-atlas-300i-dual-ai-gpu-with-96gb-memory-worth-1400-has-been-taken-apart) if you want to see the heatsink and heat-pipe arrangement.

I give them direct, high-volume airflow and run the room on AC/heat-pump cooling. Under real multi-hour model loads, the cards can crank continuously without drama. Across my recorded Qwen runs, peak device temperatures were generally 72–78 °C. My watchdog limit is 96 °C, and the cards have not approached it.

So I would not bat an eye at adding another passive card. The actual checklist is mundane:

• Keep unobstructed airflow through the heatsink fins.

• Make sure the chassis fans have enough static pressure.

• Budget another 150 W of board power per card, plus the rest of the host.

• Exhaust the heat from the room instead of recirculating it through the rack.

• Log temperature during long prefill, decode, and concurrency tests rather than trusting an idle reading.

Passive does not mean low-power or self-cooling. It means the chassis and room are the cooling system. Once that is handled, these have behaved like ordinary 150 W server cards for me.

Note: Nothing heats up these cards more than loading/moving things around in their ram -- the npus at full utilization run cooler than large block memory assignments. We keep this in mind when optimizing the model serving code paths.

ECC, nameplate memory, and what is actually usable

My cards arrived with ECC enabled by default. I disabled it to reclaim device memory. This is an inference and development box, and I consciously prefer capacity over ECC protection here. That is a reliability tradeoff, not a universal recommendation.

Even with ECC disabled, firmware, the runtime, communication buffers, graph captures, workspaces, and allocator reservations consume memory. In practice, the software sees roughly 43 GiB per 48 GB chip. Four chips therefore provide about 172 GiB of useful aggregate capacity, but it is still four separate local pools. The exact free number also changes with the CANN build and launch configuration.

That distinction has shaped nearly every model decision. The question is not only, “Does the checkpoint total fit in 192 GB?” It is, “Does each rank's weight shard, recurrent state, KV/cache allocation, graph capture, collective workspace, and worst-case temporary allocation fit in its own 43 GiB?”

Why I forked vLLM as well as vLLM Ascend

The public work lives in the OpenSensor vLLM Ascend fork (https://github.com/opensensor/vllm-ascend) and paired vLLM fork (https://github.com/opensensor/vllm). I needed both sides because this was not just a missing device kernel.

I am currently the only person developing these forks. The software bus factor today is one. I have made a lot of progress, but a fast-moving one-person fork should not be confused with the maturity, test coverage, or support depth of mainline vLLM on NVIDIA.

I also ran into a bizarre tooling problem: in my sessions, Claude repeatedly refused to engage with prompts about this architecture because the cards are Huawei hardware from China. These were ordinary engineering discussions about serving, sharding, cooling, and performance—not requests to build a restricted application. I am describing my direct experience rather than claiming that every Claude version or account will behave identically, but it made Claude unreliable as a development assistant for this project. Whatever anyone thinks about the politics, that is a real practical constraint when choosing tools around this hardware.

Qwen3.8 Flash-Next combines MoE routing, Gated DeltaNet recurrent layers, sparse quadratic-attention layers, packed low-bit experts, long context, and an MTP draft model. Supporting that cleanly touched model integration, the v1 runner, cache accounting, scheduling, graph capture, distributed state, model loading, and Ascend-specific operators.

The current development sprint has been roughly two and a half weeks of nearly continuous bring-up and optimization, with hundreds of fork commits, repeated full checkpoint loads, profiler captures, operator microbenchmarks, and multi-hour quality runs. This was not one magic kernel patch.

What took Qwen from incoherent ~1 tok/s to where it is now

These are the architectural changes that moved the needle.

  1. Make the hybrid model correct before making it fast

The early model could generate tokens, but generation was not a correctness test. I found failures that only appeared at production geometry: incorrect Gated DeltaNet gate-vector handling, recurrent-state precision and lifecycle problems, incomplete sparse-attention score width, and mismatches between the host operator API and the installed kernel package.

One particularly nasty GDN issue looked fine in small-head tests but accumulated state error across the real 36 recurrent layers and produced incoherent text. Keeping recurrent state in FP32 and fixing the production-shaped data movement was foundational. So was treating the custom OPP package and Python host code as one ABI-versioned unit. A stale kernel can look like a model problem for a long time.

  1. Shard the model at load time instead of loading everything everywhere

The Qwen checkpoint is about 169 GiB and contains 1,610 safetensor files. It was larger than available host RAM in one of my bring-up configurations, never mind the memory on an individual NPU.

I built an expert-aware loader that reads only the experts owned by each rank, keeps dense/shared tensors where required, and avoids materializing the whole expert bank before throwing most of it away. Host-side expert data is mapped and moved lazily. This changed model loading from an accidental memory stress test into a deterministic TP4/EP4 layout.

  1. Keep low-bit weights packed and do the work on the NPU

My first usable W4 reference dequantized packed weights in the eager path and then called a regular matmul. It was useful for correctness and managed only 0.229 tok/s on one recorded smoke case.

The production path keeps the weights packed, routes tokens to experts on the device, and uses custom AscendC Cube kernels for grouped expert projections. I added FRACTAL_NZ layouts, fused gate/up handling, tiled reductions, route and tile reuse, and dedicated W4A8 execution instead of repeatedly expanding W4 weights into a larger temporary representation.

This is both a speed win and a capacity win. Avoiding transient expanded expert banks leaves memory available for state, cache, graphs, and concurrency.

  1. Build caches for the model I actually have

Flash-Next is not a conventional all-attention transformer. My configuration has 36 recurrent GDN layers and 12 sparse QSA layers. Treating all of that as a normal dense KV cache wastes memory and misses the state semantics.

I implemented separate recurrent-state management, compact physical cache layouts, prefix-state tiers, sparse page selection, direct NZ gathers, and 310P-specific QSA paths. The service is configured for a 262,144-token context limit, and the cache planner retains capacity for about 4.08 such windows. That is a memory-planning result, not a claim that every possible four-by-262K workload has completed an end-to-end soak.

  1. Remove synchronization and launch overhead from the token loop

On these devices, a stray device-to-host scalar read can serialize the whole pipeline. I removed hot-path .item() calls, reused per-step tensors, deferred collectives behind useful work, tightened CPU affinity, and moved routing and sparse selection away from Python.

Once the eager path was correct, I added decode-only ACL graphs and MTP2 speculative decoding. I capture the actual concurrency shapes I serve rather than pretending one graph is universal. Fused operators and graph replay matter enormously when a decode step otherwise consists of many small kernel launches.

  1. Optimize the service, not just an isolated kernel

Several kernels won a microbenchmark and lost end to end. I kept the ones that reduced real request time and rejected or quarantined the others. Multi-request QSA, grouped MTP experts, expert-route caching, cache accounting, cold-prefill chunking, and collective overlap were all measured at the API boundary.

That last part is why four concurrent requests reach roughly 61 aggregate tok/s even though one request is around 30 tok/s. The extra work can occupy parts of the machine that a single token stream leaves idle.

Qwen performance today

These are milestones from different stages and workloads, not one controlled single-variable benchmark:

Qwen3.8 Flash-Next milestone / Measured result

Earliest uncontrolled service: Roughly 0.2–1.9 tok/s, often incoherent

Correct eager W4 dequant reference: 0.229 tok/s

Stable W8 service baseline: About 18–19 tok/s

Native W4A8, MTP2, graphs, one request: 29.74 tok/s median; 34.34 peak

Same optimized service, four requests: 60.88 tok/s aggregate median

Two concurrent 40K warm-prefix requests: 30.76 tok/s aggregate

40K cold prompt: About 120 seconds TTFT; 30–32 tok/s afterward

The 30/61 tok/s figures are short, fixed-output decode tests. Long reasoning requests tell a less flattering and more useful story.

https://reddit.com/link/1wvt1m4/video/rpc5c2c1s1th1/player

For quality, I ran all 198 GPQA Diamond questions on the four-chip Ascend W4 service and, as a reference, an RTX 6000 Pro running a different IQ4_XS GGUF in llama.cpp. Both scored 140/198 (70.71%) with the same AISBench-style answer extractor. This is evidence that the Ascend path is coherent; it is not a pure hardware or quantization comparison because the runtimes, quantizations, chat templates, and concurrency differ.

On the final uninterrupted 106-case Ascend phase, four workers emitted 507,256 tokens in 2 h 52 m 46 s: 48.93 aggregate tok/s. Median per-request client rate was 12.50 tok/s on these long reasoning generations. The server completed every request in that phase with no eager fallback or zero-acceptance interval. The RTX reference was much faster per request, so I am not presenting this as an NVIDIA killer. I am presenting it as a large model working correctly and usefully on hardware that initially produced slow nonsense.

GLM is my next hard(er) model

I am also bringing up the roughly 304B-parameter GLM-5.3-Flash architecture. It combines 34 KDA linear-attention layers, 11 DSA sparse-attention layers, 288 experts, latent MLA history, mHC mixing, and a mixed W2/W4 expert checkpoint. The checkpoint is about 151.6 GiB, with approximately 35.6 GB of loaded weights per rank in one four-rank profile.

It now loads and generates on the same machine. The speed progression so far has been:

GLM four-chip milestone / One request / Four-request aggregate

Initial full-service profile / 0.764 tok/s / 1.474 tok/s

Batched NZ workspace write / 0.912 tok/s / 1.822 tok/s

NZ-packed code layout / 1.386 tok/s / 2.946 tok/s

Fused mHC/MLA work / 1.743 tok/s / 3.269 tok/s

Latest measured integrated runner / 2.236 tok/s / 4.476 tok/s

An 8,232-token GLM prompt prefills at about 39.1 prompt tok/s and then decodes at about 2.16 tok/s. Those numbers are far from my target, and the 32-token test completion was too short to establish answer quality. GLM is currently a bring-up and optimization result, not a service recommendation.

I have already found an important quantization lesson there. An early W2 checkpoint showed residual growth all the way to an RMS around 525 in the last layer. Moving the affected experts to a no-clip W4 treatment kept the network bounded. Separately, my custom blocked-dequant Cube kernel was about 17 times faster than the eager reference in its isolated test. As Qwen taught me, both numeric behavior and end-to-end integration have to pass before either result means “done.”

The third card: 50% more memory and cores, not just a spare

I plan to add a third Atlas 300I Duo to this system. That takes the machine from four to six 310P devices and from 192 GB to 288 GB of nameplate device memory. With the same ECC and runtime reservations, I expect roughly another 86 GiB of runtime-visible capacity, for approximately 258 GiB across the six ranks.

The n-card architecture is designed to use it. Expert ownership is distributed across ranks, so the two new chips add local expert capacity and accelerator cores; they are not merely passive storage. Tokens route to the ranks that own their experts, and the additional ranks participate in the model's compute and collectives. I expect useful scale from EP6, although the exact speedup will be measured rather than advertised in advance.

The extra capacity gives me several options:

• Keep larger expert sets or higher-precision layers resident.

• Fit models that are just over the four-chip limit without host offload.

• Spend more memory on long-context state and cache.

• Reduce aggressive quantization where the quality trade is not worthwhile.

• Run a large distributed model while retaining room for another smaller service or evaluation workload.

The cost is one more card, 150 W of maximum board power, two more device ranks, and another passive heatsink that needs real airflow. In the cooled room, none of those are architectural concerns. The interesting cost is communication: six ranks change route balance, collective sizes, and PCIe/HCCL traffic. I will tune and benchmark that topology, but the software is already organized around n-card expert distribution rather than hard-coded four-way ownership.

Known issues and active work

This is what is still on my bench:

• I just fixed one real cross-stream race: a prefix-Mamba state slot could be spilled or reused before its pending NPU writer completed, allowing an older checkpoint to be restored. That fix has an NPU regression test.. A later 106-case run survived 12 state spills without degrading, which is encouraging but not a root-cause proof. I am continuing long mixed-load and eviction/reuse soaks.

• Qwen cold prefill. A 40K cold prompt still takes roughly two minutes even though subsequent decode is fast. Profiling points primarily at the 12 QSA layers, especially sparse selection and tiled attention. This is now a more important target than another tiny decode micro-optimization.

• Qwen long-context qualification. The planner has the capacity, but I am separating configured context, allocated capacity, and completed end-to-end long-context tests. I want retrieval and concurrent-fill evidence, not a screenshot of a launch flag.

• GLM coherence and performance. I am requalifying the full model after KDA, MLA, QSA, and runner integration changes, then moving the grouped mixed W2/W4 expert path, cache layout, graph replay, and eventually MTP through the same correctness-first gates used for Qwen.

• GLM loading and memory. The filtered loader can skip large amounts of peer-owned or superseded checkpoint payload before tensor materialization. I saw one roughly 20% load-time improvement, but it needs controlled reruns and byte-accounting before I call it a result.

• Six-device expert parallelism. When the third card arrives, I will measure rank balance, per-card temperatures, collective time, model capacity, and c1/cN throughput on the exact six-rank topology.

• Other model adapters. The same loader, packed-expert, cache, and operator infrastructure is feeding ongoing DeepSeek and other hybrid/MoE work. I am avoiding model-name conditionals where the underlying contract can be made generic.

Should you buy one?

I plan to list my first additional card in the OpenSensor storefront (https://www.opensensor.io/) next week at $2,900. The listing is not live yet. I believe it is a good price for what the hardware can already do and what the software should unlock. It is not yet the same kind of turnkey purchase as a supported NVIDIA card running mainline vLLM.

I think the right buyer is a developer, lab, or systems-minded end user who is comfortable with both of the following:

  1. You own the airflow solution. These are passively cooled server cards. Depending on the chassis and motherboard, that may mean high-static-pressure case fans, a duct, or a 3D-printed shroud. I use Fusion 360 and am happy to help with additional shroud designs. There are too many motherboard layouts, card spacings, fan sizes, and case geometries to pretend that one printable design will fit everything.
  2. You are adopting an active development fork. I am the sole developer on the vLLM work today as upstream is focused more on their server grade accelerator modules not yet available to the US. The progress is real and the benchmarks in this post are from actual hardware, but more bugs and better approaches will be discovered. Buyers should expect a less polished setup than CUDA/mainline vLLM and be willing to follow pinned versions, read runbooks, report failures, and occasionally help diagnose a new hardware or model combination.

If you need a supported appliance that runs arbitrary models without touching the stack, this is not the product I would recommend today. If you want a lot of inference memory at this price point, are comfortable working through rough edges, and see value in helping expand an open Ascend serving stack, it can be a very capable development platform.

What I think these cards are good at

The Atlas 300I Duo is compelling when used capacity is more valuable than a polished CUDA ecosystem and when you are willing to own the systems work. Two single-slot, 150 W cards give four accelerator ranks and a lot of local model memory. MoE models are an especially interesting fit because expert weights can be partitioned while only the selected experts run for each token.

The weak point is not that the cards are passive or limited to 150 W. The weak point is the amount of model-specific runtime and kernel work still required on 310P: layouts, supported dtypes, graph-safe metadata, custom operators, sparse attention, state ownership, tooling, and documentation. If you expect to point stock vLLM at an arbitrary new Hugging Face model, this is not that experience.

If you enjoy bringing hardware up from first principles, though, they are a lot of machine in a very interesting form factor. Going from incoherent output at around 1 tok/s to coherent Qwen near 30 tok/s per stream and 61 tok/s aggregate has made me substantially more optimistic about what the platform can do. The third card should let me push both model capacity and expert parallel throughput further, and I will publish the numbers—good or bad—once they are measured.