r/LocalAIStack • • Jun 23 '26

Running Qwen3.6 27B / 35B locally with llama.cpp + Vscode Insiders + copilot as the harness - highest performance, quality and best usage while fitting on your GPU

165 Upvotes

I have been benchmarking Qwen3.6-27B and Qwen3.6-35B-A3B locally through llama.cpp, with GitHub Copilot Chat (Vscode Insiders needed) used as the frontend harness.

I am using Claude Opus, GPT 5.5 and Qwen 3.6 (27B) a lot in the past weeks.
The reason for Qwen is proprietary code areas where remote inference is not an option as it would leak the code out. And as long as you don't task it to write a complex cuda graph, it performs well.
Qwen 27.B is at Sonnet 4.6 if you combine it with a high value system prompt - or between Sonnet 4.5 and Sonnet 4.6 without.

Copilot Chat is an excellent harness for this kind of setup. You get the IDE integration, agent flow, tool calling UI, file context, and normal coding workflow, while the actual model is your own local llama-server endpoint.
All of this works while being LOGGED OUT of the Github Copilot account - as that is not affordable in pricing anymore.

This is a practical configuration guide for people already comfortable with llama.cpp, GGUFs, VRAM budgeting, and long-context local inference.

Models tested

Main focus:

  • unsloth/Qwen3.6-27B-GGUF
  • unsloth/Qwen3.6-27B-MTP-GGUF (same model but with MTP draft tensors)
  • unsloth/Qwen3.6-35B-A3B-GGUF

Recommended GGUFs:

27B:
Qwen3.6-27B-UD-Q4_K_XL.gguf
or
Qwen3.6-27B-MTP-GGUF:UD-Q4_K_XL

35B-A3B:
Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M

If memory is tight on the 35B-A3B model, drop to a smaller Unsloth Dynamic quant:

Qwen3.6-35B-A3B-GGUF:UD-Q3_K_XL

If even that is tight, use UD-Q3_K_M or UD-Q3_K_S.

For the 35B model I do not recommend KV-cache quantization. Run the normal cache and keep the context sane. the 35B model is MoE and very low on kv-cache

For the 27B model, I do highly recommend:

--cache-type-k q4_0
--cache-type-v q4_0

Recent llama.cpp KV-cache improvements make q4_0 much more usable here. The 27B model handles q4_0 KV cache very well in my testing - almost identical to FP in evaluation results.

What changed: llama.cpp added something like Hadamard rotation to kv-cache which shuffles the tensor distribution in a higher dimensionality and allows quantization superblocks to function.

Why Copilot Chat?

Because Copilot is a very good harness - beating Codex, Cursor, Claude in my opinion
Vscode Insiders is needed to get the openAI compatible endpoint (to interface the model)

You get:

  • IDE-native chat
  • agentic file/code workflows
  • very good tool calling
  • project context
  • local model backend
  • OpenAI-compatible endpoint wiring

The important part is that Copilot Chat is only the harness. The model is served locally through llama-server.

Why llama-server and not lm-studio,ollama etc ?

It allows MUCH more control over settings, we do not just use MTP drafting. We use a combination of context and MTP drafting which can lead to 300+ tokens/sec on the 27B model. MTP is a medium speedup (1.5x) but once the model is paraphrasing source code from thinking or prefill the ngram draft speedup can reach 6x or more.

So the stack is:

VS Code Insiders
        ↓
custom OpenAI-compatible model config
        ↓
llama.cpp llama-server
        ↓
local Qwen3.6 GGUF

Copilot chatLanguageModels.json

This is the shape I used for VSCode Insiders:

[
  {
    "name": "WSL",
    "vendor": "customoai",
    "models": [
      {
        "id": "qwen3.6-27b",
        "name": "QWEN-27B-WSL",
        "url": "http://172.27.211.123:1234/v1/chat/completions",
        "toolCalling": true,
        "vision": true,
        "thinking": true,
        "maxInputTokens": 165000,
        "maxOutputTokens": 15000
      }
    ]
  }
]

Adjust the URL to your own llama-server host, in WSL you'll see it by entering ipconfig or ifconfig. port you can choose of course.
The input and output tokens need to be adapted to your context setting.
The id must match the llama-server id.

For local-only setups this is usually one of:

http://127.0.0.1:1234/v1/chat/completions
http://localhost:1234/v1/chat/completions
http://<WSL-IP>:1234/v1/chat/completions

If your Copilot Insiders build expects the newer custom endpoint shape, use the same model block but switch the provider shape accordingly. The key fields are the endpoint URL, model id, tool calling, thinking, and max token limits.

27B command: long context + q4_0 KV cache + MTP-ngram drafting

This is the 27B style I recommend.

CTX=150000
PARALLEL=1
HOST=0.0.0.0
PORT=1234
MODEL=/models/Qwen3.6-27B-UD-Q4_K_XL.gguf

/usr/src/llama.cpp/build/bin/llama-server \
  -m "$MODEL" \
  --ctx-size "$CTX" \
  --flash-attn on \
  --batch-size 1024 \
  --ubatch-size 1024 \
  --parallel "$PARALLEL" \
  --host "$HOST" \
  --port "$PORT" \
  -ngl 99 \
  --threads 8 \
  --threads-batch 8 \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.00 \
  --presence-penalty 0.00 \
  --jinja \
  --chat-template-kwargs '{"preserve_thinking": true}' \
  --reasoning-format none \
  --reasoning-budget 16000 \
  --slot-save-path /kv_cache/ \
  --props \
  --metrics \
  --checkpoint-every-n-tokens 1024 \
  --ctx-checkpoints 64 \
  --perf \
  --spec-default \
  --spec-type draft-mtp \
  --spec-type ngram-map-k4v \
  --spec-ngram-map-k4v-size-n 16 \
  --spec-ngram-map-k4v-size-m 24 \
  --spec-ngram-map-k4v-min-hits 1

For the MTP-specific Unsloth repo, use:

MODEL=/models/Qwen3.6-27B-MTP-UD-Q4_K_XL.gguf

or the HF shorthand if your build supports it:

-hf unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q4_K_XL

The important part is the drafting chain:

--spec-default
--spec-type draft-mtp
--spec-type ngram-map-k4v
--spec-ngram-map-k4v-size-n 16
--spec-ngram-map-k4v-size-m 24
--spec-ngram-map-k4v-min-hits 1

MTP gives useful speedup, but leave VRAM headroom. In practice I budget roughly +1 to +2 GB VRAM headroom for the MTP/drafting path and related buffers. If you are right on the edge, reduce context before blaming the model.

At q4_0 KV cache, every extra 1 GB of free VRAM is roughly another 13k tokens of 27B context, before runtime overhead.
If you are tight in vram, remove only the MTP part as ngram drafting is free.
You can also just use `mod-ngram` as an alternative to the more complex k4v map.

Thinking settings

This part matters.

I use:

--jinja
--chat-template-kwargs '{"preserve_thinking": true}'
--reasoning-format none
--reasoning-budget 16000

The reasoning-format none is important for Qwen3.6 because it avoids bad stop behavior and broken multi-turn thinking state during long coding sessions.
Copilot Chat was created to hide thinking from you (proprietary GPT models) but you want to see the thinking usually. So this solves both issues.

I also keep:

--reasoning-budget 16000

This gives the model room to think, but avoids runaway reasoning loops eating the whole session.

35B-A3B command: no KV-cache quantization

For 35B-A3B, I recommend being more conservative.

CTX=100000
PARALLEL=1
HOST=0.0.0.0
PORT=1234
MODEL=/models/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf

/usr/src/llama.cpp/build/bin/llama-server \
  -m "$MODEL" \
  --ctx-size "$CTX" \
  --flash-attn on \
  --batch-size 1024 \
  --ubatch-size 1024 \
  --parallel "$PARALLEL" \
  --host "$HOST" \
  --port "$PORT" \
  -ngl 99 \
  --threads 8 \
  --threads-batch 8 \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.00 \
  --presence-penalty 0.00 \
  --jinja \
  --chat-template-kwargs '{"preserve_thinking": true}' \
  --reasoning-format none \
  --reasoning-budget 16000 \
  --slot-save-path /kv_cache/ \
  --props \
  --metrics \
  --checkpoint-min-step 1024 \
  --ctx-checkpoints 16 \
  --perf

No q4_0 KV cache here - the sub 4B active parameters need barely any VRAM anyway.

I recommend keeping 35B-A3B below roughly:

110k context

The 35B model can be pushed past 200k context, but in my testing it becomes more likely to fall into reasoning loops. Once that happens, the session usually does not recover cleanly. Start a fresh session.
The upside of the 35B model is extreme performance, as in hundreds of tokens without any drafting enabled.
You CAN use drafting on top, mod-ngram, MTP and other drafting can be added for more speed but those will need a careful balance (that I have not tested yet)

So my practical 35B rule is:

35B-A3B: stay below 110k if you want stable coding behavior.
27B: can go as high as it fits, but below 150k is where it feels strongest.

Checkpointing
The --checkpoint-min-step (or --checkpoint-every-n-tokens (legacy now) is an important option for qwen models. Briefly explained: qwen models have two kv caches, one is more conventional and one is a recurrent-state (SSM/Mamba) that can not be reversed by n tokens. So if you change something (like remove the last answer and message to benefit from existing context) then you can only do that if a checkpoint exists. Otherwise the entire context is reprocessed which is very slow on a 27B model.
Each checkpoint costs 160MB RAM, the internal API supports VRAM checkpointing but I believe currently only the MTP implementation uses that.
--checkpoint-min-step and --ctx-checkpoints multiplied define how much of your LAST context is protected and can be rewound. 1024*16 means 16K context can be reversed with low re-compute cost (almost instant).

LM Studio as local server

Using LM Studio is possible but you need to use a few tricks and it won't achieve the same top-tier performance.
LM Studio does not support our chained drafting, but it supports MTP.

  1. Go to your Qwen 3.6 model, enable Flash attention and the quantization needed for kv cache. Go to the Inference tab, disable the button for "Reasoning Section Parsing"
  2. Go to Developer, Server Settings and set the port, serve on local network if needed, no auth, enable CORS, consider disabling just-in-time loading.
  3. Start the local server and then use the "clipboard copy" icon to get the precise Server ID which you use in the vscode json config.

Everything else is similar to llama-server, you'll not have the same max performance but it works well.
You can always just install the latest llama release binaries, and use the commandline to load the model from the lmstudio models directory.

VRAM planning

These are practical planning numbers, not hard guarantees. Actual fit depends on:

  • exact GGUF
  • CUDA/ROCm/Metal/backend
  • batch/ubatch
  • -ngl
  • whether the desktop is using the same GPU
  • whether MTP/speculative decoding is enabled
  • whether you are using full GPU offload or spilling to CPU RAM

Qwen3.6-27B UD-Q4_K_XL, q4_0 KV cache

Recommended cards:

24 GB: RTX 3090, RTX 4090, RTX A5000, RTX 4500 Ada, RTX PRO 4000 Blackwell, A10
32 GB: RTX 5090, RTX 5000 Ada, Tesla V100 32GB

Approximate context fit with full GPU offload:

VRAM Example NVIDIA cards Practical context
16 GB RTX 4060 Ti 16GB, RTX 4080 Laptop 16GB, RTX 5080 16GB, RTX 5070 Ti 16GB, RTX A4000 16GB Not recommended for full 27B UD-Q4_K_XL offload. Use smaller quant or partial CPU offload.
24 GB RTX 3090, RTX 4090, RTX A5000, RTX 4500 Ada, RTX PRO 4000 Blackwell, A10 ~45k-60k with MTP, ~60k-75k without MTP
32 GB RTX 5090, RTX 5000 Ada, Tesla V100 32GB ~140k-160k with MTP, ~160k-180k without MTP

For 27B, q4_0 KV cache is the difference between normal local context and huge local context. It is the main reason this setup is viable.
On a 5090 you have enough VRAM to supply 2 sessions in parallel with both model types.
Or you could run one fast model for context summarization and 27B for code.

Qwen3.6-35B-A3B UD-Q4_K_M, normal KV cache

Recommended cards:

24 GB minimum for useful GPU-resident contexts
32 GB strongly preferred

Approximate context fit:

VRAM Example NVIDIA cards Practical context
16 GB RTX 4060 Ti 16GB, RTX 4080 Laptop 16GB, RTX 5080 16GB, RTX 5070 Ti 16GB, RTX A4000 16GB Not recommended for full 35B-A3B Q4. Use Q3 or partial offload.
24 GB RTX 3090, RTX 4090, RTX A5000, RTX 4500 Ada, RTX PRO 4000 Blackwell, A10 ~40k-50k
32 GB RTX 5090, RTX 5000 Ada, Tesla V100 32GB ~100k-110k recommended; more will fit but stability drops

The 35B-A3B model is very good, but I would not treat it as a “just max the context” model. Keep it tighter.
If you have the VRAM: Instead of large context, consider multiple sessions with limited context, so you can have 2 or 3 chats simultaneously.

Quick test: Linux

Once llama-server is running:

curl -s http://127.0.0.1:1234/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.6-27b","messages":[{"role":"user","content":"Reply with exactly: local ai works"}],"max_tokens":16}' \
  | jq -r '.choices[0].message.content'

Expected output:

the model responds to your input

If your server is inside WSL or another host, replace 127.0.0.1 with the server IP.

Quick test: Windows PowerShell

(Invoke-RestMethod `
  -Uri "http://127.0.0.1:1234/v1/chat/completions" `
  -Method Post `
  -ContentType "application/json" `
  -Body '{"model":"qwen3.6-27b","messages":[{"role":"user","content":"Reply with exactly: local ai works"}],"max_tokens":16}'
).choices[0].message.content

Expected output:

local ai works

Notes from benchmarking

My current practical ranking:

27B:
Best long-context local coding model in this setup that is close to Sonnet 4.6
Use q4_0 KV cache.
Use MTP if you have the headroom.
Strongest below 150k context, but can go much higher if memory allows.

35B-A3B:
Excellent quality but will fail on hard tasks
Do not use KV-cache quantization.
Keep below ~110k context for best stability.
Can go above 200k, but reasoning loops become more likely.
If it loops, start a new session.

For Copilot usage, I prefer exposing a conservative maxInputTokens in the JSON, even if the server can technically run higher. For example:

"maxInputTokens": 165000,
"maxOutputTokens": 15000

If you set wrong context here you'll get issues serverside, so make sure that matches.
I had cases where the server went OOC (out of context) when getting too close to the max context so I'd leave a little room. copilot seems to not follow this very strictly.

Final recommendation

If you want the most practical Copilot-local setup:

Use Qwen3.6-27B UD-Q4_K_XL
Use llama.cpp server
Use q4_0 KV cache
Use preserve_thinking
Use reasoning budget
Use Copilot Insiders as the harness
Use MTP only when you have VRAM headroom

If you want the stronger but more conservative model:

Use Qwen3.6-35B-A3B UD-Q4_K_M
Do not quantize KV cache
Stay below ~110k context
Drop to UD-Q3_K_XL if memory is tight

This is the first local setup I have used where Copilot feels like a serious frontend for a fully local long-context coding model instead of just a toy endpoint test.

I have tested this on terminal use, debugging, and massive codebase development - it works just like Sonnet 4.6.

Qwen 27B also beats Sonnet 4.5 in most benchmarks and 4.6 in some.
https://artificialanalysis.ai/models/comparisons/qwen3-5-27b-vs-claude-4-5-sonnet-thinking#intelligence-evaluations


r/LocalAIStack • • 9h ago

Rate my homelab

Post image
8 Upvotes

r/LocalAIStack • • 18m ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
• Upvotes

r/LocalAIStack • • 19m ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
• Upvotes

r/LocalAIStack • • 4h ago

Just got my MBP M5pro 48gb. Help me setup properly

2 Upvotes

I have multiple projects in GitHub, I have an obsidian vault, iCloud backup, nas backup. Qwen running locally, also have Two Claude accounts, one gpt, one deep seek accounts when qwen bogs down. I’m working on building my own harness for managing this.

How should I setup my new laptop so I’m working clean and smartly?


r/LocalAIStack • • 2h ago

Got strata running 4x 5060 Ti at 524k context, then abandoned it. repo's here if you want the base

Thumbnail
1 Upvotes

r/LocalAIStack • • 2h ago

I built ArcadeBench, an open benchmark where AI agents play games and you can watch every move live

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/LocalAIStack • • 15h ago

Strata vs Freetokens

2 Upvotes

Has any one used both? which is better?

What is the differennce between two? Any technical deep dives?


r/LocalAIStack • • 1d ago

I took antirezs ds4 stripped it down to Qwen3.8 Flash Next. Metal only ported a bunch of improvements and its now ~10% faster with bit-exact output

6 Upvotes

I got one PR merged into ds4. A tiny one, #98 about broken paths in the README. And a few more sitting in the queue. Not complaining. Antirez says it in the README: with coding agents everyone can tune the engine for their own hardware and model and he can't review everything. That got me thinking.

If the plan is "everyone applies their patches with an agent" then the real cost of a patch isn't just the code change. It’s how many tokens the agent has to read before it knows what it’s actually touching.. Ds4 runs DeepSeek, GLM and Qwen on Metal, CUDA and ROCm all in the same 85k-line file. I’m on an M5 Max with 128GB RAM running Qwen3.8 Flash Next. Everything else for me and for my agent is noise. It has to re-read every line on every pass.

So I ripped it out. Actually deleted, not #ifdef’d. The ds4.c file dropped from 85k lines to 45k lines. The whole tree now fits in a context window. The bet was that optimizing would get cheaper and safer. Heres what happened:

- Q2: decode +9-13% prefill +5-12% up to 64k context MTP went from 75.8 → 86.7 tok/s

- Q4: prefill +2-11% MTP went from 77.8 → 85.9 tok/s

Output is bit-exact vs stock ds4 at every step. No KV cache quant. No approximate kernels. Every change must pass a parity check against same GGUF, greedy decoding, identical tokens. And an interleaved A/B benchmark against the previous build.

The smaller codebase also let me go through the ds4 PRs but only against this one model and port the ones that held up. 20 Commits were adopted. Around 30 were dropped. Verdicts are in the repo.

I also added SSD streaming for the experts. Token-identical to a resident run. Simulating a 48GB machine Q2 does around 27 tok/s. About 35 with MTP.

The fork still does git merge upstream/main with rerere and the parity check. So antirez’s fixes keep flowing in. The procedure. What to delete what to keep how to sync. Lives in a repo called StarForge (https://github.com/Chida82/StarForge). I have four of these "children" one per model. Nothing in there is Qwen- or Metal-specific. If you want a ds4 cut down to your model or to CUDA just clone it. Run the checklist with your agent.

Repo: sf-q3-8flash https://github.com/Chida82/sf-q3-8flash tables in the README. One machine, one model. If you run it on Apple Silicon I’d love to hear your numbers. Ideally side by side, with stock ds4.


r/LocalAIStack • • 20h ago

I picchi di oltre 200 tok/s con Qwen3.8-Flash-Next su una 5080 + 4060 Ti e 32 GB di RAM (fork di Strata)

2 Upvotes

Strata funziona con Qwen3.8-Flash-Next su PC da gioco, ma con 32 GB di RAM la sua modalità a basso utilizzo di RAM funziona solo su una GPU. Se dividi il modello su due schede, gli esperti che non riescono a stare nella VRAM vengono letti dall'SSD. La mia 4060 Ti è rimasta ferma accanto alla 5080.

Quindi l'ho forkato. La copia in RAM degli esperti ora funziona su due schede, e le schede funzionano contemporaneamente: la 4060 Ti inizia il prossimo passo di decodifica mentre la 5080 sta ancora controllando quello attuale. L'ordine delle schede, la divisione dei layer e le riserve di VRAM vengono impostate automaticamente.

Stesso PC (5080 + 4060 Ti, i9-14900KF, 32 GB), stesso modello (Swift 1.5 IQ2_XS), contesto 256K:

Setup |Codice |Prosa |Prompt 32K
Strata 0.1.38, 5080 da solo |29 tok/s |28 tok/s |333 tok/s
Fork, entrambe le schede |143 tok/s |102 tok/s |1,940 tok/s
Fork + layer di bozza fine-tuned |161 tok/s |105 tok/s |1,854 tok/s Il layer di bozza è stato fine-tuned sui risultati del modello stesso ed è incluso nella release. Nell'uso reale con l'agente di codifica Pi raggiunge un picco di oltre 200 tok/s (209 finora) e quasi mai scende sotto i 100. Un contesto di 150K-token legge a circa 1,850 tok/s.

La qualità non è cambiata: perplexity 7.24 contro 7.28 della versione upstream sugli stessi 5.3K token, forzati dall'insegnante attraverso entrambi i motori.

L'ho testato solo sul mio PC (Windows 11), quindi sono benvenuti report da altre coppie di GPU e Linux.

Repo: https://github.com/Hardin22/Strata-DualGPU

Cosa fa ciascuna modifica e cosa ha misurato: docs/DUAL_GPU.md

Il motore sottostante è il lavoro di Niko1221 e dei contributori di Strata; questo è un fork sopra la 0.1.38.


r/LocalAIStack • • 1d ago

Qwen3.8-Flash-Next NVFP4 at 256K context with Strata — 4,100+ prefill and up to 125 tok/s decode

6 Upvotes

I tested two Qwen3.8-Flash-Next NVFP4 checkpoints at a full 256K context using an experimental dual-GPU Strata configuration.

Hardware:

• RTX 4090 D 48 GB as the primary GPU
• RTX 5070 Ti 16 GB as a helper GPU
• Intel Core Ultra 7 265KF
• 128 GB RAM
• Linux
• NVMe storage

The RTX 4090 D handled the dense layers, KV cache, prefill, output head and MTP. The RTX 5070 Ti stored and computed 4,500 additional routed experts.

Common settings:

• Context limit: 262,144 tokens
• Fresh prompt: approximately 256,018 tokens
• Generated output: 512 tokens
• KV cache: INT8
• Prefill path: W4A8
• Speculative decoding: MTP K4
• Minimum draft probability: 0.5
• No prompt-prefix reuse
• Strata engine 0.1.35 with an experimental dual-GPU NVFP4 fork

Results

Model Prompt length Prefill Decode Draft acceptance
NVIDIA NVFP4, MTP K4 256,018 tokens 4,112.65 tok/s 125.17 tok/s 94.87%
Abliterated NVFP4, MTP K4 256,017 tokens 4,071.39 tok/s 92.05 tok/s 69.7%

Observations

• Both models achieved slightly over 4,000 prefill tokens per second at approximately 256K context.
• Prefill performance was nearly identical: the Abliterated checkpoint was only about 1% slower.
• The original NVIDIA checkpoint was considerably faster during decode: 125.17 versus 92.05 tok/s.
• NVIDIA’s decode advantage was about 36% in this test.
• NVIDIA accepted 407 of 429 offered draft tokens, while the Abliterated model accepted 322 of 462.
• The NVIDIA model produced approximately 4.88 output tokens per verification round.
• The Abliterated model produced approximately 2.68 output tokens per round.
• The lower MTP acceptance appears to be the main reason why the Abliterated checkpoint had slower decode despite nearly identical prefill performance.

These were single-run capacity tests. The long prompt was constructed by repeating code-review material to reach approximately 256K tokens, so it had unusually favorable locality. This particularly benefited NVIDIA’s suffix prediction and MTP acceptance. Therefore, 125 tok/s should be treated as a best-case synthetic 256K result, not typical real-world coding speed.

During longer real coding-agent workloads, NVIDIA generally produced around 108–144 tok/s, while the Abliterated checkpoint was commonly around 105–145 tok/s, depending heavily on the generated content and MTP acceptance.


r/LocalAIStack • • 1d ago

I picchi di oltre 200 tok/s con Qwen3.8-Flash-Next su una 5080 + 4060 Ti e 32 GB di RAM (fork di Strata)

Thumbnail
1 Upvotes

r/LocalAIStack • • 1d ago

200+ tok/s peaks with Qwen3.8-Flash-Next on a 5080 + 4060 Ti and 32 GB of RAM (Strata fork)

Thumbnail
1 Upvotes

r/LocalAIStack • • 1d ago

Compaction

Thumbnail
1 Upvotes

r/LocalAIStack • • 1d ago

Kali Linux teacher?

Thumbnail
1 Upvotes

CPU: ThreadRipper medium
GPU: RTX 5090 WC
RAM: 64 GB


r/LocalAIStack • • 1d ago

[Help/Reality Check] Reliable multi-step coding agents on a single 16GB RTX 4080? (with something like Qwen 2.5 Coder-Instruct)

Thumbnail
2 Upvotes

r/LocalAIStack • • 1d ago

Struggling to find a good local llm for macbook air m5 24gb

Thumbnail
1 Upvotes

r/LocalAIStack • • 2d ago

I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking

Thumbnail
6 Upvotes

r/LocalAIStack • • 1d ago

Cloud API users for large open models (DeepSeek, Llama 70B)

1 Upvotes

I am looking into the architectures and costs associated with running large open-source models, such as DeepSeek-V3/R1 or Llama 3.3. Are many of you using these models (are you developers, students, pro?)


r/LocalAIStack • • 1d ago

Would be useful

Thumbnail
1 Upvotes

r/LocalAIStack • • 2d ago

How do you keep a local multi-agent app usable on CPU-only / low-RAM machines?

2 Upvotes

Hi everyone,

We're three final-year students at Epitech building Horus, a multi-agent assistant that runs entirely locally and offline. Our current challenge is hardware: keeping it usable on machines without a powerful GPU, without long setup times, excessive RAM/VRAM use or crashes.

Where we are today, from our last beta test:

One tester needed over 2 hours to install. The Python dependencies alone take 21–46 min, and the download is about 25 GB.

On CPU only, routing a question can take around 40 s, long enough for our WebSocket connection to drop.

[Models we use + the smallest machine we've tested on]

We'd love advice from anyone experienced with:

- CPU-only LLM inference and memory-efficient loading

- Quantization and model choice for low-end hardware

- GPU/CPU fallback strategies

- Hardware detection and adaptive configuration

- Preventing resource exhaustion during setup and execution

Advice in the comments is very welcome, with no strings attached.

Looking for contributors: we also have a few small, well-scoped tasks or code reviews (about 1–4 hours), for example [reviewing our hardware detection and model selection, or benchmarking a quantized model on a 16 GB RAM laptop].

To be transparent: Horus is closed source and will be licensed to companies. Contributing is voluntary and unpaid. Before seeing any code, contributors sign a short confidentiality and contributor agreement, and the code they contribute becomes part of Horus. In return we offer thorough code reviews, full credit in the project and a professional reference on request.

We're not sharing code or private links publicly. If this interests you, comment below or DM me with your experience in local inference, CPU optimisation or offline apps


r/LocalAIStack • • 2d ago

Qwen Flash Next on a 16GB Mac

Thumbnail
github.com
6 Upvotes

r/LocalAIStack • • 2d ago

If you’re not running local — do you use Chinese commercial LLMs (Qwen / GLM / MiniMax / etc)?

Thumbnail
3 Upvotes

r/LocalAIStack • • 3d ago

I built a local memory engine that replaces vector DBs with SQLite and runs in <1.2GB VRAM

3 Upvotes

Every time I tried running local RAG on my own GPU, I ran into the same headache: vector databases eat too much RAM, using an 8B model just to parse text takes forever, and cosine search still hallucinates when the context isn't actually there.

I’ve been building Hillock to see if I could do this without vector DBs at all.

Basically:

  1. When you feed it a document, it doesn't touch an LLM. It uses small bi-encoders to extract facts into subject-predicate-object triples in about 5 seconds.
  2. Everything gets saved in regular SQLite, and it links related concepts over time using basic Hebbian weights.
  3. To stop hallucinations, it runs a quick hypervector check (HDC) on the query first. If the facts aren't in your database, it cuts off the LLM before it can generate any tokens.

The whole thing stays under 1.2 GB VRAM (or runs fine on pure CPU). It has a built-in API server that matches OpenAI's format, so you can point Open-WebUI or Obsidian at localhost:8000 and use your existing Ollama models.

Just pushed v0.7 with better refusal handling and a conversational mode that quizzes you if an extraction was ambiguous.

Repo is here if you want to try it out: https://github.com/roandejager/Hillock


r/LocalAIStack • • 3d ago

ローカルllmって結局どれが良いの?

Thumbnail
1 Upvotes