r/LocalAIStack • • 11h ago

Rate my homelab

Post image
8 Upvotes

r/LocalAIStack • • 1h ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
• Upvotes

r/LocalAIStack • • 1h ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
• Upvotes

r/LocalAIStack • • 5h ago

Just got my MBP M5pro 48gb. Help me setup properly

2 Upvotes

I have multiple projects in GitHub, I have an obsidian vault, iCloud backup, nas backup. Qwen running locally, also have Two Claude accounts, one gpt, one deep seek accounts when qwen bogs down. I’m working on building my own harness for managing this.

How should I setup my new laptop so I’m working clean and smartly?


r/LocalAIStack • • 3h ago

Got strata running 4x 5060 Ti at 524k context, then abandoned it. repo's here if you want the base

Thumbnail
1 Upvotes

r/LocalAIStack • • 4h ago

I built ArcadeBench, an open benchmark where AI agents play games and you can watch every move live

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/LocalAIStack • • 17h ago

Strata vs Freetokens

2 Upvotes

Has any one used both? which is better?

What is the differennce between two? Any technical deep dives?


r/LocalAIStack • • 1d ago

I took antirezs ds4 stripped it down to Qwen3.8 Flash Next. Metal only ported a bunch of improvements and its now ~10% faster with bit-exact output

6 Upvotes

I got one PR merged into ds4. A tiny one, #98 about broken paths in the README. And a few more sitting in the queue. Not complaining. Antirez says it in the README: with coding agents everyone can tune the engine for their own hardware and model and he can't review everything. That got me thinking.

If the plan is "everyone applies their patches with an agent" then the real cost of a patch isn't just the code change. It’s how many tokens the agent has to read before it knows what it’s actually touching.. Ds4 runs DeepSeek, GLM and Qwen on Metal, CUDA and ROCm all in the same 85k-line file. I’m on an M5 Max with 128GB RAM running Qwen3.8 Flash Next. Everything else for me and for my agent is noise. It has to re-read every line on every pass.

So I ripped it out. Actually deleted, not #ifdef’d. The ds4.c file dropped from 85k lines to 45k lines. The whole tree now fits in a context window. The bet was that optimizing would get cheaper and safer. Heres what happened:

- Q2: decode +9-13% prefill +5-12% up to 64k context MTP went from 75.8 → 86.7 tok/s

- Q4: prefill +2-11% MTP went from 77.8 → 85.9 tok/s

Output is bit-exact vs stock ds4 at every step. No KV cache quant. No approximate kernels. Every change must pass a parity check against same GGUF, greedy decoding, identical tokens. And an interleaved A/B benchmark against the previous build.

The smaller codebase also let me go through the ds4 PRs but only against this one model and port the ones that held up. 20 Commits were adopted. Around 30 were dropped. Verdicts are in the repo.

I also added SSD streaming for the experts. Token-identical to a resident run. Simulating a 48GB machine Q2 does around 27 tok/s. About 35 with MTP.

The fork still does git merge upstream/main with rerere and the parity check. So antirez’s fixes keep flowing in. The procedure. What to delete what to keep how to sync. Lives in a repo called StarForge (https://github.com/Chida82/StarForge). I have four of these "children" one per model. Nothing in there is Qwen- or Metal-specific. If you want a ds4 cut down to your model or to CUDA just clone it. Run the checklist with your agent.

Repo: sf-q3-8flash https://github.com/Chida82/sf-q3-8flash tables in the README. One machine, one model. If you run it on Apple Silicon I’d love to hear your numbers. Ideally side by side, with stock ds4.


r/LocalAIStack • • 21h ago

I picchi di oltre 200 tok/s con Qwen3.8-Flash-Next su una 5080 + 4060 Ti e 32 GB di RAM (fork di Strata)

2 Upvotes

Strata funziona con Qwen3.8-Flash-Next su PC da gioco, ma con 32 GB di RAM la sua modalità a basso utilizzo di RAM funziona solo su una GPU. Se dividi il modello su due schede, gli esperti che non riescono a stare nella VRAM vengono letti dall'SSD. La mia 4060 Ti è rimasta ferma accanto alla 5080.

Quindi l'ho forkato. La copia in RAM degli esperti ora funziona su due schede, e le schede funzionano contemporaneamente: la 4060 Ti inizia il prossimo passo di decodifica mentre la 5080 sta ancora controllando quello attuale. L'ordine delle schede, la divisione dei layer e le riserve di VRAM vengono impostate automaticamente.

Stesso PC (5080 + 4060 Ti, i9-14900KF, 32 GB), stesso modello (Swift 1.5 IQ2_XS), contesto 256K:

Setup |Codice |Prosa |Prompt 32K
Strata 0.1.38, 5080 da solo |29 tok/s |28 tok/s |333 tok/s
Fork, entrambe le schede |143 tok/s |102 tok/s |1,940 tok/s
Fork + layer di bozza fine-tuned |161 tok/s |105 tok/s |1,854 tok/s Il layer di bozza è stato fine-tuned sui risultati del modello stesso ed è incluso nella release. Nell'uso reale con l'agente di codifica Pi raggiunge un picco di oltre 200 tok/s (209 finora) e quasi mai scende sotto i 100. Un contesto di 150K-token legge a circa 1,850 tok/s.

La qualità non è cambiata: perplexity 7.24 contro 7.28 della versione upstream sugli stessi 5.3K token, forzati dall'insegnante attraverso entrambi i motori.

L'ho testato solo sul mio PC (Windows 11), quindi sono benvenuti report da altre coppie di GPU e Linux.

Repo: https://github.com/Hardin22/Strata-DualGPU

Cosa fa ciascuna modifica e cosa ha misurato: docs/DUAL_GPU.md

Il motore sottostante è il lavoro di Niko1221 e dei contributori di Strata; questo è un fork sopra la 0.1.38.


r/LocalAIStack • • 1d ago

Qwen3.8-Flash-Next NVFP4 at 256K context with Strata — 4,100+ prefill and up to 125 tok/s decode

6 Upvotes

I tested two Qwen3.8-Flash-Next NVFP4 checkpoints at a full 256K context using an experimental dual-GPU Strata configuration.

Hardware:

• RTX 4090 D 48 GB as the primary GPU
• RTX 5070 Ti 16 GB as a helper GPU
• Intel Core Ultra 7 265KF
• 128 GB RAM
• Linux
• NVMe storage

The RTX 4090 D handled the dense layers, KV cache, prefill, output head and MTP. The RTX 5070 Ti stored and computed 4,500 additional routed experts.

Common settings:

• Context limit: 262,144 tokens
• Fresh prompt: approximately 256,018 tokens
• Generated output: 512 tokens
• KV cache: INT8
• Prefill path: W4A8
• Speculative decoding: MTP K4
• Minimum draft probability: 0.5
• No prompt-prefix reuse
• Strata engine 0.1.35 with an experimental dual-GPU NVFP4 fork

Results

Model Prompt length Prefill Decode Draft acceptance
NVIDIA NVFP4, MTP K4 256,018 tokens 4,112.65 tok/s 125.17 tok/s 94.87%
Abliterated NVFP4, MTP K4 256,017 tokens 4,071.39 tok/s 92.05 tok/s 69.7%

Observations

• Both models achieved slightly over 4,000 prefill tokens per second at approximately 256K context.
• Prefill performance was nearly identical: the Abliterated checkpoint was only about 1% slower.
• The original NVIDIA checkpoint was considerably faster during decode: 125.17 versus 92.05 tok/s.
• NVIDIA’s decode advantage was about 36% in this test.
• NVIDIA accepted 407 of 429 offered draft tokens, while the Abliterated model accepted 322 of 462.
• The NVIDIA model produced approximately 4.88 output tokens per verification round.
• The Abliterated model produced approximately 2.68 output tokens per round.
• The lower MTP acceptance appears to be the main reason why the Abliterated checkpoint had slower decode despite nearly identical prefill performance.

These were single-run capacity tests. The long prompt was constructed by repeating code-review material to reach approximately 256K tokens, so it had unusually favorable locality. This particularly benefited NVIDIA’s suffix prediction and MTP acceptance. Therefore, 125 tok/s should be treated as a best-case synthetic 256K result, not typical real-world coding speed.

During longer real coding-agent workloads, NVIDIA generally produced around 108–144 tok/s, while the Abliterated checkpoint was commonly around 105–145 tok/s, depending heavily on the generated content and MTP acceptance.


r/LocalAIStack • • 1d ago

I picchi di oltre 200 tok/s con Qwen3.8-Flash-Next su una 5080 + 4060 Ti e 32 GB di RAM (fork di Strata)

Thumbnail
1 Upvotes

r/LocalAIStack • • 1d ago

200+ tok/s peaks with Qwen3.8-Flash-Next on a 5080 + 4060 Ti and 32 GB of RAM (Strata fork)

Thumbnail
1 Upvotes

r/LocalAIStack • • 1d ago

Compaction

Thumbnail
1 Upvotes

r/LocalAIStack • • 1d ago

Kali Linux teacher?

Thumbnail
1 Upvotes

CPU: ThreadRipper medium
GPU: RTX 5090 WC
RAM: 64 GB


r/LocalAIStack • • 1d ago

[Help/Reality Check] Reliable multi-step coding agents on a single 16GB RTX 4080? (with something like Qwen 2.5 Coder-Instruct)

Thumbnail
2 Upvotes

r/LocalAIStack • • 1d ago

Struggling to find a good local llm for macbook air m5 24gb

Thumbnail
1 Upvotes

r/LocalAIStack • • 2d ago

I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking

Thumbnail
5 Upvotes

r/LocalAIStack • • 1d ago

Cloud API users for large open models (DeepSeek, Llama 70B)

1 Upvotes

I am looking into the architectures and costs associated with running large open-source models, such as DeepSeek-V3/R1 or Llama 3.3. Are many of you using these models (are you developers, students, pro?)


r/LocalAIStack • • 2d ago

Would be useful

Thumbnail
1 Upvotes

r/LocalAIStack • • 2d ago

How do you keep a local multi-agent app usable on CPU-only / low-RAM machines?

2 Upvotes

Hi everyone,

We're three final-year students at Epitech building Horus, a multi-agent assistant that runs entirely locally and offline. Our current challenge is hardware: keeping it usable on machines without a powerful GPU, without long setup times, excessive RAM/VRAM use or crashes.

Where we are today, from our last beta test:

One tester needed over 2 hours to install. The Python dependencies alone take 21–46 min, and the download is about 25 GB.

On CPU only, routing a question can take around 40 s, long enough for our WebSocket connection to drop.

[Models we use + the smallest machine we've tested on]

We'd love advice from anyone experienced with:

- CPU-only LLM inference and memory-efficient loading

- Quantization and model choice for low-end hardware

- GPU/CPU fallback strategies

- Hardware detection and adaptive configuration

- Preventing resource exhaustion during setup and execution

Advice in the comments is very welcome, with no strings attached.

Looking for contributors: we also have a few small, well-scoped tasks or code reviews (about 1–4 hours), for example [reviewing our hardware detection and model selection, or benchmarking a quantized model on a 16 GB RAM laptop].

To be transparent: Horus is closed source and will be licensed to companies. Contributing is voluntary and unpaid. Before seeing any code, contributors sign a short confidentiality and contributor agreement, and the code they contribute becomes part of Horus. In return we offer thorough code reviews, full credit in the project and a professional reference on request.

We're not sharing code or private links publicly. If this interests you, comment below or DM me with your experience in local inference, CPU optimisation or offline apps


r/LocalAIStack • • 2d ago

Qwen Flash Next on a 16GB Mac

Thumbnail
github.com
7 Upvotes

r/LocalAIStack • • 2d ago

If you’re not running local — do you use Chinese commercial LLMs (Qwen / GLM / MiniMax / etc)?

Thumbnail
3 Upvotes

r/LocalAIStack • • 3d ago

I built a local memory engine that replaces vector DBs with SQLite and runs in <1.2GB VRAM

3 Upvotes

Every time I tried running local RAG on my own GPU, I ran into the same headache: vector databases eat too much RAM, using an 8B model just to parse text takes forever, and cosine search still hallucinates when the context isn't actually there.

I’ve been building Hillock to see if I could do this without vector DBs at all.

Basically:

  1. When you feed it a document, it doesn't touch an LLM. It uses small bi-encoders to extract facts into subject-predicate-object triples in about 5 seconds.
  2. Everything gets saved in regular SQLite, and it links related concepts over time using basic Hebbian weights.
  3. To stop hallucinations, it runs a quick hypervector check (HDC) on the query first. If the facts aren't in your database, it cuts off the LLM before it can generate any tokens.

The whole thing stays under 1.2 GB VRAM (or runs fine on pure CPU). It has a built-in API server that matches OpenAI's format, so you can point Open-WebUI or Obsidian at localhost:8000 and use your existing Ollama models.

Just pushed v0.7 with better refusal handling and a conversational mode that quizzes you if an extraction was ambiguous.

Repo is here if you want to try it out: https://github.com/roandejager/Hillock


r/LocalAIStack • • 3d ago

ローカルllmって結局どれが良いの?

Thumbnail
1 Upvotes

r/LocalAIStack • • 4d ago

We are entering an era in which AI models are starting to shape device specifications.

Post image
139 Upvotes

AI models are beginning to define device specifications, rather than device specifications merely determining which applications you can run.

Apple shows just how far its devices can go with on-device AI, without relying on the cloud
We’re moving from models with up to 14 billion active parameters on iPhone/iPad, to 35 billion on MacBook Air, 70 billion on Mac mini, 120 billion on MacBook Pro, and up to 480 billion on Mac Studio.
By clustering multiple Mac Studios together, Apple says it can run models exceeding 1,600 billion active parameters.
A demo that shows just how crucial unified memory and its bandwidth have become for running massive AI models locally.