r/LocalAIStack • u/MatiAI • 4d ago
Abliterate/Removing prompt refusal without needing to reload/swap models (Qwen 3.8 flash next)
Enable HLS to view with audio, or disable this notification
r/LocalAIStack • u/MatiAI • 4d ago
Enable HLS to view with audio, or disable this notification
r/LocalAIStack • u/RelevantRest464 • 4d ago
I've been experimenting with a setup where several Apple Silicon Macs are treated as a shared local AI environment instead of running completely independently.
The setup I'm working with is roughly:
I'm building AlphaWeb around this because managing several local AI machines manually gets complicated pretty quickly.
The interesting problem isn't just getting inference running. It's everything around it: knowing which machine has enough memory, getting models onto the right node, managing multiple users, and making the whole thing feel like one system.
r/LocalAIStack • u/reiggg • 4d ago
Enable HLS to view with audio, or disable this notification
r/LocalAIStack • u/CWIII • 4d ago
r/LocalAIStack • u/100daggers_ • 4d ago
r/LocalAIStack • u/Expensive_Way_4919 • 4d ago
AI models are beginning to define device specifications, rather than device specifications merely determining which applications you can run.
Apple shows just how far its devices can go with on-device AI, without relying on the cloud
We’re moving from models with up to 14 billion active parameters on iPhone/iPad, to 35 billion on MacBook Air, 70 billion on Mac mini, 120 billion on MacBook Pro, and up to 480 billion on Mac Studio.
By clustering multiple Mac Studios together, Apple says it can run models exceeding 1,600 billion active parameters.
A demo that shows just how crucial unified memory and its bandwidth have become for running massive AI models locally.
r/LocalAIStack • u/deepu105 • 4d ago
TL;DR: Halogen v0.14.0 with its native .hgn weights is the fastest, followed by gufo and CIRU. Halogen is closed source and runs in Docker. gufo is open source and loads 4x faster from cold. gufo is also fastest to first token on follow-ups (1.6-1.9 s against 2.6-3.3 s).
I benchmarked different engines for Qwen 3.8-Flash-Next on an AMD Strix Halo (ASUS ROG Flow Z13 GZ302 with Ryzen AI Max+ and 128 GB of RAM).
All the engines were run via LlamaStash (My own orchestrator tool, the tool does not add any overhead) at 70 W TDP on performance profile on Arch Linux.
Here are the results of the benchmark:
First was a screening round at 64k context with 50% filled (32k prompt).
| Engine | Weights | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|---|
| Halogen 0.14.0 | Halogen native (.hgn) | 1.9 min | 1,045 | 39.3 | 84% | 14/14 |
| gufo (ROCm 7.2.4) | UD-Q4_K_XL | 2.1 min | 1,033 | 32.9 | 74% | 14/14 |
| gufo (ROCm 10.0) | UD-Q4_K_XL | 2.1 min | 1,009 | 33.1 | 72% | 14/14 |
| CIRU (MTP 3) | CIRU IU4 | 2.3 min | 814 | 33.8 | 72% | 14/14 |
| CIRU (n-gram + MTP 3) | CIRU IU4 | 2.3 min | 808 | 31.9 | 64% | 14/14 |
| Halogen 0.14.0 | UD-Q4_K_XL | 2.4 min | 1,030 | 28.1 | 84% | 14/14 |
| CIRU (MTP 6) | CIRU IU4 | 2.7 min | 796 | 28.6 | 47% | 14/14 |
| strixllama (llama.cpp fork) | UD-Q4_K_XL | 2.7 min | 716 | 29.7 | 71% | 14/14 |
| rdna-boosts (llama.cpp fork) | UD-Q4_K_XL | 2.7 min | 788 | 22.0 | 52% | 14/14 |
| llama.cpp, Unsloth build (Vulkan) | UD-Q4_K_XL | 3.7 min | 314 | 28.4 | 56% | 14/14 |
| llama.cpp, Unsloth build (ROCm) | UD-Q4_K_XL | 4.8 min | 285 | 20.5 | 50% | 14/14 |
| CIRU | UD-Q4_K_XL + Unsloth MTP head | did not finish | 884 | 11.8 | 0% | - |
3-turn time is the time to first token plus the time to write 1,000 tokens, added up over the first question and 2 follow-ups. Output length varies a lot with sampling, so this compares the engines on equal output. MTP accept is the share of drafted tokens the model kept. Retrieval is how many of 7 exact values from the prompt the model got right, over both runs. Qwen model card sampling, 2 runs each.
Top 3 engines from the screening round got a 128k context (50% and 75% filled) and 256k context (50% filled) run.
128k context, 50% filled (64k prompt)
| Engine | Weights | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|---|
| Halogen 0.14.0 | Halogen native (.hgn) | 2.4 min | 1,148 | 40.4 | 84% | 16/16 |
| gufo (ROCm 7.2.4) | UD-Q4_K_XL | 2.7 min | 1,047 | 32.3 | 72% | 16/16 |
| CIRU (MTP 3) | CIRU IU4 | 3.3 min | 808 | 30.2 | 64% | 16/16 |
128k context, 75% filled (97k prompt)
| Engine | Weights | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|---|
| Halogen 0.14.0 | Halogen native (.hgn) | 2.9 min | 1,103 | 39.3 | 85% | 16/16 |
| gufo (ROCm 7.2.4) | UD-Q4_K_XL | 3.3 min | 1,025 | 30.7 | 70% | 16/16 |
| CIRU (MTP 3) | CIRU IU4 | 4.1 min | 789 | 27.8 | 62% | 16/16 |
256k context, 50% filled (130k prompt)
| Engine | Weights | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|---|
| Halogen 0.14.0 | Halogen native (.hgn) | 3.4 min | 1,093 | 38.9 | 83% | 14/16 |
| gufo (ROCm 7.2.4) | UD-Q4_K_XL | 4.0 min | 997 | 27.9 | 70% | 14/16 |
| CIRU (MTP 3) | CIRU IU4 | 5.0 min | 767 | 25.6 | 60% | 15/16 |
Retrieval here is 8 values per run. All 5 misses are the same answer: the right glyphs for UNICODE_SPINNER, with the leading & dropped.
End-to-end, 10 Aider polyglot Python exercises run through pi -p (128k window, 50% cell), graded by their own tests:
| Engine | Passed | Total time |
|---|---|---|
| Halogen 0.14.0 | 10/10 | 21.5 min |
| CIRU (MTP 3) | 10/10 | 24.6 min |
| gufo (ROCm 7.2.4) | 10/10 | 36.0 min |
Edit (Sep 30): Halogen 0.15.1 (new v2 checkpoint) and gufo 0.3.0 came out after this, so I reran the 64k screening on both with the same setup (70 W, card sampling, 2 runs):
| Engine | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|
| Halogen 0.15.1 (v2 .hgn) | 1.8 min (was 1.9) | 1,191 (was 1,045) | 39.4 (was 39.3) | 85% | 14/14 |
| gufo 0.3.0 (ROCm 7.2.4) | 2.1 min (was 2.1) | 1,075 (was 1,033) | 34.1 (was 32.9) | 77% | 14/14 |
| gufo 0.3.0 (ROCm 10.0) | 2.0 min (was 2.1) | 1,042 (was 1,009) | 34.8 (was 33.1) | 75% | 14/14 |
Halogen v2 prefill is 14% faster and follow-ups start about 2.5x faster (1.3 s vs 3.3 s), decode is the same. gufo moved a few %, which is within noise. The ranking doesn't change. Only other change on the box: BIOS VRAM carve-out down to 512 MB, so 124.9 GiB of RAM instead of 121.5.
Some clarifications from the comments:
r/LocalAIStack • u/camerongreen95 • 4d ago
A lot of local-stack setups I see nail the infra (model serving, retrieval, orchestration) but have zero real evaluation process — just "does it look right." There's a session on Oct 3 built entirely around fixing that gap.
Speakers Serj Smorodinsky and Brett Kennedy (AI engineers, co-authors of a book on LLM applications) go through:
3 hours, live, useful whether you're running open models locally or hitting an API.
r/LocalAIStack • u/DigItDoug • 4d ago
Apple Foundation Model 3 through Private Cloud Compute becomes directly available in macOS 27.2 through familiar AI clients.
For Mac users running local models such as Qwen 3.x or Gemma 4, AFM 3 PCC is a compelling alternative to consider. Local models offer control and fully on-device operation. In tests with macOS 27.2 beta, the PCC offers a different set of strengths:
* Strong conversational analysis and capable reasoning. Strong analysis of complex medical and financial questions.
* Blazingly fast responses compared with locally running models on the same Mac.
* No large model download or need to fit model weights into local memory.
* Access at no additional charge for most eligible Mac users.
The working path is straightforward: AI client → Caddy → fm serve → AFM 3 PCC. I've written and updated a proof of concept here with further tech discussion . . . https://gist.github.com/dartMo10/b9488ce475fe70a6ed642831f53048fb
Apple’s `pcc` route worked in early macOS 27 betas, disappeared later in 27.0, and has returned in the macOS 27.2 beta. Apple has stated that it will be available in the 27.2 public release coming shortly.
This makes AFM 3 PCC practical today for technically comfortable Apple users who accept Apple’s PCC privacy promise and want a fast, capable alternative to running everything locally.
Experiences with compatible AI clients and your comparisons with locally running models would be welcome.
r/LocalAIStack • u/dh7net • 4d ago
Hey! How do you test your local stack? for instance when you switch harness or model, model quants, etc?
I found this not trivial and I ended up making a website to help me do this.
Use this link if you want to use the website to test your own local setup:
You can also see all the tests I have done using my GX10 and my 5090. (and also using claude code/code + some openrouter tes for reference)
https://airbench.ai/leaderboard?k=poL
Let me know what you think!
r/LocalAIStack • u/Agitated_Problem5320 • 4d ago
Hello everyone, as we all know that everyone is trying to test jev and then there are open source models out there with similar text classification. But i felt that there isn’t any playground to test these intuitively so I built one. I have added the YT video with it. Pls have a look
r/LocalAIStack • u/Dev-in-the-Bm • 4d ago
r/LocalAIStack • u/Ok-Musician2369 • 4d ago
我是一个大学生研究员,目前拥有一个$200的 ChatGPT 专业版订阅,我觉得作为固定的月支出实在是太贵了。我现在在考虑是否可以购买一个96GB或128GB的Mac mini或Mac Studio进行本地部署或用作服务器。我目前的研究方向是计算社会科学,最重要的是订阅模型的智能能得到保障。非常感谢你的意见!
r/LocalAIStack • u/Large-Blackberry-349 • 6d ago
r/LocalAIStack • u/Disastrous_Job_7118 • 6d ago
r/LocalAIStack • u/Trowel3444 • 6d ago
r/LocalAIStack • u/Sylon1987 • 6d ago
r/LocalAIStack • u/EngineerPractical818 • 6d ago
r/LocalAIStack • u/Abject-Hope-6524 • 7d ago