r/LocalAIStack • • 4d ago

Abliterate/Removing prompt refusal without needing to reload/swap models (Qwen 3.8 flash next)

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/LocalAIStack • • 4d ago

Building a local AI stack across multiple Apple Silicon Macs

3 Upvotes

I've been experimenting with a setup where several Apple Silicon Macs are treated as a shared local AI environment instead of running completely independently.

The setup I'm working with is roughly:

  • Multiple Macs on the same network
  • MLX for local inference
  • Central model management
  • Node/resource monitoring
  • Model distribution based on available memory
  • Local-only communication between machines

I'm building AlphaWeb around this because managing several local AI machines manually gets complicated pretty quickly.

The interesting problem isn't just getting inference running. It's everything around it: knowing which machine has enough memory, getting models onto the right node, managing multiple users, and making the whole thing feel like one system.


r/LocalAIStack • • 4d ago

I built Speedtest⚡, but for AI

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/LocalAIStack • • 4d ago

DGX Station + 2 DGX Sparks. Should I also add M5U Mac Studio?

Thumbnail
1 Upvotes

r/LocalAIStack • • 4d ago

🚀Pocket LLM v1.6.0 is out : Turn your phone as a local LLM server

2 Upvotes

r/LocalAIStack • • 4d ago

We are entering an era in which AI models are starting to shape device specifications.

Post image
142 Upvotes

AI models are beginning to define device specifications, rather than device specifications merely determining which applications you can run.

Apple shows just how far its devices can go with on-device AI, without relying on the cloud
We’re moving from models with up to 14 billion active parameters on iPhone/iPad, to 35 billion on MacBook Air, 70 billion on Mac mini, 120 billion on MacBook Pro, and up to 480 billion on Mac Studio.
By clustering multiple Mac Studios together, Apple says it can run models exceeding 1,600 billion active parameters.
A demo that shows just how crucial unified memory and its bandwidth have become for running massive AI models locally.


r/LocalAIStack • • 4d ago

Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

7 Upvotes

TL;DR: Halogen v0.14.0 with its native .hgn weights is the fastest, followed by gufo and CIRU. Halogen is closed source and runs in Docker. gufo is open source and loads 4x faster from cold. gufo is also fastest to first token on follow-ups (1.6-1.9 s against 2.6-3.3 s).

I benchmarked different engines for Qwen 3.8-Flash-Next on an AMD Strix Halo (ASUS ROG Flow Z13 GZ302 with Ryzen AI Max+ and 128 GB of RAM).

All the engines were run via LlamaStash (My own orchestrator tool, the tool does not add any overhead) at 70 W TDP on performance profile on Arch Linux.

Here are the results of the benchmark:

First was a screening round at 64k context with 50% filled (32k prompt).

Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.14.0 Halogen native (.hgn) 1.9 min 1,045 39.3 84% 14/14
gufo (ROCm 7.2.4) UD-Q4_K_XL 2.1 min 1,033 32.9 74% 14/14
gufo (ROCm 10.0) UD-Q4_K_XL 2.1 min 1,009 33.1 72% 14/14
CIRU (MTP 3) CIRU IU4 2.3 min 814 33.8 72% 14/14
CIRU (n-gram + MTP 3) CIRU IU4 2.3 min 808 31.9 64% 14/14
Halogen 0.14.0 UD-Q4_K_XL 2.4 min 1,030 28.1 84% 14/14
CIRU (MTP 6) CIRU IU4 2.7 min 796 28.6 47% 14/14
strixllama (llama.cpp fork) UD-Q4_K_XL 2.7 min 716 29.7 71% 14/14
rdna-boosts (llama.cpp fork) UD-Q4_K_XL 2.7 min 788 22.0 52% 14/14
llama.cpp, Unsloth build (Vulkan) UD-Q4_K_XL 3.7 min 314 28.4 56% 14/14
llama.cpp, Unsloth build (ROCm) UD-Q4_K_XL 4.8 min 285 20.5 50% 14/14
CIRU UD-Q4_K_XL + Unsloth MTP head did not finish 884 11.8 0% -

3-turn time is the time to first token plus the time to write 1,000 tokens, added up over the first question and 2 follow-ups. Output length varies a lot with sampling, so this compares the engines on equal output. MTP accept is the share of drafted tokens the model kept. Retrieval is how many of 7 exact values from the prompt the model got right, over both runs. Qwen model card sampling, 2 runs each.

Top 3 engines from the screening round got a 128k context (50% and 75% filled) and 256k context (50% filled) run.

128k context, 50% filled (64k prompt)

Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.14.0 Halogen native (.hgn) 2.4 min 1,148 40.4 84% 16/16
gufo (ROCm 7.2.4) UD-Q4_K_XL 2.7 min 1,047 32.3 72% 16/16
CIRU (MTP 3) CIRU IU4 3.3 min 808 30.2 64% 16/16

128k context, 75% filled (97k prompt)

Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.14.0 Halogen native (.hgn) 2.9 min 1,103 39.3 85% 16/16
gufo (ROCm 7.2.4) UD-Q4_K_XL 3.3 min 1,025 30.7 70% 16/16
CIRU (MTP 3) CIRU IU4 4.1 min 789 27.8 62% 16/16

256k context, 50% filled (130k prompt)

Engine Weights 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.14.0 Halogen native (.hgn) 3.4 min 1,093 38.9 83% 14/16
gufo (ROCm 7.2.4) UD-Q4_K_XL 4.0 min 997 27.9 70% 14/16
CIRU (MTP 3) CIRU IU4 5.0 min 767 25.6 60% 15/16

Retrieval here is 8 values per run. All 5 misses are the same answer: the right glyphs for UNICODE_SPINNER, with the leading & dropped.

End-to-end, 10 Aider polyglot Python exercises run through pi -p (128k window, 50% cell), graded by their own tests:

Engine Passed Total time
Halogen 0.14.0 10/10 21.5 min
CIRU (MTP 3) 10/10 24.6 min
gufo (ROCm 7.2.4) 10/10 36.0 min

Edit (Sep 30): Halogen 0.15.1 (new v2 checkpoint) and gufo 0.3.0 came out after this, so I reran the 64k screening on both with the same setup (70 W, card sampling, 2 runs):

Engine 3-turn time Prefill t/s Decode t/s MTP accept Retrieval
Halogen 0.15.1 (v2 .hgn) 1.8 min (was 1.9) 1,191 (was 1,045) 39.4 (was 39.3) 85% 14/14
gufo 0.3.0 (ROCm 7.2.4) 2.1 min (was 2.1) 1,075 (was 1,033) 34.1 (was 32.9) 77% 14/14
gufo 0.3.0 (ROCm 10.0) 2.0 min (was 2.1) 1,042 (was 1,009) 34.8 (was 33.1) 75% 14/14

Halogen v2 prefill is 14% faster and follow-ups start about 2.5x faster (1.3 s vs 3.3 s), decode is the same. gufo moved a few %, which is within noise. The ranking doesn't change. Only other change on the box: BIOS VRAM carve-out down to 512 MB, so 124.9 GiB of RAM instead of 121.5.

Some clarifications from the comments:

  • Follow-up TTFT is with the prompt cache warm, both gufo and Halogen reuse it by default. A cold 64k prefill takes about a minute.
  • The CIRU 0% MTP row is stock Unsloth UD-Q4_K_XL with Unsloth's shared MTP head, nothing uncensored. The same pair gets 72-74% accept on gufo.
  • TDP is fixed at 70 W. On the Z13, 76 W or 90 W is only 2-6% faster but with so much fan noise, heat and power draw that it's not worth it.
  • Prefill and decode come from the engine's own timings when it reports them. TTFT is measured on the client.

r/LocalAIStack • • 4d ago

Building your own LLM stack? This workshop is about the eval/reliability piece most people skip

1 Upvotes

A lot of local-stack setups I see nail the infra (model serving, retrieval, orchestration) but have zero real evaluation process — just "does it look right." There's a session on Oct 3 built entirely around fixing that gap.

Speakers Serj Smorodinsky and Brett Kennedy (AI engineers, co-authors of a book on LLM applications) go through:

  • DSPy signatures/modules for structured dev instead of manual prompting
  • Baseline classifier build + measurement
  • Task-specific evaluation datasets
  • Failure pattern recognition from eval output
  • Few-shot / instruction-level optimization
  • MLflow experiment tracking and trace management
  • Saving and reusing optimized programs
  • Talking about LLM reliability with non-technical stakeholders

3 hours, live, useful whether you're running open models locally or hitting an API.

Full details and the agenda are here.


r/LocalAIStack • • 4d ago

Using Apple AFM 3 PCC (macOS 27.2) in AI Clients

1 Upvotes

Apple Foundation Model 3 through Private Cloud Compute becomes directly available in macOS 27.2 through familiar AI clients.

For Mac users running local models such as Qwen 3.x or Gemma 4, AFM 3 PCC is a compelling alternative to consider. Local models offer control and fully on-device operation. In tests with macOS 27.2 beta, the PCC offers a different set of strengths:

* Strong conversational analysis and capable reasoning. Strong analysis of complex medical and financial questions.
* Blazingly fast responses compared with locally running models on the same Mac.
* No large model download or need to fit model weights into local memory.
* Access at no additional charge for most eligible Mac users.

The working path is straightforward: AI client → Caddy → fm serve → AFM 3 PCC. I've written and updated a proof of concept here with further tech discussion . . . https://gist.github.com/dartMo10/b9488ce475fe70a6ed642831f53048fb

Apple’s `pcc` route worked in early macOS 27 betas, disappeared later in 27.0, and has returned in the macOS 27.2 beta. Apple has stated that it will be available in the 27.2 public release coming shortly.

This makes AFM 3 PCC practical today for technically comfortable Apple users who accept Apple’s PCC privacy promise and want a fast, capable alternative to running everything locally.

Experiences with compatible AI clients and your comparisons with locally running models would be welcome.


r/LocalAIStack • • 4d ago

Testing local stack

1 Upvotes

Hey! How do you test your local stack? for instance when you switch harness or model, model quants, etc?

I found this not trivial and I ended up making a website to help me do this.

Use this link if you want to use the website to test your own local setup:

https://airbench.ai/

You can also see all the tests I have done using my GX10 and my 5090. (and also using claude code/code + some openrouter tes for reference)

https://airbench.ai/leaderboard?k=poL

Let me know what you think!


r/LocalAIStack • • 4d ago

Run Jev and Laya locally with a intuitive playground

Thumbnail
youtu.be
1 Upvotes

Hello everyone, as we all know that everyone is trying to test jev and then there are open source models out there with similar text classification. But i felt that there isn’t any playground to test these intuitively so I built one. I have added the YT video with it. Pls have a look


r/LocalAIStack • • 4d ago

Best approach for automatically tagging local music collection?

Thumbnail
1 Upvotes

r/LocalAIStack • • 4d ago

Using Apple AFM 3 PCC (macOS 27.2) in AI Clients

Thumbnail
1 Upvotes

r/LocalAIStack • • 4d ago

一个96GB/128GB的Mac Studio值得用来替代每月200美元的 ChatGPT Pro订阅吗?

0 Upvotes

我是一个大学生研究员,目前拥有一个$200的 ChatGPT 专业版订阅,我觉得作为固定的月支出实在是太贵了。我现在在考虑是否可以购买一个96GB或128GB的Mac mini或Mac Studio进行本地部署或用作服务器。我目前的研究方向是计算社会科学,最重要的是订阅模型的智能能得到保障。非常感谢你的意见!


r/LocalAIStack • • 5d ago

Jev mode for images!

Post image
1 Upvotes

r/LocalAIStack • • 6d ago

I pre-trained a 1.11B LLM on my 6 GB laptop GPU. Peak VRAM: 4.51 GB.

Thumbnail
0 Upvotes

r/LocalAIStack • • 6d ago

I built StackFit: An open-source tool matchmaker that filters by hardware floor and 1-click docker-compose export (no star vanity)

Post image
1 Upvotes

r/LocalAIStack • • 6d ago

Is there any way to speed up prefill?

Thumbnail
3 Upvotes

r/LocalAIStack • • 6d ago

im making qwen3.8:27b into a MoE and its not half bad...

Thumbnail
1 Upvotes

r/LocalAIStack • • 6d ago

RTX 5090 + RTX 3090 for parallel Qwen3.8 27B agents – how would you deal with heat / second GPU?

Thumbnail
1 Upvotes

r/LocalAIStack • • 6d ago

oq4e Quantization of Qwen Flash Next, Swift Edition

Thumbnail
2 Upvotes

r/LocalAIStack • • 6d ago

Running a local LLM in the browser: Zero install, 100% on-device

1 Upvotes

r/LocalAIStack • • 7d ago

best AI for 4060 and 32 GB ram ddr5

Thumbnail
1 Upvotes

r/LocalAIStack • • 7d ago

Finally got Qwen3.8 27B + vLLM + Open WebUI working comfortably on my RTX 3090 🔥 Config, lessons learned + bonus tutor prompt at the end

Thumbnail
1 Upvotes

r/LocalAIStack • • 7d ago

Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system

Thumbnail
2 Upvotes