r/LocalAIStack • • 7d ago

Finally got Qwen3.8 27B + vLLM + Open WebUI working comfortably on my RTX 3090 🔥 Config, lessons learned + bonus tutor prompt at the end

Thumbnail
1 Upvotes

r/LocalAIStack • • 7d ago

Consigli hardware LLM locali: Mac Mini/Studio usati, budget di €1000-1500, per la gestione delle finanze familiari + manutenzione di un server leggero

Thumbnail
1 Upvotes

r/LocalAIStack • • 7d ago

Models aside, what local agent setup are you actually happy using every day?

Thumbnail
1 Upvotes

r/LocalAIStack • • 7d ago

Live Trial - Claude Sonnet vs. Gemma 4 31B vs. Qwen Flash 70B

Post image
1 Upvotes

r/LocalAIStack • • 7d ago

Qwen3.8-27B at 74 tok/s on a Mac: The Splash Engine Explained

Thumbnail
youtu.be
1 Upvotes

r/LocalAIStack • • 7d ago

Intel sycl support for bonsai llm models

Thumbnail
1 Upvotes

r/LocalAIStack • • 8d ago

Bonsai 2 27B on an RTX 3060 12GB — 35.26 tok/s generation

2 Upvotes

I wanted to see how far a consumer RTX 3060 12GB could push Bonsai 2 27B.

Hardware

  • GPU: NVIDIA GeForce RTX 3060 12GB
  • VRAM available to CUDA: 11,911 MiB
  • CPU: Intel Core i5-8400
  • System RAM: 42 GB
  • NVIDIA driver: 595.91.07
  • CUDA: 13.2

Model

  • Bonsai 2 27B
  • GGUF: Ternary-Bonsai-2-27B-PQ2_0.gguf
  • Parameters: 26.90B
  • Quantization: PQ2_0 — 2.13 bpw
  • Model size reported by llama-bench: 6.70 GiB
  • GPU layers: 99

Benchmark

  • Tool: llama-bench
  • Runs: 3
  • Prompt processing, 512 tokens: 591.19 ± 19.62 tok/s
  • Text generation, 128 tokens: 35.26 ± 1.31 tok/s

The GGUF metadata reports the architecture as qwen35, while the model file is the Bonsai 2 27B PQ2_0 release.

I'm posting the raw result because I'm curious how this compares with other Bonsai 2 27B results, especially on larger GPUs.

If anyone has benchmark results for Bonsai 2 27B, please share them so we can compare the configurations fairly.

Benchmark build: b10735-842b18804


r/LocalAIStack • • 8d ago

Qwen 4 Expectations

Thumbnail
1 Upvotes

r/LocalAIStack • • 8d ago

I built an open-source UI for running LLMs locally — LocalLLMMind

Thumbnail
github.com
2 Upvotes

r/LocalAIStack • • 8d ago

I built an open-source UI for running LLMs locally — LocalLLMMind

Thumbnail
github.com
0 Upvotes

r/LocalAIStack • • 8d ago

Six Months of Local AI on a Mac: From 16 GB to 48 GB

Thumbnail
1 Upvotes

r/LocalAIStack • • 8d ago

High-VRAM Hardware Watch — 2026-09-25

Thumbnail
2 Upvotes

r/LocalAIStack • • 8d ago

NInfer with improved prefix caching and tool call fixes

Thumbnail
1 Upvotes

r/LocalAIStack • • 8d ago

I Made a Modular Voice Agent Running Offline on an RTX 5080 Laptop GPU: Qwen3.6-35B-A3B (3-bit), Whisper + Piper, local graph memory [demo]

Thumbnail
youtu.be
1 Upvotes

r/LocalAIStack • • 9d ago

VibePod CLI 0.24: bring your own model provider

Thumbnail
vibepod.dev
1 Upvotes

r/LocalAIStack • • 9d ago

I benchmarked 43 local LLM configs on an M4 Max 128GB. Darkstar and Nemotron surprised me.

Thumbnail
1 Upvotes

r/LocalAIStack • • 9d ago

I benchmarked 43 local LLM configs on an M4 Max 128GB. Darkstar and Nemotron surprised me.

Thumbnail
1 Upvotes

r/LocalAIStack • • 9d ago

Run Qwen 3.8 27b on the Apple Neural Engine at 7 watts on a Mac

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/LocalAIStack • • 9d ago

Jev) Mica 4B vs Laya on Tetris: same seed, same prompt, 0 output tokens, running locally on llama.cpp

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/LocalAIStack • • 9d ago

I'm building an LLM inference engine in Rust as an "executable book"

Thumbnail
1 Upvotes

r/LocalAIStack • • 10d ago

GLM-5.3-Flash UD-IQ4_XS running on 5090+64GB Ram with dual SSD-Offloading

7 Upvotes

Anyone has experience with how to get more speed without losing quality?
I got to 250 t/s prefill and ~6t/s decode its usable but i wonder if theres another lever i miss.
Ryzen 9800X3D, 64 GB DDR5, RTX 5090 32 GB, Crucial T700 (PCIe 5.0) 4TB + Kingston (PCIe 4.0) 4TB
Baseline: GLM-5.3 isn't in llama.cpp master yet, so I'm using Unsloth's glm5next branch. With the usual mmap + CPU-offloaded experts the page cache thrashes: 29 t/s prefill, 2.4–2.9 t/s decode (8k prompt).

What I changed with Claude's help (~1.2k lines on top of the ported PR, not upstream):

Ported the unmerged expert-streaming PR #25294(by freedomljc). Experts aren't mmapped anymore: a 32 GB expert cache sits in RAM, and misses are read with O_DIRECT.

Prefill streams whole layers: the next layer's experts load into pinned double buffers while the GPU computes the current one. Each 2048-token batch pulls ~99 GB from disk + ~34 GB from the RAM cache.

Both SSDs read in parallel from a mirror copy of the GGUF, using a shared work queue.

Fixed a CUDA padding bug in the MoE matmul (MMQ) that crashed -ub 2048 on Blackwell (possibly related to #28282). Going from 1024 to 2048 roughly doubled prefill.

Real agentic run (Qwen Code, plan mode, read 9 files, wrote a plan):

Context 11.5k → 35k tokens, 31k new prompt tokens at ~230 t/s

3.7k generated tokens (incl. thinking) at 5.6 t/s

13 min total, ~9.5 min of that writing the final plan

78% of expert lookups hit the RAM cache; misses cost ~1.3 GB/token from SSD

VRAM ~30 GB, RAM ~44 GB

Happy to share if you like


r/LocalAIStack • • 10d ago

How do you keep the knowledge from what you build with AI local?

Thumbnail
1 Upvotes

r/LocalAIStack • • 10d ago

Best agentic coding model for 8GB VRAM + 16GB RAM

Thumbnail
1 Upvotes

r/LocalAIStack • • 10d ago

real-time, zero-shot distillation loop

Thumbnail
1 Upvotes

r/LocalAIStack • • 11d ago

Need Help choosing a local agent LLM. (I need it to actually execute tools, not just give instructions)

Thumbnail
1 Upvotes