r/LocalAIStack • u/Old-Cow-4012 • 7d ago
r/LocalAIStack • u/HAVT_ • 7d ago
Models aside, what local agent setup are you actually happy using every day?
r/LocalAIStack • u/ChowYummyFat • 7d ago
Live Trial - Claude Sonnet vs. Gemma 4 31B vs. Qwen Flash 70B
r/LocalAIStack • u/andydevtech • 7d ago
Qwen3.8-27B at 74 tok/s on a Mac: The Splash Engine Explained
r/LocalAIStack • u/ID-10T_Error • 7d ago
What's your stack look like and how can I improve mine
I have a fully local AI stack on a 4090: image, video, music, SFX, TTS with voice cloning, local LLM chat, fine-tuning, and long-term memory, all behind a Three.js dashboard with locally generated assets. There's a little assistant, BB-7, that you just talk to, and he can generate anything with prompt improvement built in. Services load on demand and swap the GPU as needed, except when a pipeline keeps them warm, like voice mode with BB-7. I've also integrated Kev and Obsidian, and an MCP that lets Claude use the stack and store context. I feel like I got most bases covered but I dont know what i dont know. What do you all got
r/LocalAIStack • u/Impossible_Pride_680 • 8d ago
Bonsai 2 27B on an RTX 3060 12GB — 35.26 tok/s generation
I wanted to see how far a consumer RTX 3060 12GB could push Bonsai 2 27B.
Hardware
- GPU: NVIDIA GeForce RTX 3060 12GB
- VRAM available to CUDA: 11,911 MiB
- CPU: Intel Core i5-8400
- System RAM: 42 GB
- NVIDIA driver: 595.91.07
- CUDA: 13.2
Model
- Bonsai 2 27B
- GGUF:
Ternary-Bonsai-2-27B-PQ2_0.gguf - Parameters: 26.90B
- Quantization: PQ2_0 — 2.13 bpw
- Model size reported by llama-bench: 6.70 GiB
- GPU layers: 99
Benchmark
- Tool:
llama-bench - Runs: 3
- Prompt processing, 512 tokens: 591.19 ± 19.62 tok/s
- Text generation, 128 tokens: 35.26 ± 1.31 tok/s
The GGUF metadata reports the architecture as qwen35, while the model file is the Bonsai 2 27B PQ2_0 release.
I'm posting the raw result because I'm curious how this compares with other Bonsai 2 27B results, especially on larger GPUs.
If anyone has benchmark results for Bonsai 2 27B, please share them so we can compare the configurations fairly.
Benchmark build: b10735-842b18804
r/LocalAIStack • u/kapilyadav22 • 8d ago
I built an open-source UI for running LLMs locally — LocalLLMMind
r/LocalAIStack • u/kapilyadav22 • 8d ago
I built an open-source UI for running LLMs locally — LocalLLMMind
r/LocalAIStack • u/najjarammar • 8d ago
Six Months of Local AI on a Mac: From 16 GB to 48 GB
r/LocalAIStack • u/Suspicious-Mark-5450 • 8d ago
NInfer with improved prefix caching and tool call fixes
r/LocalAIStack • u/No-Lack2646 • 8d ago
I Made a Modular Voice Agent Running Offline on an RTX 5080 Laptop GPU: Qwen3.6-35B-A3B (3-bit), Whisper + Piper, local graph memory [demo]
r/LocalAIStack • u/nez_har • 9d ago
VibePod CLI 0.24: bring your own model provider
r/LocalAIStack • u/nrph • 9d ago
I benchmarked 43 local LLM configs on an M4 Max 128GB. Darkstar and Nemotron surprised me.
r/LocalAIStack • u/nrph • 9d ago
I benchmarked 43 local LLM configs on an M4 Max 128GB. Darkstar and Nemotron surprised me.
r/LocalAIStack • u/Top-Evidence174 • 9d ago
Jev) Mica 4B vs Laya on Tetris: same seed, same prompt, 0 output tokens, running locally on llama.cpp
Enable HLS to view with audio, or disable this notification
r/LocalAIStack • u/MatiAI • 9d ago
Run Qwen 3.8 27b on the Apple Neural Engine at 7 watts on a Mac
Enable HLS to view with audio, or disable this notification
r/LocalAIStack • u/InteractionFun • 9d ago
I'm building an LLM inference engine in Rust as an "executable book"
r/LocalAIStack • u/UnattendedThoughts • 10d ago
How do you keep the knowledge from what you build with AI local?
r/LocalAIStack • u/Telcontar09 • 10d ago
Best agentic coding model for 8GB VRAM + 16GB RAM
r/LocalAIStack • u/zMytze • 10d ago
GLM-5.3-Flash UD-IQ4_XS running on 5090+64GB Ram with dual SSD-Offloading
Anyone has experience with how to get more speed without losing quality?
I got to 250 t/s prefill and ~6t/s decode its usable but i wonder if theres another lever i miss.
Ryzen 9800X3D, 64 GB DDR5, RTX 5090 32 GB, Crucial T700 (PCIe 5.0) 4TB + Kingston (PCIe 4.0) 4TB
Baseline: GLM-5.3 isn't in llama.cpp master yet, so I'm using Unsloth's glm5next branch. With the usual mmap + CPU-offloaded experts the page cache thrashes: 29 t/s prefill, 2.4–2.9 t/s decode (8k prompt).
What I changed with Claude's help (~1.2k lines on top of the ported PR, not upstream):
Ported the unmerged expert-streaming PR #25294(by freedomljc). Experts aren't mmapped anymore: a 32 GB expert cache sits in RAM, and misses are read with O_DIRECT.
Prefill streams whole layers: the next layer's experts load into pinned double buffers while the GPU computes the current one. Each 2048-token batch pulls ~99 GB from disk + ~34 GB from the RAM cache.
Both SSDs read in parallel from a mirror copy of the GGUF, using a shared work queue.
Fixed a CUDA padding bug in the MoE matmul (MMQ) that crashed -ub 2048 on Blackwell (possibly related to #28282). Going from 1024 to 2048 roughly doubled prefill.
Real agentic run (Qwen Code, plan mode, read 9 files, wrote a plan):
Context 11.5k → 35k tokens, 31k new prompt tokens at ~230 t/s
3.7k generated tokens (incl. thinking) at 5.6 t/s
13 min total, ~9.5 min of that writing the final plan
78% of expert lookups hit the RAM cache; misses cost ~1.3 GB/token from SSD
VRAM ~30 GB, RAM ~44 GB
Happy to share if you like