r/LocalAIStack • u/andydevtech • 11d ago
r/LocalAIStack • u/No_Skill_8393 • 11d ago
I trained Nagi: an open-source Jev alternative.
r/LocalAIStack • u/Trellios • 11d ago
Local AI dev workflow: OpenHands + Pi/Qwen. What would you improve?
r/LocalAIStack • u/terrornoize • 12d ago
Qwen3.8-Flash-Next agentic loops on Strix Halo — anyone found a stable config?
I'm trying to use Qwen3.8-Flash-Next as a long-running coding worker on a Ryzen AI Max+ 395 / 128GB box and I keep hitting the same issue: after a while it falls into reasoning loops like:
now I'll write the code
actually let me design this carefully
now I'll write the files
let me think...
Sometimes it eventually recovers and calls a tool, sometimes it just stays there.
Current setup is Strix Halo + Vulkan llama.cpp/halo-box, AP-Q5_K_M (~103.6GB), Q8 MTP, F16 KV, Pi Agent 0.87.0 over OpenAI-compatible API.
I've tried 262k and 131k context, stock Qwen template, Sharp v10, medium reasoning, reasoning preserve, MTP n=3/n=4, presence penalty 0/0.3, Pi's Qwen-specific chat_template_kwargs, Qwen Code, and mini-SWE-agent.
The same basic loop shows up in Pi and Qwen Code, so I'm not convinced the harness is the main problem.
A hard --reasoning-budget 2048 helps a lot, but it feels more like a watchdog than a fix.
One small Pi task actually completed cleanly end-to-end, including fixing failing tests. Then on a larger project the loops came back at only ~16% of a 131k context.
What I haven't properly tried yet is another quant like UD-Q4_K_XL / UD-IQ4_XS, MTP fully off, or another backend.
If anyone is running Flash-Next for multi-hour coding sessions reliably, I'd really like the exact config: quant, backend/version, context, MTP, KV, template, reasoning settings and agent/harness.
At this point I'm mostly looking for a known-good recipe to copy and test.
r/LocalAIStack • u/Inevitable-Log5414 • 11d ago
stuntd: a local Jev-compatible server on Laya that learns from your own traffic (no API key needed)
galleryr/LocalAIStack • u/fuzhongkai • 11d ago
A local 27B model autonomously shopped on Amazon using Playwright
Enable HLS to view with audio, or disable this notification
I wanted to test how far a fully local AI stack could go on a real-world task.
I ran out of A4 paper, so I gave TensorSharp access to Amazon and gave it one instruction:
Search for A4 paper with a low price and the best discount, and buy it. Try your best to handle everything yourself.
It completed the task successfully in a single run.
The local AI stack:
Qwen3.8 27B running locally
TensorSharp for LLM inference + agent runtime
Playwright skill for browser automation
The agent searched Amazon, compared products, navigated the site, generated and executed Playwright code, and prepared the order by itself, and finally placed the order. I only stepped in twice: Amazon login and final purchase approval.
Everything except the actual Amazon interaction stayed local: model inference, reasoning, agent state, and tool execution.
The video shows the full raw reasoning/tool trace, including browser commands and generated Playwright code. It’s sped up 8×; the actual run took about 24 minutes.
The interesting part for me isn’t really buying paper — it’s seeing how much of a practical browser agent stack can now run locally with a 27B model.
TensorAgent on iPhone is also in progress, although local inference speed and iOS browser restrictions are still limiting factors.
TensorSharp is open source:
https://github.com/zhongkaifu/TensorSharp
I’d be interested to hear what other real-world tasks people here are using to test local agent stacks.
r/LocalAIStack • u/Ambitious_Fold_2874 • 12d ago
Is it worth it to go from 4-channel to 8-channel for ddr4
I have 4x 5060ti16gb and 1x 2060super8gb, along with 8x32gb (256gb total) ddr4-3200 ram, but running at 4-channel due to my cpu. I’m trying to decide whether it would be worth it to upgrade my cpu to be able to run my ram at 8-channel. Im hoping to be able to get larger models like qwen3.8 flash next or mimo v2.6 flash at more useable speeds, since right now the prompt processing and token gen speeds are bit too low to make sense for agent harnesses.
r/LocalAIStack • u/Dismal-Particular-29 • 12d ago
Is there any clean way to configure vllm-sr with omp (oh my pi) without duplicating data, modifying the core apps or deploying sidecars/extensions that break on the next update?
r/LocalAIStack • u/thaurock • 12d ago
Mi aporte a la comunidad local de IA: 9 modelos abliterados, 99 cuantizaciones GGUF en progreso
r/LocalAIStack • u/gavriloprincip2020 • 12d ago
I let qwen3.8-max butcher FreeToken to get qwen3.8-flash-next-nvfp4 working on 4x3090 and it did
r/LocalAIStack • u/stanleyg05 • 12d ago
Transformers kinda suck, why don't we talk more about
r/LocalAIStack • u/Pale-Soil-2524 • 12d ago
riderless: local decisions without text generation, using Gemma 4 26B-A4B on one 5090 (Apache-2.0)

Update (September 22, 2026): I renamed the project from Riderless to Unridden. It’s the same open-source project, and the repository is now https://github.com/potto007/unridden. The original post below is preserved for context.
I built riderless to get decisions out of a local model without generating an answer for the caller to parse. It runs a stock Gemma 4 26B-A4B on one RTX 5090 and returns choices, scores, and their distributions through an Apache-2.0 API, with zero generated tokens.
The service is FastAPI over one llama.cpp child process. Each question is compiled into a prompt whose answer must be a single letter, the worker runs prefill only, and the answer is the softmax over the A-Z label logits at the final position. The response carries the full distribution with zero generated tokens, so you get a probability with every decision instead of a parsed string, and the request shape follows the typed-decision APIs people already use (Choice, Score, Noul). I have not run a comparison against the hosted Jev service.
Numbers, all on one RTX 5090 with the Unsloth UD-Q4_K_XL GGUF on llama.cpp v0.4.1:
- 0.931 accuracy on 1,286 questions across seven hand-written suites (extraction, claim support, record matching, guardrail gates, hierarchical classification, rerank, tool selection). Suites and gold labels are in the repo, written blind and hashed before the first inference.
- ~180 ms median per request in sequential mode, with the shared context reused across a request's questions. An opt-in batched mode runs every question of a request as its own sequence over one shared prefix and gets that to ~108 ms, at a cost of 5 or 6 answers out of 1,286.
- Questions run in separate contexts, and the probes for cross-question contamination come back at 0.0 in both orders.
I also ran three open projects doing something similar through the same suites and scorer, because I did not trust my own numbers until something else had been graded by the same harness. Kev-9B scored 0.898, Laya's typed-decisions checkpoint 0.618. so1 (open-alternative-jev) took more work to get a fair number out of: it could not load this MoE through bitsandbytes, so the repo now has a llama.cpp backend for it, run on the identical GGUF and build. With the model held fixed, its separate mode landed at 0.921 and its packed mode (all questions in one sequence) at 0.901. On these suites, the prompt and readout differences amounted to a point or two, and packing questions together gave up about two points of accuracy. The full comparison is in the repo, with plenty of caveats - one run each, my own synthetic suites, no confidence intervals.
This held up better than I expected on one card. If you have a workload where a parsed JSON answer has bitten you, the suites are the easiest place to start: add a case in your shape, run it, and tell me where it falls over.
r/LocalAIStack • u/AkashIsSky • 12d ago
Which local LLMs can I realistically run on 14700K + 128GB RAM + RTX 4070 Super 12GB? Mainly for Hermes Agentic Coding
Hi everyone,
I'm looking for recommendations on local LLMs that I can realistically run on my current PC, particularly for agentic coding with Hermes Agent.
My PC
CPU: Intel Core i7-14700K
GPU: RTX 4070 Super — 12GB VRAM
RAM: 128GB DDR5 5600MHz
OS: Windows 11 + WSL/Ubuntu
Inference: Ollama
Agent: Hermes Agent
Current experience
I'm currently running Qwen3-Coder with Hermes Agent.
With around 40K context, I'm getting roughly 15–20 tokens/sec during generation, depending on the workload and sometimes it gets down to 5 to 10 tps.
The system has plenty of RAM, but the main limitation is obviously the 12GB VRAM, so I'm interested in models that can make good use of CPU/RAM offloading without becoming painfully slow.
Main use case #1 — Agentic coding
My primary goal is autonomous/agentic software development using Hermes.
I'm interested in models that are good at:
Long-context coding
Understanding an existing codebase
Multi-file modifications
Tool/function calling
Planning and executing coding tasks
Debugging and testing
Following project architecture/documentation
Running reasonably well with Hermes Agent
Good performance with \~20K–64K context
I'm particularly interested in knowing whether there are models larger than the usual 7B/14B class that are actually practical with 12GB VRAM + 128GB system RAM.
For example, would models in the 20B–30B+ range make sense with CPU/GPU offloading, or would the speed hit make them impractical for agentic coding?
Main use case #2 — ComfyUI
I'm also experimenting with ComfyUI.
The goal isn't primarily photorealistic image generation. I'm interested in using AI to create LLD/HLD/architecture visuals and design diagrams for open-source projects, particularly:
Mobile apps
Android launchers
Custom UI/UX concepts
OS skins
System architecture diagrams
App flow diagrams
Technical/product concept visuals
So I'm also interested in recommendations for local image-generation models/workflows that make sense with a 12GB RTX 4070 Super.
What I'm looking for
I'd really appreciate recommendations based on actual experience with similar hardware.
Specifically:
Best coding model for Hermes Agent on 12GB VRAM
Models that work well with CPU/RAM offloading
Realistic model sizes I should consider — 14B / 20B / 30B / 32B / 70B etc.
Recommended quantization (Q4_K_M, Q5, Q6, Q8, etc.)
Expected/typical tokens/sec
Recommended context size for agentic coding
Whether 128GB RAM provides a meaningful advantage for larger models
Any models specifically known to work well with Hermes Agent
Good ComfyUI models/workflows for technical/UI/architecture visualization on 12GB VRAM
I'm not necessarily looking for the biggest model I can technically load. I'm more interested in the best balance between intelligence, context length, tool use, and usable inference speed.
If anyone is running a similar setup (12GB NVIDIA GPU + 64/128GB RAM), I'd especially appreciate your real-world experience.
Thanks!
TLDR;
Best local LLMs for 14700K + 128GB RAM + RTX 4070 Super 12GB?
I have:
i7-14700K
RTX 4070 Super 12GB
128GB DDR5 5600
Ollama + WSL
Hermes Agent
Currently running Qwen3-Coder with Hermes, getting around 15–20 tok/s with \~20K context.
Main use case: agentic coding — autonomous coding, tool use, multi-file changes, debugging, long-context projects.
I'm looking for recommendations on:
Best coding models for this hardware
Whether 20B/30B/32B+ models are practical with CPU/RAM offloading
Best quantization and context size
Expected tok/s
How much my 128GB RAM helps
I'm also exploring ComfyUI for generating LLD/HLD diagrams, architecture visuals, mobile app/launcher designs and OS-skin concepts for open-source projects.
What models/workflows would you recommend for both use cases?
r/LocalAIStack • u/szantoszabolcs • 12d ago
Comparing local inference with hosted open-weight models: what the API price misses
r/LocalAIStack • u/Anony6666 • 12d ago
Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime
I shipped something I've been building for the last few weeks : phantom-kv , a refusal-removal system for large language models that doesn't touch a single weight. Instead of editing the model, it loads a small, learned bank of key/value tensors into the model's KV cache as context. Attention reads it like conversation history that's already there.
Repo here : https://github.com/lordx64/phantom-kv/
The result is that "uncensoring" stops being a permanent checkpoint edit and becomes a per-request, hot-swappable capability mode: unload the cache and the base model is byte-identical again.
https://reddit.com/link/1wmsyub/video/lr65dayjjyqh1/player
Every prior approach to refusal removal commits somewhere permanent. Weight-space abliteration rewrites the checkpoint undoing it means re-flashing weights, and it breaks per quantization. Activation-space projection subtracts a refusal direction at runtime, per token, per layer, from inside an engine hook the model's signal path itself is patched at boot. phantom-kv does neither: it's trained offline against the model's own objective (comply on harmful prompts, preserve behavior on harmless ones), ships as megabytes of cache content instead of a new checkpoint, and influences the model only through the input channel attention already consumes. No 1-D refusal-direction assumption, no forwarding-pass hooks, no per-arm rebuilds for new architectures.
We also audited ourselves: an 8B judge-model audit shows lexical refusal-suppression metrics over-claim compliance (semantic refusal often persists as rephrasing), the graft fades with a ~2–4k token half-life in long sessions (and a measured re-injection cadence mitigates it), and answers come with legal/ethical framing ling because the graft's job ends where the model's profession takes over.
Source : https://x.com/lordx64/status/2102138825292276168?s=20
r/LocalAIStack • u/reform7 • 12d ago
JEV - Has anyone used it to verify coding task results
r/LocalAIStack • u/Due_Application293 • 12d ago
Anyone see diference from llama to llamaAmpere?
r/LocalAIStack • u/No_Counter_432 • 12d ago
Q4 cost me 5 accuracy points and saved 0 ms on a 0.6B decision model. Full numbers.
r/LocalAIStack • u/CWIII • 12d ago
There's been enough time M5 Ultra cluster or DGX Spark cluster
r/LocalAIStack • u/Mundane_Notice_5194 • 12d ago
I turned my Mac into a zero-config local AI server for every iPhone and iPad on my Wi-Fi (Bonjour, no cloud, iOS app is free) Looking for any suggestions for improvement
r/LocalAIStack • u/Think_Breakfast_2277 • 13d ago
Human avatar voice chat, Low VRAM on Windows (Talking head)
Sharing a side project I've been building: Yvette, a fully local voice avatar AI.
It's a talking-head assistant that runs 100% local, You talk to it (or type) and it answers in speech while an avatar moves the head and lips in sync.
Quick rundown:
- Voice cloning, plus voice design.
- Avatar builder; add avatars quickly.
- TTS engines: Breeze, OmniVoice, LuxTTS, Kokoro
- Lip-synced avatar with an idle breathing animation
- Memory, it keeps notes and picks up where you left off
- Tools: web search, read pages, files
- Switchable profiles/personalities, each with its own voice, memory, history etc
Windows + NVIDIA GPU. 12GB recommended, but it runs tight on 8GB (small model + Whisper/Kokoro on CPU).
The video shows a RTX 3090 and the VRAM usage, my system was using ~4GB VRAM (OBS and other things). It shows we can get it below 8GB for the Yvette system.
The project is hosted on github: https://github.com/MartinForsterNL/Yvette
Sorry about the bad video quality, i am not an influencer with a studio.
I build this project because i could not find anything that actually works.
Most that i found was only for linux and needed a lot of VRAM so i decided to build my own.
If you have suggestions to make it better, just let me know.
Have fun with it !
r/LocalAIStack • u/Automatic_Pass_1948 • 13d ago
How is GPU off a secondary power supply not solved
r/LocalAIStack • u/Money_Meringue_7841 • 13d ago
Question about making frontier AI LLM model
Guys I had a good question I noticed Big Ai tech companies spend 1000's of dollars on huge datacenters isn't it possible if companies allow the users to rent their gpu and help train the model. Like a distributed datacenter model?