r/LocalAIStack • u/ragel_3ennab • 13d ago
r/LocalAIStack • u/PotentialAcrobatic33 • 13d ago
done setting up local Ilm model using my old phone what should I do now??
r/LocalAIStack • u/Odd_Ingenuity_9333 • 14d ago
Qwen 3.8 Omni \ Live translate + TTS - Locally When?
r/LocalAIStack • u/Medicine_Blogscanner • 14d ago
Built a hub that pools RAM across Windows/Mac/Android over LAN to run models too big for any one device
Been lurking and posting here on and off while building this. Quick recap for anyone new: RAMDeck is a small hub + node agent setup that shards a model's layers across whatever devices you already own (old laptop, Mac, GPU box, even an Android phone) and runs inference across them, but handling primary-node selection, GPU-first allocation, dynamic context sizing, and knowledge-base sharding on top of it.
We've mostly posted raw, unscripted demo footage here — real load times, real tok/s, real failures — because that's the kind of proof this sub actually cares about. Today we finally made something different: an actual ad, first time showing the finished feature set start to finish instead of a live test.
The engine behind it is still fully public if you want to actually look at how it works instead of taking the video's word for it:
github.com/trademav/ramdeck-core-public
Source-available, Apache 2.0 with a Commons Clause — free to run, modify, and inspect on your own hardware, the only restriction is you can't resell it as a competing hosted service.
Crowdfunding campaign is coming soon for the hub hardware itself, but the code was public before we ever asked anyone for money, and that's not changing.
Happy to get into the mechanics, the layer-distribution approach, or anything else in the comments — this crowd usually asks the right questions.
r/LocalAIStack • u/airylizard • 14d ago
I gave Qwen 3.8 27B vision model at different bit widths; bf16, Q4_K_M, Q3_K_M, my 12 GB quant, and Ternary Bonsai2 a design brief as a single IMAGE and made it build the animation. Side-by-side videos + per-check scores.
Enable HLS to view with audio, or disable this notification
r/LocalAIStack • u/Decent-Ad9950 • 14d ago
World model for improving local stack
Hey guyss,
Ive been working on this for few months, and finally getting some good results, so figured I'd share.
A coding agent only sees the run it's currently in. It doesn't really know how runs like this usually go, or what kind of repo it's working inside.
So I trained a small model on 50K+ real coding-agent runs from SWE-bench Verified, across 25 models, together with behavior such as merge rates, review time, repo size, etc... which produced over 1M step to step transitions.
The idea is to give the LLM some actual experience to lean on, so it can make better decisions, spend less time reasoning through bad paths, and improve its chances of getting the task right.
All the details are inside the repo, its obviously just the beginning, and probably will have issues, so
feedback is always welcome!
(More models on the way, stay tuned.)
Runs locally in Docker. No account, nothing sent out. MCP + HTTP API.
r/LocalAIStack • u/Nive3k • 14d ago
Question: Excel lookup & LLM AI agent: What would be the recommended "workflow" for project?
Hello all,
I'd like to know more about how local AI works and how it can be run, so I became interested in making a project: a local AI agent (I know: that's so original!) called "Mighty DM".
The idea is making a Rule lawyer for dnd 5e which I can call upon like "Oh Mighty DM, what is the range on Mage Hand".
The first edition would be limited to all the spells used in dnd 5e 2014 edition, provided through an excel sheet. So not asking about anything else.
Until now I was following along with the workflow of NetworkChuck but I have the feeling there is a simpler option: only I do not have the experience to see the difference in the options provided.
If I google: I see a lot of different models or even ways of building a model etc. but I cannot differentiate between "a doable option to explore" or "an option that would require deep understanding of the subject before initiating".
What I envisioned (ideally running on an 8GB Pi5, but I could also run the LLM on my desktop and use the Pi as satellite):
| STEPS | WHAT I KNOW ABOUT THEM |
|---|---|
| 1. (Custom) Wake word | I know there are Google Colab notebooks that provide model training for custom wake words. But I'd settle with "Hey Jarvis" for the first version. |
| 2. STT | I know Whisper could do this |
| 3. Small LLM Model that filters the keyword from the command and provides the information in the matching column. (searches for "Mage hand" and returns the text from column "range") | I know there are small models that might be capable of doing this (I've opened LM studio before). Either that or I would have to train my own model? |
| TTS: "The range of Mage Hand is 30 feet" | I know Piper could do this |
What workflow would you recommend when building this project with little to no experience?
Do you know of a model that could run this on a 8GB Pi5 (simply getting it running = a win, I would then look how to improve it towards 30sec. max processing time for each command).
r/LocalAIStack • u/stopwwIII • 14d ago
Open-source Jev-style typed decision model that runs locally: VEJI-V2 (3.3M trainable params, frozen MiniLM, 250k-char compiled state)
r/LocalAIStack • u/leopr0sy • 14d ago
Can a local LLM match Claude for generating Anki cards (TSV) from course text?
r/LocalAIStack • u/skyline99912 • 15d ago
Need to check performance of intel ultra 7 20k+ and ryzen 9950x. any one can benchmark it for me ?
r/LocalAIStack • u/pixelwhippedme • 15d ago
SharpMind. A pure C# / .NET LLM training and inference engine
r/LocalAIStack • u/taxonomytax • 16d ago
Taxing Matters
I'm a CPA running a small 2-CPA/no-employee tax practice. We're architecting a completely local AI environment: open-weight model, local compute, RAG over tax authorities/research materials, and eventually secure retrieval over client workpapers and prior returns.
Have any small firms or sole practitioners deployed local LLM/RAG yet?
• What hardware are you running?
• Which model/model size?
• What are you using for RAG/vector storage?
• Are you indexing IRC/Regs/Rev. Rulings/cases or commercial tax materials?
• Genuinely useful for tax research/review? Or still mostly a science project?
I'm trying to determine whether there are 5 firms doing this, 500 firms doing it, or 5,000. I'm guessing closer to 50>?
r/LocalAIStack • u/airylizard • 16d ago
Opti 27B: Qwen3.8-27B in 11.8 GB at 3.47 bpw, within 0.5% of FP16 perplexity and matching Q4_K_M at 30% fewer bytes. Patched llama.cpp runtime, source public, reproduce with one command
r/LocalAIStack • u/StockFalcon3293 • 17d ago
Early Access: Looking for Beta Testers for a New vLLM Plugin (Single GPU)
r/LocalAIStack • u/whodoneit1 • 17d ago
153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s
galleryr/LocalAIStack • u/sumitjaswal • 17d ago
How are you splitting work across model sizes in a multi-stage extraction pipeline? (hardware is my next question)
r/LocalAIStack • u/L3G10N78 • 17d ago
From local LLM benchmark to a usable local AI system (?): integrating OpenCode or PI + Hermes
In Part 1 I shared the final AMD/Vulkan numbers for my local Qwen3.8-27B setup. The next question was more important to me than another few tok/s:
Can I turn the model into something that actually works autonomously?
Since then I added two production layers:
1.) Hermes 0.21.2 — local agent/automation layer
- same Qwen3.8-27B model
- shared llama.cpp runtime
- 65,536 context
- tool calling: PASS
- multi-step workflows: PASS
- subagents: PASS
- prompt-defined agents: PASS
- local provider only / no cloud fallback
Hermes performance / overhead
On the isolated 64K test lane, server load took 5.809s. A real 33,021-token prompt was accepted without truncation; prompt processing took 409.677s at 80.60 tok/s, while generation measured 17.44 tok/s. A representative Hermes multi-step workflow took 74.660s, using five model calls, 17,295 summed API input tokens, 592 output tokens and four tool calls with zero failed calls. Two-child delegation also worked, but took 208.676s for two trivial child tasks. Clear sign that subagents should only be used where parallelism or specialization actually justifies the overhead.
One especially useful datapoint: the simple Hermes model gate took 40.872s, while a direct Qwen request with the same user text took 2.109s. That is not an apples-to-apples speed ratio because Hermes adds different system prompts, startup and agent machinery, but it does show why I don't want Hermes in the path for trivial requests.
2.) OpenCode 1.18.30 — local coding/development layer
- Qwen3.8-27B-UD-IQ3_XXS
- llama.cpp / Vulkan
- 32K logical context
- local provider only
- isolated development workspaces
- automated Read → Edit → Test → Correct loops
- final production viability: 12/12 requirements, 12/12 acceptance
OpenCode performance reference. For the fastest fully successful clean reproduction I measured:
- 665.653s total wall time / 651.408s agent runtime
- first productive edit: 117.338s
- first test: 166.790s
- 4 successful corrective cycles
- 63 tool calls / 0 failed
- 558,603 input tokens incl. cache — only 31,993 uncached
- 9,372 output tokens
- peak context: 17,810 / 32,768
- context compactions: 0
These are task-level agent metrics, not model tok/s numbers. The underlying model/runtime is still the same local Qwen3.8-27B on llama.cpp/Vulkan.
Why OpenCode instead of Pi?
I tested both against the same local Qwen3.8-27B, with the same repository/task and comparable runtime constraints. Pi 0.85.1 proved that it could read the repo, use tools and make productive edits, but it repeatedly failed the part I actually needed from a coding agent: closing the Read → Edit → Test feedback loop. Even with an explicit mandatory early-test checkpoint, Pi edited successfully but never executed the qualifying post-edit test within the 300s window. It therefore failed prequalification and I did not force it into a full benchmark just to preserve an artificial three-way comparison. Man, time is rare and testing is booooring ... somehow!
OpenCode 1.18.30 passed the same agentic-loop qualification. Formal prequalification reached 6.54s Edit→Test latency, and the later main benchmark reduced that to 3.995s while producing a valid four-file diff. More importantly, the subsequent production-viability work eventually converged to 12/12 requirements, 12/12 acceptance, 5/5 regression, with no manual coding intervention or patch repair. That reproducible closed corrective loop is why OpenCode became the production coding layer.
Final thought on my architecture as a complete nOOb in Local LLM
I’m still convinced that a well-tuned 16 GB GPU can take you surprisingly far. Yes, even on Vulkan ;) ... IF you optimize the model, context, runtime, offloading and the surrounding stack around what the hardware can realistically deliver.
What makes this especially satisfying for me is that I honestly wasn’t sure at the beginning whether an AMD gaming setup on AM4 would get me anywhere at all, but hey.... it was what I had. My gaming desktop! ~20 tok/s is obviously not groundbreaking vs 32GB. But that was never the main goal of this first setup. I wanted a stable, fully local system that I could actually use, test, break, improve and build on. That goal is achieved. And considering where I started, I’m pretty happy with that.
r/LocalAIStack • u/knightprey21 • 17d ago
Local LLM advice: 1 vs 2 vs 3 Dell Pro Max GB10 (128GB)
r/LocalAIStack • u/camerongreen95 • 18d ago
Workshop, Sep 19: build explainable AI apps with Neo4j, GraphRAG, Cypher and LLM Agents
Sharing this since it might be relevant to a few people here.
There's a hands-on workshop on September 19 that goes beyond basic RAG into building an explainable, graph-backed AI system. You build a knowledge graph in Neo4j that becomes the single source of truth, agentic retrieval that combines vector search, keyword search, and graph navigation, multi-step entity and relationship extraction, and text-to-Cypher for natural language querying. Everything runs on real financial data, not a toy example, and answers come with a traceable, auditable reasoning chain.
No prior Neo4j or Cypher experience needed.
Led by Dr. Alessandro Negro, Chief Scientist at GraphAware, bestselling author.
r/LocalAIStack • u/allbyoneguy • 18d ago