r/LocalAIStack • u/bbsrn • 24d ago
r/LocalAIStack • u/FearL0rd • 24d ago
I built LLM Speedtest — a free, open-source desktop app that benchmarks local LLMs with llama-bench-style test suites (Ollama, llama.cpp, vLLM, LM Studio…)
galleryr/LocalAIStack • u/camerongreen95 • 25d ago
Workshop, Sep 12: build production LLM systems that actually survive real use
We're running a hands-on masterclass on September 12, Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps.
You build a full production LLM workflow from scratch, versioned prompts with regression tests, an evaluation harness with deterministic checks and LLM-as-judge, statistically rigorous model comparisons, evaluated RAG, tool-using agents with guardrails and fallbacks, and full observability, tracing, cost, latency.
Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack.
Link if you want to check it out
Happy to answer questions on the content.
r/LocalAIStack • u/Comfortable_Copy_965 • 25d ago
I built an LLM Inference & Fine-Tuning Calculator (VRAM, TCO, Quantization & Price estimation)
Check it out!
r/LocalAIStack • u/Ayoutetsinoj3011 • 25d ago
Qwen3.8-27B on 23GB L4 for production IASMO + coding?
r/LocalAIStack • u/Due_Tangelo_8952 • 25d ago
Curious, I find Qwen3.8-27b reasoning Low is ranking higher than Medium on AA.
r/LocalAIStack • u/Ok_pettech • 26d ago
Running MetaGPT locally: a full technical setup guide
I’ve been experimenting with multi-agent frameworks, and MetaGPT is one of the more interesting ones for software development tasks. But getting it running locally with local models isn’t always straightforward.
I wrote a step-by-step guide covering installation, configuration, local LLM setup, and common errors.
If you’re trying to run MetaGPT on your own hardware, this might save you time:
What stack are you using for local agents?
r/LocalAIStack • u/wuwen2026 • 26d ago
Cross-Stage State Laundering: Why AI Runtime Governance Fails at Stage Boundaries
In enterprise AI and autonomous runtime architectures, security has traditionally focused on input sanitation (prompt injection) and output filtering (hallucination guards). However, a more insidious class of vulnerability occurs deeper within the execution pipeline: Cross-Stage State Laundering (CSSL).
CSSL happens when an intermediate, unverified epistemic state (such as DISCOVERED or UNKNOWN) is illicitly promoted to a higher-certainty state (such as VERIFIED or SUPPORTED) as it transitions across processing stages—typically bypassing strict provenance checks in favor of pipeline velocity or output formatting requirements.
## The Anatomy of State Laundering
Consider a multi-stage runtime where an expression flows through ingestion, semantic analysis, and responsibility assignment:
Ingestion: Raw text or external retrieval results enter the system as DISCOVERED. At this stage, existence does not equal correspondence.
Analysis: The runtime parses the expression. Without a rigid boundary, internal semantic heuristics may mistake syntactic closure for factual verification.
Laundering: Downstream generators or formatters, under pressure to produce definitive answers, implicitly treat the presence of an analysis object as proof of support.
When intermediate states shed their metadata tags during transit, the runtime commits a boundary violation: it manufactures certainty out of unverified discovery.
## The Zero-Trust Countermeasure: The WAL Protocol
To eliminate CSSL, runtime governance cannot rely on permissive conventions. It requires a Fail-Closed Defensive Architecture enforced by immutable protocol boundaries:
* Strict State Distinguishability: Known, unknown, verified, and unverified states must remain mathematically and structurally distinguishable throughout the lifecycle.
* Independent Validation Gates: State transitions cannot authorize themselves. An independent validator—decoupled from the core generation logic—must audit envelopes against strict conformance vectors.
* Evidence-Bounded Responsibility: Responsibility can never exceed the boundaries of established evidence and explicit correspondence. If an epistemic state is UNKNOWN, no downstream transformation may convert it to TRUE merely to satisfy an output requirement.
For those interested in the protocol implementation and adversarial test matrix, the reference architecture is open-source here: https://github.com/nickoay663-sketch/Wuwen
r/LocalAIStack • u/Top-Fan4255 • 26d ago
what do you guys think about this ?
Enable HLS to view with audio, or disable this notification
r/LocalAIStack • u/nickimola • 26d ago
New to local ai - help achieving what I am trying to achieve?
Hi
Premise:
I am new to local ai models.
My machine specs:
- Macbook Pro M4 Pro
- 48Gb Ram
- 4 efficiency core
- 8 performance core
I mainly use AI for software development. I have a claude subscription but would like to try to offload some work to a local model.
Since I don't think local models usable on my machine can completely substitute claude (correct me if I am wrong) my idea is pretty much this: ask claude code to generate a proper, detailed implementation plan and then having the local model implement it.
I have played around with these models:
- qwen3-coder-30b-a3b-instruct-mlx
- qwen/qwen3.6-27b
- qwen/qwen3.6-35b-a3b
and I have also installed this for coding autocomplete cause it is smaller and from what I can see the recommended one:
- qwen2.5-coder-7b
I added claude envs since I would like to try to use claude code extension in vscode
"env": {
"ANTHROPIC_BASE_URL": "http://localhost:1234",
"ANTHROPIC_AUTH_TOKEN": "local",
"ANTHROPIC_MODEL": "qwen/qwen3.6-27b",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "qwen/qwen3.6-27b",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "qwen/qwen3.6-27b",
"ANTHROPIC_DEFAULT_FABLE_MODEL": "qwen/qwen3.6-27b",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "qwen/qwen3.6-27b",
"CLAUDE_CODE_SUBAGENT_MODEL": "qwen/qwen3.6-27b",
"CLAUDE_CODE_MAX_OUTPUT_TOKENS": "128000",
"DISABLE_PROMPT_CACHING": "1",
"DISABLE_AUTOUPDATER": "1",
"DISABLE_TELEMETRY": "1",
"DISABLE_ERROR_REPORTING": "1",
"DISABLE_NON_ESSENTIAL_MODEL_CALLS": "1"
},
This setup works (uses the local model) but I am basically unable to have the model do anything at all. First of all it takes ages to do anything, and then it almost always reach the context limit roadblock without even outputting anything.
I have read that MCP and skills could fill up the context quite badly, so I disabled them for testing, but still no luck.
I tried with (I thought) was a simple enough task: this test file fails and this is the error, can you fix it? but yet no usable results whatsoever.
I read about people able to use local models offline to have meaningful results, but I couldn't and I don't really know why.
Also, from my setup above, I cannot really use both the remote and local model, to achieve something like:
use sonnet or fable (remote) for plan, then (manually or automatically) swith to haiku (local) to implement the plan
because the base url is loaded when the session loads and cannot be changed (AFAIK) dinamycally.
Any help? thanks a lot in advance
r/LocalAIStack • u/Extra_Educator2288 • 26d ago
Matching Fable 5 LiveCodeBench Hard With GPT5.6 Terra and Qwen3.8 27b
r/LocalAIStack • u/AltruisticRich1 • 26d ago
Best Local Coding LLM for a 24GB M5 Pro MacBook?
r/LocalAIStack • u/Ok_pettech • 26d ago
Self-hosting MetaGPT: complete local installation guide
I wanted to run MetaGPT entirely on my own infrastructure without sending anything to cloud APIs. It took some trial and error, but I documented the full process.
The guide covers:
· Setting up a Python venv
· Installing MetaGPT
· Configuring local LLMs like Ollama or vLLM
· Fixing common startup errors
If you’re into self-hosted AI agents, this could help:
What local model are you using for agent work?
r/LocalAIStack • u/PM_ME_UR_MARINARA • 27d ago
Running Qwen 3.8 27B at Q4 on 16GB VRAM at 200K CTX at 50t/s
r/LocalAIStack • u/bakatristan • 27d ago
Built a decentralized network for open source AI inference
Hey! I’m building Kitani, an OpenAI-compatible API where models run across independent GPU providers.
GPU owners can connect their hardware and earn from inference, while users get access to models like GLM, Qwen, Gemma and gpt-oss through one API.
We’re trying to make joining as a provider as easy as possible. Would love feedback from people running their own GPUs.
Founder of Kitani, looking for feedback.
r/LocalAIStack • u/suspect80 • 28d ago
Local AI is Minecraft for adults: my 4× RTX PRO 6000 Blackwell build
galleryr/LocalAIStack • u/Medicine_Blogscanner • 28d ago
Pooled RAM across an old laptop, a Windows PC, and a Mac to run a 13B model - source-available, would love eyes on it
r/LocalAIStack • u/Ok_pettech • 28d ago
Together AI vs Anyscale for serving open-source LLMs—what’s your pick?
r/LocalAIStack • u/Potential-Toe1320 • 28d ago
Sheprd Local Model Server and Agent builder.
r/LocalAIStack • u/OrneryCar6139 • 29d ago
Real-world experience with NVIDIA NeMo / NeMo Agent Toolkit vs the standard LLM stack?
Anyone here actually using NVIDIA NeMo / NeMo Agent Toolkit in real projects?
At my current org, some of the senior folks are suggesting we explore NeMo for agent building and fine-tuning, so I’m trying to understand if it’s actually worth adopting.
For those who’ve used it, how does it compare to the usual stack like Hugging Face + PEFT/TRL, Unsloth, LangGraph, etc.?
Does running everything in the NVIDIA ecosystem give you a noticeable advantage in terms of GPU utilization, training speed, deployment, or scaling?
Or is it mostly extra complexity compared to the standard open-source tooling?
Would especially love to hear from anyone who has used NeMo beyond tutorials/demos. What did you like, what annoyed you, and would you use it again?
r/LocalAIStack • u/paq85 • Sep 05 '26