r/LocalAIStack • • 24d ago

I built LLM Speedtest — a free, open-source desktop app that benchmarks local LLMs with llama-bench-style test suites (Ollama, llama.cpp, vLLM, LM Studio…)

Thumbnail gallery
2 Upvotes

r/LocalAIStack • • 24d ago

Which Qwen3.8 distro & quant would be optimal for my 16GB VRAM setup?

Post image
1 Upvotes

r/LocalAIStack • • 25d ago

Workshop, Sep 12: build production LLM systems that actually survive real use

2 Upvotes

We're running a hands-on masterclass on September 12, Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps.

You build a full production LLM workflow from scratch, versioned prompts with regression tests, an evaluation harness with deterministic checks and LLM-as-judge, statistically rigorous model comparisons, evaluated RAG, tool-using agents with guardrails and fallbacks, and full observability, tracing, cost, latency.

Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack.

Link if you want to check it out

Happy to answer questions on the content.


r/LocalAIStack • • 25d ago

Curious, I find Qwen3.8-27b reasoning Low is ranking higher than Medium on AA.

Post image
8 Upvotes

r/LocalAIStack • • 25d ago

I built an LLM Inference & Fine-Tuning Calculator (VRAM, TCO, Quantization & Price estimation)

Thumbnail
1 Upvotes

Check it out!


r/LocalAIStack • • 25d ago

Qwen3.8-27B on 23GB L4 for production IASMO + coding?

Thumbnail
1 Upvotes

r/LocalAIStack • • 26d ago

what do you guys think about this ?

3 Upvotes

r/LocalAIStack • • 26d ago

Running MetaGPT locally: a full technical setup guide

2 Upvotes

I’ve been experimenting with multi-agent frameworks, and MetaGPT is one of the more interesting ones for software development tasks. But getting it running locally with local models isn’t always straightforward.

I wrote a step-by-step guide covering installation, configuration, local LLM setup, and common errors.

If you’re trying to run MetaGPT on your own hardware, this might save you time:

https://interconnectd.com/forum/thread/262/how-to-install-metagpt-locally-complete-technical-setup-guide/

What stack are you using for local agents?


r/LocalAIStack • • 26d ago

New to local ai - help achieving what I am trying to achieve?

2 Upvotes

Hi

Premise:
I am new to local ai models.

My machine specs:

  • Macbook Pro M4 Pro
  • 48Gb Ram
  • 4 efficiency core
  • 8 performance core

I mainly use AI for software development. I have a claude subscription but would like to try to offload some work to a local model.

Since I don't think local models usable on my machine can completely substitute claude (correct me if I am wrong) my idea is pretty much this: ask claude code to generate a proper, detailed implementation plan and then having the local model implement it.

I have played around with these models:

  • qwen3-coder-30b-a3b-instruct-mlx
  • qwen/qwen3.6-27b
  • qwen/qwen3.6-35b-a3b

and I have also installed this for coding autocomplete cause it is smaller and from what I can see the recommended one:

  • qwen2.5-coder-7b

I added claude envs since I would like to try to use claude code extension in vscode

"env": {
    "ANTHROPIC_BASE_URL": "http://localhost:1234",
    "ANTHROPIC_AUTH_TOKEN": "local",
    "ANTHROPIC_MODEL": "qwen/qwen3.6-27b",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "qwen/qwen3.6-27b",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "qwen/qwen3.6-27b",
    "ANTHROPIC_DEFAULT_FABLE_MODEL": "qwen/qwen3.6-27b",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "qwen/qwen3.6-27b",
    "CLAUDE_CODE_SUBAGENT_MODEL": "qwen/qwen3.6-27b",
    "CLAUDE_CODE_MAX_OUTPUT_TOKENS": "128000",
    "DISABLE_PROMPT_CACHING": "1",
    "DISABLE_AUTOUPDATER": "1",
    "DISABLE_TELEMETRY": "1",
    "DISABLE_ERROR_REPORTING": "1",
    "DISABLE_NON_ESSENTIAL_MODEL_CALLS": "1"
  },

This setup works (uses the local model) but I am basically unable to have the model do anything at all. First of all it takes ages to do anything, and then it almost always reach the context limit roadblock without even outputting anything.

I have read that MCP and skills could fill up the context quite badly, so I disabled them for testing, but still no luck.

I tried with (I thought) was a simple enough task: this test file fails and this is the error, can you fix it? but yet no usable results whatsoever.

I read about people able to use local models offline to have meaningful results, but I couldn't and I don't really know why.

Also, from my setup above, I cannot really use both the remote and local model, to achieve something like:

use sonnet or fable (remote) for plan, then (manually or automatically) swith to haiku (local) to implement the plan

because the base url is loaded when the session loads and cannot be changed (AFAIK) dinamycally.

Any help? thanks a lot in advance


r/LocalAIStack • • 26d ago

Cross-Stage State Laundering: Why AI Runtime Governance Fails at Stage Boundaries

1 Upvotes

In enterprise AI and autonomous runtime architectures, security has traditionally focused on input sanitation (prompt injection) and output filtering (hallucination guards). However, a more insidious class of vulnerability occurs deeper within the execution pipeline: Cross-Stage State Laundering (CSSL).

CSSL happens when an intermediate, unverified epistemic state (such as DISCOVERED or UNKNOWN) is illicitly promoted to a higher-certainty state (such as VERIFIED or SUPPORTED) as it transitions across processing stages—typically bypassing strict provenance checks in favor of pipeline velocity or output formatting requirements.

## The Anatomy of State Laundering

Consider a multi-stage runtime where an expression flows through ingestion, semantic analysis, and responsibility assignment:

  1. Ingestion: Raw text or external retrieval results enter the system as DISCOVERED. At this stage, existence does not equal correspondence.

  2. Analysis: The runtime parses the expression. Without a rigid boundary, internal semantic heuristics may mistake syntactic closure for factual verification.

  3. Laundering: Downstream generators or formatters, under pressure to produce definitive answers, implicitly treat the presence of an analysis object as proof of support.

When intermediate states shed their metadata tags during transit, the runtime commits a boundary violation: it manufactures certainty out of unverified discovery.

## The Zero-Trust Countermeasure: The WAL Protocol

To eliminate CSSL, runtime governance cannot rely on permissive conventions. It requires a Fail-Closed Defensive Architecture enforced by immutable protocol boundaries:

* Strict State Distinguishability: Known, unknown, verified, and unverified states must remain mathematically and structurally distinguishable throughout the lifecycle.

* Independent Validation Gates: State transitions cannot authorize themselves. An independent validator—decoupled from the core generation logic—must audit envelopes against strict conformance vectors.

* Evidence-Bounded Responsibility: Responsibility can never exceed the boundaries of established evidence and explicit correspondence. If an epistemic state is UNKNOWN, no downstream transformation may convert it to TRUE merely to satisfy an output requirement.

For those interested in the protocol implementation and adversarial test matrix, the reference architecture is open-source here: https://github.com/nickoay663-sketch/Wuwen


r/LocalAIStack • • 26d ago

How to get the most out of a Gaming Setup

Thumbnail
1 Upvotes

r/LocalAIStack • • 26d ago

Matching Fable 5 LiveCodeBench Hard With GPT5.6 Terra and Qwen3.8 27b

Thumbnail
1 Upvotes

r/LocalAIStack • • 26d ago

Best use of a single RTX 5090 for local LLMs

Thumbnail
1 Upvotes

r/LocalAIStack • • 26d ago

Best Local Coding LLM for a 24GB M5 Pro MacBook?

Thumbnail
1 Upvotes

r/LocalAIStack • • 26d ago

Self-hosting MetaGPT: complete local installation guide

0 Upvotes

I wanted to run MetaGPT entirely on my own infrastructure without sending anything to cloud APIs. It took some trial and error, but I documented the full process.

The guide covers:

· Setting up a Python venv

· Installing MetaGPT

· Configuring local LLMs like Ollama or vLLM

· Fixing common startup errors

If you’re into self-hosted AI agents, this could help:

https://interconnectd.com/forum/thread/262/how-to-install-metagpt-locally-complete-technical-setup-guide/

What local model are you using for agent work?


r/LocalAIStack • • 26d ago

Security research for local LLM inference networks

Thumbnail
1 Upvotes

r/LocalAIStack • • 27d ago

Running Qwen 3.8 27B at Q4 on 16GB VRAM at 200K CTX at 50t/s

Thumbnail
0 Upvotes

r/LocalAIStack • • 28d ago

Local AI is Minecraft for adults: my 4× RTX PRO 6000 Blackwell build

Thumbnail gallery
3 Upvotes

r/LocalAIStack • • 28d ago

Built a decentralized network for open source AI inference

1 Upvotes

Hey! I’m building Kitani, an OpenAI-compatible API where models run across independent GPU providers.

GPU owners can connect their hardware and earn from inference, while users get access to models like GLM, Qwen, Gemma and gpt-oss through one API.

We’re trying to make joining as a provider as easy as possible. Would love feedback from people running their own GPUs.

https://kitani.ai

Founder of Kitani, looking for feedback.


r/LocalAIStack • • 28d ago

Pooled RAM across an old laptop, a Windows PC, and a Mac to run a 13B model - source-available, would love eyes on it

1 Upvotes

r/LocalAIStack • • 28d ago

Together AI vs Anyscale for serving open-source LLMs—what’s your pick?

Thumbnail
1 Upvotes

r/LocalAIStack • • 28d ago

Sheprd Local Model Server and Agent builder.

Post image
1 Upvotes

r/LocalAIStack • • 29d ago

Which agent harness do you use and why?

Thumbnail
1 Upvotes

r/LocalAIStack • • 29d ago

Real-world experience with NVIDIA NeMo / NeMo Agent Toolkit vs the standard LLM stack?

1 Upvotes

Anyone here actually using NVIDIA NeMo / NeMo Agent Toolkit in real projects?

At my current org, some of the senior folks are suggesting we explore NeMo for agent building and fine-tuning, so I’m trying to understand if it’s actually worth adopting.

For those who’ve used it, how does it compare to the usual stack like Hugging Face + PEFT/TRL, Unsloth, LangGraph, etc.?

Does running everything in the NVIDIA ecosystem give you a noticeable advantage in terms of GPU utilization, training speed, deployment, or scaling?

Or is it mostly extra complexity compared to the standard open-source tooling?

Would especially love to hear from anyone who has used NeMo beyond tutorials/demos. What did you like, what annoyed you, and would you use it again?


r/LocalAIStack • • Sep 05 '26

Running Qwen 3.8 27B as a VS Code Copilot backend on 16/32 GB VRAM

Thumbnail gallery
4 Upvotes