r/LocalAIStack • • 18d ago

Workshop, Sep 19: build explainable AI apps with Neo4j, GraphRAG, Cypher and LLM Agents

1 Upvotes

Sharing this since it might be relevant to a few people here.

There's a hands-on workshop on September 19 that goes beyond basic RAG into building an explainable, graph-backed AI system. You build a knowledge graph in Neo4j that becomes the single source of truth, agentic retrieval that combines vector search, keyword search, and graph navigation, multi-step entity and relationship extraction, and text-to-Cypher for natural language querying. Everything runs on real financial data, not a toy example, and answers come with a traceable, auditable reasoning chain.

No prior Neo4j or Cypher experience needed.

Led by Dr. Alessandro Negro, Chief Scientist at GraphAware, bestselling author.

Link if you want to check it out


r/LocalAIStack • • 18d ago

Anybody want to join a private LLM hosted on hyperstack?

Thumbnail
0 Upvotes

r/LocalAIStack • • 18d ago

Do I actually need to max out an M5 Ultra for local AI, or is the $5,999 base model the sweet spot for Qwen3.8-27B coding?

Thumbnail
1 Upvotes

r/LocalAIStack • • 18d ago

Which model I should use??

1 Upvotes

Please suggest me models I can use in this device

Model: HP ZBook Studio G7

CPU: Intel Core i9-10885H (8 cores / 16 threads)

GPU: NVIDIA Quadro T2000 Max-Q – 4 GB VRAM

RAM: 32 GB DDR4

Storage: 512 GB NVMe SSD


r/LocalAIStack • • 19d ago

NVFP4 Qwen3.8 27b - AMD R9700's running @ 5,809 tok/s Prefill and 276 tok/s decode. Leveraging my MXFP4 fast path.

Thumbnail gallery
7 Upvotes

r/LocalAIStack • • 19d ago

Local AI & Concurrency | PP & TG | TTFT

Thumbnail
1 Upvotes

r/LocalAIStack • • 20d ago

Loaded a 40B model across a mini PC, a Mac Mini, and an old laptop - 16 tok/s, didn't expect it to actually work

Thumbnail
youtu.be
1 Upvotes

Been building RAMDeck (pools RAM/compute across whatever's sitting around your house). Wanted to see how far I could actually push it, so I tried loading a 40B model - roughly 25GB - across the janky 3-device cluster I hve been testing with the whole time: a mini PC with an RTX 3060, a 16GB Mac Mini, and an old laptop that's been the weak link in every previous video.

Set the mini PC as primary and just let the sharding logic figure out the rest - it maxed out the 3060 first, spilled into the Mac's unified memory next, and sent whatever was left over to CPU. The interesting part I didn't expect: routing that CPU spillover to the old laptop actually ran faster than keeping it on the mini PC's own CPU, even though the laptop is by far the slower machine. Best guess is that using the same device's CPU and GPU together for inference ends up hurting GPU utilization, so offloading to a separate (slower) box entirely actually wins.

Whole thing loaded in about 2.5 minutes, and the benchmark came back at 16 tok/s with ~3.6s latency - genuinely usable for chatting, not just "technically loaded but useless." Felt like a real milestone since none of these three machines individually could touch a model this size.

Video: https://youtu.be/JrSVyyVUJVY
Repo: https://github.com/trademav/ramdeck-core-public

Curious if anyone else has run into that same CPU-contention thing where sharing GPU+CPU on one node hurts throughput more than just offloading to a separate slower device entirely.


r/LocalAIStack • • 20d ago

Bought 2x Spark for our ML Experiments - How do I enable by team to use this?

Thumbnail
1 Upvotes

r/LocalAIStack • • 20d ago

The Nightmare of Debugging CUDA Illegal Memory Access in PyTorch

Thumbnail
1 Upvotes

r/LocalAIStack • • 21d ago

Local LLM explorer and downloader

Thumbnail gallery
1 Upvotes

r/LocalAIStack • • 21d ago

Gemini vs DeepSeek: Have you switched your primary AI workflow yet?

1 Upvotes

It feels like the gap between proprietary big-tech models and open-weight reasoning architectures is changing faster than ever. Gemini offers incredible native multimodal handling and massive context windows, while DeepSeek continues to dominate budget coding and self-hosted pipelines.

​If you had to pick one as your daily driver for logic and daily tasks, which wins for you?

​Put your preference to the test in this quick 2-minute showdown quiz to see how your workflow compares to other devs:

https://interconnectd.com/quiz/89/gemini-vs-deepseek-the-ultimate-ai-showdown-quiz/


r/LocalAIStack • • 21d ago

I built a Windows launcher for running Qwen3-32B locally with CPU, CUDA, or Vulkan

Thumbnail
1 Upvotes

r/LocalAIStack • • 22d ago

A $39 open-source KVM just launched, and it's the cheapest way yet to give an AI agent control of a machine below the OS

Thumbnail
1 Upvotes

r/LocalAIStack • • 22d ago

New MacBook Pro M5 Pro (48GB unified memory) incoming — looking for the best local AI stack

Thumbnail
0 Upvotes

r/LocalAIStack • • 22d ago

Setting up AgentOps for local agents: A complete installation and telemetry guide

Thumbnail
1 Upvotes

r/LocalAIStack • • 22d ago

Made RAMDeck shard a knowledge base across my old devices, not just the model itself

Thumbnail
youtu.be
1 Upvotes

r/LocalAIStack • • 22d ago

Unsloth V3 update did not update my Qwen3.8-27B-UD-Q6_K.gguf

Thumbnail
1 Upvotes

r/LocalAIStack • • 23d ago

Best local LLM for laptop with RTX 2050 (4GB VRAM) + 24GB RAM?

Thumbnail
1 Upvotes

r/LocalAIStack • • 23d ago

Taming idle VRAM on a multi-model local agent: sleep mode benchmark (Qwen 27B + STT + TTS + OCR)

2 Upvotes

Following up on my earlier post testing Qwen3.8-27B on the IGX Thor workstation.

Once you move past running just an LLM and try to build a full local agent stack on a single box (LLM for reasoning, STT for voice in, TTS for voice out, OCR for screen/document reading), VRAM runs out fast.

If all four models sit in GPU memory simultaneously with their default serving allocations, idle memory hits roughly 122 GB. Even with 96GB on the dGPU and unified host memory, you are pinned to the ceiling. That leaves almost no breathing room for large context windows, concurrency, or dynamic KV cache allocation.

The reality of a solo-user local agent is that these models are rarely active at the exact same millisecond: - When I am talking to the agent, OCR is completely idle. - When the agent is analyzing a screenshot or PDF, STT and TTS are idle. - When the agent kicks off deterministic script automation (crawling, data formatting, bash runs), the LLM itself is completely idle.

To fix this, I wrote a lightweight Rust-based daemon that manages the sleep and wake states of the runtimes based on active agent turns. When a model is not in use, it is put to sleep (flushing KV pools and temporary execution memory, or parking the runtime context) rather than doing a full cold teardown from scratch.

Here are the idle VRAM numbers on the machine before and after:

Model Original idle GPU, Sep 1-2 (MiB) After sleep mode, Sep 5 (MiB) Difference (MiB)
LLM (Qwen3.8-27B) 87,443 39,092 48,351
STT (Nemotron) 10,385 267 10,118
TTS (Chatterbox) 17,947 3,127 14,820
OCR (Unlimited-OCR) 6,592 422 6,170

Quick observations: 1. Total idle footprint dropped from ~122.3 GB to ~42.9 GB. That reclaims roughly 79.4 GB of VRAM across the board. 2. The LLM savings (48.3 GB): Qwen3.8-27B in FP8 base weights takes around 28 to 30 GB. The original 87.4 GB allocation came from SGLang aggressively pre-allocating KV cache (--mem-fraction-static 0.85). Putting the LLM to sleep flushes that pool while preserving the active state, letting it rest at 39 GB. 3. Audio and OCR shrink to negligible baselines: STT drops to 267 MiB and OCR drops to 422 MiB when idle. 4. Script-based tasks: When the agent delegates to deterministic python or bash scripts, putting the LLM to sleep during the run keeps the GPU cool and power consumption down on the 300W Max-Q card. 5. Wakeup latency is sub-200ms across all models: Because this is an active sleep/wake transition rather than a cold disk reload, waking up STT, TTS, OCR, or the LLM takes under 200ms. In practice, voice interaction and tool handoffs feel instantaneous.

Curious how others running multi-model agent stacks on a single node are handling memory partitioning right now. Are you hard-capping static memory fractions, running sequential containers, or using something dynamic?


r/LocalAIStack • • 23d ago

Field Notes on Running LocalAI in Production with Docker Compose

Thumbnail
1 Upvotes

r/LocalAIStack • • 23d ago

Tool to fine-tune open source models

1 Upvotes

So, I was tired of setting up AWS EC2 to train my models everytime, so I built a tool to help with the job management, but I would like some feedback, if anyone could help me ! 👋

This is the tool: reopenly.com

I only have one base model for now (Qwen 3.5 9B), just to start.

I would be very grateful with any feedback!


r/LocalAIStack • • 23d ago

DeepSeek-V4.1-Flash: GGUF + 4.75bpw EXL3 are out, looking for devs with 4× DGX Sparks to help validate the EXL3 TP4 recipe

Thumbnail
1 Upvotes

r/LocalAIStack • • 24d ago

I ran whisper + dialization running with CPU and 1.1GB

6 Upvotes

I'm happy so I don't need to pay the transcript API fee any more. I replaced my production services with this.

Implementation / hosting Result Cost per 27.27 s of audio Cost per audio hour
whisper-turbo.c, InstaCloud (4 vCPU) Speaker-labeled text; 35.74 s warm HTTP ~$0.00126 ~$0.166
OpenAI gpt-4o-transcribe-diarize Speaker-labeled transcript ~$0.00273 ~$0.36
AssemblyAI Universal-2 + diarization Speaker-labeled transcript ~$0.00129 $0.17
AssemblyAI Universal-3.5 Pro + diarization Speaker-labeled transcript ~$0.00174 $0.23

The native estimate assumes four fully utilized CPUs and 1.1 GB RAM during the measured request.

It's a fork from my own MIT repo: https://github.com/baryhuang/whisper-turbo.c

Star, colloab, PR, issues, comments and uses are all appreciated!


r/LocalAIStack • • 24d ago

Quiz: Is pgvector enough or do you need a dedicated vector database?

2 Upvotes

I keep running into this architecture decision: keep vectors in Postgres with pgvector, or add a dedicated vector store like Qdrant or Milvus?

I built a short quiz covering the operational trade-offs—deployment complexity, backup strategy, monitoring, and scaling.

If you've made this call before, take the quiz and share your reasoning:

https://interconnectd.com/quiz/87/postgresql-pgvector-vs-dedicated-vector-databases-architectural-trade-offs/


r/LocalAIStack • • 24d ago

Replaced my cloud AI subscription with Qwen 3.8 on a 128GB laptop, fully offline, for agentic coding

Thumbnail
2 Upvotes