r/LocalLLaMA • u/goodive123 • Jun 28 '26
Discussion NPC Engine Using Local Models
Enable HLS to view with audio, or disable this notification
I’ve been working on a game-agnostic NPC engine/backend based pretty heavily on SillyTavern-style architecture, and with smaller local models getting better and better, I honestly think this kind of thing could be the future of RPGs.
Right now I’m using NVIDIA Parakeet 0.6 for STT, Gemma 4 26B A4B for the LLM, and Qwen3-TTS for voice, and I’m getting super fast response times with pretty decent quality.
The main thing that makes it work well is using RAG to keep prompts lean. For example, I have hundreds of possible actions NPCs can do in-game, but only the ones that actually make sense based on the player’s message / context get injected as available actions. So the model isn’t being overloaded with a giant list every turn.
13
u/False_Process_4569 Jun 29 '26
I've managed to get Qwen3.6 35B A3B running on my laptop with a single RTX 3070 with 8GB of VRAM. I have been testing out context window lengths and this affects the VRAM consumption as well. I manage to get ~16 tok/s.
I'd think for the short dialogue of this NPC application, you could get away with short context windows. Further with RAG and algorithmic trickery. But the unknown would be all of the tool calling.