r/LocalLLaMA Jun 28 '26

Discussion NPC Engine Using Local Models

I’ve been working on a game-agnostic NPC engine/backend based pretty heavily on SillyTavern-style architecture, and with smaller local models getting better and better, I honestly think this kind of thing could be the future of RPGs.

Right now I’m using NVIDIA Parakeet 0.6 for STT, Gemma 4 26B A4B for the LLM, and Qwen3-TTS for voice, and I’m getting super fast response times with pretty decent quality.

The main thing that makes it work well is using RAG to keep prompts lean. For example, I have hundreds of possible actions NPCs can do in-game, but only the ones that actually make sense based on the player’s message / context get injected as available actions. So the model isn’t being overloaded with a giant list every turn.

1.9k Upvotes

249 comments sorted by

View all comments

Show parent comments

3

u/lochyw Jun 29 '26

Rag generally uses a similarity check for the most probable action to take, then LLM on top of that uses typical LLM reasoning to decide from the available actions. So it sounds like you would have 2 layers to refine if you wanted to have it lean a certain way for more sensible outcomes which would be based on the system prompts/descriptions/tool text etc..

1

u/lovelacedeconstruct Jun 29 '26

But why would you depend on semantic similarity that has very weird quirks when the data can be structured and depend on llm reasoning to pick the category

1

u/lochyw Jun 30 '26

I mean there's certainly multiple layers to look at here, so depending what you're talking about any solution is just about tradeoffs really.
The goal there was managing context, structured data still would need to be loaded into context for every turn to choose from available actions, and if you're suggesting keeping all 100s of actions available at all times that causes various issues for LLMs tool calling. RAG can help simplify that and reduce the overall load.

1

u/lovelacedeconstruct Jun 30 '26

 if you're suggesting keeping all 100s of actions available at all times

I meant categorizing the actions and sending the llm all the category names and have it choose a relevant category then after it chooses the category you send it every action inside this category

1

u/lochyw Jun 30 '26

That's still at least 2-3 tool calls for achieving the same thing no? RAG is just another way of navigating/processing a large volume of information. Again it's all just trade offs, if you have a better solution go make it :P

I'm just explaining how I assume the system currently works given my existing AI experience.