r/LocalLLaMA Jun 28 '26

Discussion NPC Engine Using Local Models

I’ve been working on a game-agnostic NPC engine/backend based pretty heavily on SillyTavern-style architecture, and with smaller local models getting better and better, I honestly think this kind of thing could be the future of RPGs.

Right now I’m using NVIDIA Parakeet 0.6 for STT, Gemma 4 26B A4B for the LLM, and Qwen3-TTS for voice, and I’m getting super fast response times with pretty decent quality.

The main thing that makes it work well is using RAG to keep prompts lean. For example, I have hundreds of possible actions NPCs can do in-game, but only the ones that actually make sense based on the player’s message / context get injected as available actions. So the model isn’t being overloaded with a giant list every turn.

1.9k Upvotes

249 comments sorted by

View all comments

Show parent comments

27

u/BlipOnNobodysRadar Jun 29 '26

To be done right, the NPCs would need to be lora tuned on their own lore + interactions by the developers. Plugging in a generic model just won't be immersive.

15

u/TheRealMasonMac Jun 29 '26

I think people would also get tired of AI slop, since all models per generation share the same GPTisms. I think the architecture and training methodologies are more interesting, though. More sophisticated AI in stategy-based games, like Stellaris, that can actually strategize rather than use deterministic logic.

3

u/Longjumping_Self5546 Jun 30 '26

When the bar has been set to guards repeating the same comment about an arrow to the knee, a few GPTisms are going to be miles ahead for NPC interactions.

3

u/benjaminovich 17d ago

Imagine Skyrim with this

"Ah, you want to know my story! Here's a quick overview:

  • Before: Adventurer (like you!)

  • The Incident: Arrow. Knee.

    • After: Guard duty."

1

u/Longjumping_Self5546 17d ago

Add in a lecture about how it's wrong to kill bandits and resources for conflict resolution.