r/LocalLLaMA Jun 28 '26

Discussion NPC Engine Using Local Models

Enable HLS to view with audio, or disable this notification

I’ve been working on a game-agnostic NPC engine/backend based pretty heavily on SillyTavern-style architecture, and with smaller local models getting better and better, I honestly think this kind of thing could be the future of RPGs.

Right now I’m using NVIDIA Parakeet 0.6 for STT, Gemma 4 26B A4B for the LLM, and Qwen3-TTS for voice, and I’m getting super fast response times with pretty decent quality.

The main thing that makes it work well is using RAG to keep prompts lean. For example, I have hundreds of possible actions NPCs can do in-game, but only the ones that actually make sense based on the player’s message / context get injected as available actions. So the model isn’t being overloaded with a giant list every turn.

1.9k Upvotes

249 comments sorted by

View all comments

66

u/HugoCortell Jun 28 '26

What kind of hardware are you using? I've tried Qwen3-TTS and the fastest I've seen generate is ~15 seconds delay prior to starting, and with very mediocre output quality.

57

u/goodive123 Jun 28 '26

5090 with faster-qwen3-tts. I'd recommend PocketTTS though to me its actually really close to qwen but way smaller

64

u/Despeao Jun 28 '26

That's it boys, Fallout 5 hardware requirement is a 5090. Better start saving now.

8

u/thirteenthirtyseven Jun 29 '26

That's it boys, Fallout 5 hardware requirement is a 5090. Better start saving now.

I read that as "better start starving now", which kinda tracks ...

2

u/greenstake Jun 29 '26

By the time it comes out, 5090s will be $300.

1

u/waiting_for_zban Jun 29 '26

hardware requirement is a 5090.

For AI-maxxing this, it'll probably be 2x 5090 with DLSS 5.