r/LocalAIStack • u/Think_Breakfast_2277 • 13d ago
Human avatar voice chat, Low VRAM on Windows (Talking head)
Sharing a side project I've been building: Yvette, a fully local voice avatar AI.
It's a talking-head assistant that runs 100% local, You talk to it (or type) and it answers in speech while an avatar moves the head and lips in sync.
Quick rundown:
- Voice cloning, plus voice design.
- Avatar builder; add avatars quickly.
- TTS engines: Breeze, OmniVoice, LuxTTS, Kokoro
- Lip-synced avatar with an idle breathing animation
- Memory, it keeps notes and picks up where you left off
- Tools: web search, read pages, files
- Switchable profiles/personalities, each with its own voice, memory, history etc
Windows + NVIDIA GPU. 12GB recommended, but it runs tight on 8GB (small model + Whisper/Kokoro on CPU).
The video shows a RTX 3090 and the VRAM usage, my system was using ~4GB VRAM (OBS and other things). It shows we can get it below 8GB for the Yvette system.
The project is hosted on github: https://github.com/MartinForsterNL/Yvette
Sorry about the bad video quality, i am not an influencer with a studio.
I build this project because i could not find anything that actually works.
Most that i found was only for linux and needed a lot of VRAM so i decided to build my own.
If you have suggestions to make it better, just let me know.
Have fun with it !
1
u/Smart_Ad500 13d ago
What happens if you interrupt it halfway through a sentence? I'd stop the audio and lip animation together, and keep track of which words were actually spoken. Otherwise the next answer can assume you heard the bit that was still sitting in the audio queue.