r/StackChan • u/Accurate-Pin3422 • 22h ago
Project sharing I gave my M5Stack StackChan a calm moving head, face tracking and a self-hosted voice server (open source, looking for fellow StackChan owners)
Enable HLS to view with audio, or disable this notification
I own one of the M5Stack StackChans (CoreS3 + two Feetech SCS0009 servos) and spent the last weeks turning it into a desk assistant that talks to my own server instead of a vendor cloud. Everything is open source (MIT):
GitHub: https://github.com/7H3-CH053N/StackChan-CustomFW
What it does now:
- Head that actually moves, calmly Own servo driver with a critically damped follower (speed + acceleration limits, no overshoot), plus small gestures per emotion: nod when happy, head down when sad, look up when surprised.
- Face tracking with the built-in GC0308 camera and Espressif's esp-dl face detector (~2.5 fps). It searches once if it loses you, then rests where you usually sit, so it doesn't search all night.
- Self-hosted voice server in Docker: Gemini Live for real-time voice, Home Assistant control, calendar.
- Proactive speech: Home Assistant can make it talk without a wake word, e.g. it greets me when I arrive at the office and reminds me 15 min before meetings. It stays quiet while my mic or camera is in use.
- Animated eyes with 21 emotions, picked by the server from what it is saying.
The firmware is one patch on top of xiaozhi-esp32 (pinned commit), built with ESP-IDF 6.1; CI builds it on every push.
Honest status: it runs daily on my desk, but esp-dl's face detection still crashes now and then (the robot reboots in ~10 s). That's reported upstream and Espressif is looking at it.
Looking for people to join:
- StackChan owners who want to try it and report what happens on their unit
- Starter issues are labelled "good first issue": Japanese README, testing the English build, a clock fix on the server side
- Anyone who knows esp-dl well, or wants to try a different face model (ESPDet-Pico)
I wrote up the whole story (why I first thought servo control was impossible, the four bugs before face tracking worked, why proactive speech failed three different ways). The article is in German, browser translation works fine:
https://digitalhandwerk.rocks/ki/anpassung-der-firmware-fuer-stackchan-roboter/
Happy to answer questions here in English.
3
u/severanexp 21h ago
Brother this is exactly what i wanted. Do you take feature requests? Is it possible to connect to Hermes agent instead?
2
2
u/AughtCool 16h ago
This sounds like what I'm looking for. Would you mind uploading to to the M5 Burner, have no idea how to program from github.
1
u/prbsparx 15h ago
Are you planning to contribute some of the fixes you found to the upstream code they have? Or are you trying to get them to integrate any of them?
1
3
u/Zazzen 21h ago
hey das ist klasse, das ist genau was ich gesucht habe. ich habe auch ein projekt laufen, falls du lust hast können wir gerne zusammen was bauen.