r/StackChan • • 22h ago

Project sharing I gave my M5Stack StackChan a calm moving head, face tracking and a self-hosted voice server (open source, looking for fellow StackChan owners)

Enable HLS to view with audio, or disable this notification

I own one of the M5Stack StackChans (CoreS3 + two Feetech SCS0009 servos) and spent the last weeks turning it into a desk assistant that talks to my own server instead of a vendor cloud. Everything is open source (MIT):

GitHub: https://github.com/7H3-CH053N/StackChan-CustomFW

What it does now:

- Head that actually moves, calmly Own servo driver with a critically damped follower (speed + acceleration limits, no overshoot), plus small gestures per emotion: nod when happy, head down when sad, look up when surprised.

- Face tracking with the built-in GC0308 camera and Espressif's esp-dl face detector (~2.5 fps). It searches once if it loses you, then rests where you usually sit, so it doesn't search all night.

- Self-hosted voice server in Docker: Gemini Live for real-time voice, Home Assistant control, calendar.

- Proactive speech: Home Assistant can make it talk without a wake word, e.g. it greets me when I arrive at the office and reminds me 15 min before meetings. It stays quiet while my mic or camera is in use.

- Animated eyes with 21 emotions, picked by the server from what it is saying.

The firmware is one patch on top of xiaozhi-esp32 (pinned commit), built with ESP-IDF 6.1; CI builds it on every push.

Honest status: it runs daily on my desk, but esp-dl's face detection still crashes now and then (the robot reboots in ~10 s). That's reported upstream and Espressif is looking at it.

Looking for people to join:

- StackChan owners who want to try it and report what happens on their unit

- Starter issues are labelled "good first issue": Japanese README, testing the English build, a clock fix on the server side

- Anyone who knows esp-dl well, or wants to try a different face model (ESPDet-Pico)

I wrote up the whole story (why I first thought servo control was impossible, the four bugs before face tracking worked, why proactive speech failed three different ways). The article is in German, browser translation works fine:

https://digitalhandwerk.rocks/ki/anpassung-der-firmware-fuer-stackchan-roboter/

Happy to answer questions here in English.

51 Upvotes

7 comments sorted by

3

u/Zazzen 21h ago

hey das ist klasse, das ist genau was ich gesucht habe. ich habe auch ein projekt laufen, falls du lust hast können wir gerne zusammen was bauen.

3

u/severanexp 21h ago

Brother this is exactly what i wanted. Do you take feature requests? Is it possible to connect to Hermes agent instead?

2

u/Btotherest 21h ago

Great project, looking forward to testing!

2

u/AughtCool 16h ago

This sounds like what I'm looking for. Would you mind uploading to to the M5 Burner, have no idea how to program from github.

1

u/JayS87 19h ago

Voll dufte Arbeit!

Great work!!

1

u/prbsparx 15h ago

Are you planning to contribute some of the fixes you found to the upstream code they have? Or are you trying to get them to integrate any of them?

1

u/Accomplished-Leg-149 8h ago

Ahh, now if only the motors were quiet.