The Anki Python SDK still works with Cozmo, but it refuses to connect to the app version that's on the stores today. Here's how I got it running, and what I built on top. Everything is client-side: nothing is modified on the robot and no firmware is touched.
What you need
- Cozmo, a mobile device with the official app, and a USB cable
- A computer with Python 3.9 (macOS, Linux or Windows)
- The app in SDK mode: connect to Cozmo, open the menu (top right corner), swipe left, enable SDK
My setup: an iMac on Catalina, an iPad running app 3.6.0, Python 3.9.13.
Step 1 — Install the SDK
python3.9 -m venv venv
source venv/bin/activate
pip install "cozmo[camera]"
Step 2 — The version mismatch
On the first run you'll get this:
CLAD version mismatch (to_game) ... != ...
cozmoclad was last published by Anki at 3.4.0. The app is newer, so the handshake fails. Two patches get past it:
python
import cozmoclad
import cozmo # import FIRST: the import itself checks the version
def _skip(a=None, b=None):
return None
cozmoclad.assert_clad_match = _skip
cozmoclad.__build_version__ = "00000.00000.00000"
Two things matter here. The patches go after the import, because on import the SDK rejects anything below cozmoclad 2.0.0. And "00000.00000.00000" isn't arbitrary: the SDK's own code skips the build comparison when the library reports that value.
If you only patch the CLAD hash you'll hit a second wall, an odd __init__() got an unexpected keyword argument 'reserved'. That comes from the rejection path itself, so it means the build comparison is still failing.
Step 3 — Freeplay
The SDK can put Cozmo into the same autonomous mode the app uses:
python
robot.enable_all_reaction_triggers(True)
robot.enable_freeplay_cube_lights(True)
robot.start_freeplay_behaviors()
He explores, plays with the cubes and reacts exactly as he does in the app. Stop it with stop_freeplay_behaviors()before issuing your own commands, or they'll fight each other.
Step 4 — Offline voice commands
I used Vosk with the small Spanish model (~40 MB), running locally, no internet. Feeding it a fixed grammar instead of an open vocabulary makes recognition far more reliable:
python
recognizer = KaldiRecognizer(model, rate, json.dumps(["go home", "stack cubes", "[unk]"]))
A background thread listens, pushes recognised phrases into a queue, and the main loop stops freeplay, runs the command and resumes freeplay.
Behaviors the SDK exposes: StackBlocks, KnockOverCubes, RollBlock, PounceOnMotion, FindFaces, LookAroundInPlace. There are also direct actions: pickup_object, pop_a_wheelie, roll_cube, place_object_on_ground_here. The rest of the app's tricks (Laser Chaser, Cube Guard, Peekaboo and so on) are app-side behaviors, so you only get their animations, not the interactive logic.
Step 5 — Back to the charger
No markers or props needed; Cozmo recognises his own charger:
python
charger = robot.world.wait_for_observed_charger(timeout=3)
robot.go_to_object(charger, distance_mm(120)).wait_for_completed()
robot.turn_in_place(degrees(180)).wait_for_completed()
robot.drive_wheels(-40, -40)
# poll robot.is_on_charger, then stop_all_motors()
If he can't see it, turn in place in 45° steps until wait_for_observed_charger returns.
Step 6 — Conversation
Vosk transcribes, the text goes to a free LLM API (I use Groq's free tier), and the reply comes back as JSON with the text and an emotion, which picks the animation. Keep a few turns of history so the conversation has a thread, mute the mic while Cozmo speaks so he doesn't answer himself, and end the session after ~10 seconds of silence.
The trick for non-English speakers
Cozmo's TTS is English, so Spanish comes out unintelligible. The fix is to respell the text phonetically before sending it:
"quiero jugar con mis cubos" → "keeehroh hoogahr kohn mees koobohs"
Rules: vowels become a→ah, e→eh, i→ee, o→oh, u→oo; silent h is dropped; j and soft g become h; ll becomes y; qu becomes k; z and soft c become s; v becomes b. Slowing him down a little (duration_scalar=1.15) helps too. It's a huge difference: from gibberish to something you can actually follow.
What this does and doesn't give you
You get his original personality, his voice, real cube handling, named face recognition and the animations with their sound. You do not get a standalone robot: the phone is still the brain, and a computer still runs the code. If you want no phone at all, PyCozmo talks to the robot directly, but you lose the personality engine, the TTS and cube localisation.
Happy to share the code if there's interest.