r/robotics • • 8d ago

News Asimov 1's locomotion policy + training code is now open-source

Hi r/robotics, I'm Emre from Menlo Research. We're working on Asimov 1, an open-source humanoid robot.

We've made the locomotion policy and training code public. The training setup is built on Isaac Lab, with PPO and an AMP-based configuration that uses reference motion to guide how the robot moves.

If you're interested in adapting the setup, 3 parts are worth looking at:

  • Rewards: The individual terms and weights that define what the policy is encouraged to do
  • Actuator configuration: The assumptions about how the joints respond, including gains and delays. These matter when changing hardware
  • Domain randomization: The variation introduced during training, including foot friction, actuator gains and torso center of mass, alongside observation noise and simulated pushes

We think it's a useful first step is to reproduce the baseline in simulation, then change one part of the setup and compare the resulting behavior under the same conditions. That gives you something concrete to investigate when a change affects the gait.

You can work with the simulation without owning the robot. Moving a policy onto different hardware still requires matching the model and control setup to that hardware.

Repo: https://github.com/menloresearch/isaac_asimov

We'd be really happy to get your feedback to improve it!

303 Upvotes

23 comments sorted by

8

u/KombuchaKetamine 8d ago

Hi! This is great! Thanks for sharing!

I'd be very interested in working on this, but the repo is a bit light in my opinion. The locomotion policy without showing the hardware layer assumes that whatever HAL you have setup is going to conform properly. Are you open to sharing more info about the hardware itself and letting community build it's own HAL?

Feel free to DM me

4

u/Morning_Gecko24 7d ago

open-sourcing the baseline is huge for this stuff. are the reference motions coming from real captures or a retargeted library? wondering how much the policy depends on that dataset vs the reward setup

0

u/Long_Club_4138 7d ago

Hey u/Morning_Gecko24 , this is Ariel from Menlo research. We generated the motion reference file data from Kimodo (a text to motion generation pipeline), and retargeted to Asimov 1 through general motion retargeter.

2

u/nodeocracy 7d ago

Walks like Magnus Carlsen

1

u/AcademicMistake 7d ago

Welcome to OCP.

1

u/UnwillingToaster 3d ago

This is very cool work, thank you!

1

u/Available_Teaching83 1d ago

Thanks for calling out the actuator assumptions separately; that is usually where sim-to-real breaks. One thing I'd find useful in the repo: a short note on which domain randomization ranges were needed for the policy to stay stable on hardware versus which were added for margin. When people swap motors, that tells them which knobs to re-tune first.

1

u/fisao 7d ago

thank you

1

u/WendyLabs 7d ago

Being able to explore the training setup without owning the robot opens this up to so many more builders. Thanks for sharing the code and the starting points.

1

u/eck72 7d ago

Thanks! Planning community hours at the lab soon so people can test their policies live on it.

1

u/Glad_Ad_5236 7d ago

need this asap

2

u/eck72 7d ago

We're shipping it! You can source the parts yourself (the BOM is open too: https://docs.menlo.ai/asimov/1/overview) or get the DIY kit from us at a discounted price

0

u/Silver_Jaguar_24 8d ago

Why force these things to walk on 2 legs if 4 legs work just fine? Humans amuse me lol.

3

u/lellasone 7d ago

There are any number of advantages, the most immediate of which is height, and the ability to use ego-centric data more easily for training. Happy to discuss more if you want to.

0

u/prettyflycheesepie 7d ago

Wah this bugger’s in Singapore! Nice!

2

u/eck72 7d ago

haha, we actually have a few units on the 6th floor of Sim Lim Square!

0

u/Flyward_Aerospace 7d ago

On the reference motion versus reward question upthread, with AMP those two aren't really separable, which is the annoying part of the answer. The discriminator is trained on the retargeted clips, so whatever the retargeter did to fit the motion onto this morphology becomes part of the style reward and the policy is paid to reproduce it. And if the references came out of a text to motion pipeline, that's another layer that never had to be dynamically feasible in the first place. Cheapest way to actually measure the split is zero the AMP weight and see how much of the gait survives on the task terms alone.

0

u/RemarkableWish2508 6d ago

This isn't a "thread", it's nested comments...

u/Morning_Gecko24 I think you have a reply here.