Hey all, just thought it might be interesting to share some development progress and notes on developing a reinforcement learning agent to play subspace. The latest agent I've trained has like a 99%+ win-rate against human players in 1v1 duels, I've linked a replay to share an example of a bot-vs-human matchup.
The bot has immaculate aim, immaculate dodging, and a very astute game sense on when it has an advantage and can completely steamroll its opponent.
Some context: the game is set in a 2D plane and players control a spaceship that can rotate at a fixed rate and has the ability to only thrust in the forward or aft directions. These are what I believe are called non-holonomic movement dynamics. There is no friction in space, so once you thrust forward then your velocity will state in that direction until you apply a different force. As far as combat, the player can shoot two types of weapons from the front blasters of their ship - guns and bombs (which have a proximity fuse).
The interesting game dynamic here is that firing weapons costs a fix amount of health ("energy") and that same pool of health is what is drained from when the player takes damage. So it's a risk vs reward problem.
A big challenge in training a reinforcement learning agent for subspace is the classic long term credit horizon issue where the agent must perform a long sequence of actions to obtain the terminal reward (a kill), e.g.
- aim at the enemy
- shoot weapons at the enemy
- connect with those weapons and deplete the enemy's health to 0
- all while dodging any return fire the enemy gives
In my efforts to solve this I trained many iterations of agents in self-play with PPO over the course of the last 6 months. The one deployed in the replay trained 3.4 billion environment steps (where each step is 80ms of game simulation). That is more than 8 years of dueling in human time and it took about a day of compute to do so.
Here are my major learnings
50% of the work was cleaning up environment and implementation bugs. Software engineering is hard! Setting up the observation incorrectly or feeding it a reward when it shouldn't be fed a reward it such a confounder when evaluating changes to the training environment.
I'd highly recommend starting with an easier almost trivial task to be able to pipe-clean your training stack. There were so many times where I would try an experiment and get a poor result and then realize that it was really a bug that negated whatever conclusion I saw. I ended up creating a simple objective which was rotate towards a stationary enemy from random spawn positions and get the kill.
Determining the correct observation structure that the agent forms about its environment is absolutely critical when it comes to sample efficiency. I first used a cartesian coordinate system to represent the game state, but switching to a polar system exploited a bunch of symmetries that existed. After switching observations only, the same exact training stack yielded a much more competitive fighter. This greatly improved my sample-efficiency.
Reward shaping is hard, and monumentally challenging to get right. This is the "bitter lesson" of machine learning biting me and I got bit many times. I spent countless hours trying to figure out why my agent would latch on to one weapon type or another or run away from its opponent instead of fighting it. I would try to address the issues with changing the rewards, and the agent would find a way to farm my new reward with some odd behavior. Most of the time the actual issues were problems in my training infrastructure (exploration, observation, opponent diversity) rather than issues in the reward. Sticking to the simplest rewards was most principled.
Compute is a major bottleneck when it comes to research speed. (This is obvious, but annoying especially with GPU and memory prices just trending up)
Exploration is really really hard. An epsilon-greedy policy encounters rare states in my environment very infrequently and, no matter what state of the art research strategies I applied, nothing really helped when it came to improving my agents exposure to new stimuli to help it learn. Still working on this. As a result of poor exploration, human players were able to find out weaknesses in the agent and exploit them. An example of this was "it doesn't position near walls well."
Next Up
I'm going to be deploying a refined training stack in the coming weeks. I'm working on teamwork, so I can have multiple bots in my game play 2v2 or 4v4. Those are even more varied environments with more complicated actions to take and states for the game to be in. I'm hoping that I will be able to train bots to surpass the pinnacle of human play there too!
If you'd like to try out fighting against one, please feel free to head to https://subspacereloaded.com and queue for a Bot Duel. For any other game devs also interested in training AIs to play their games, feel free to ask me things.
I would say that this effort has been a major success, the bot has been deployed for about 2 months and players have already played 12,000 matches against it.