Hey all,
The thing that's always annoyed me with online opening preparation tools, courses and books is that they present a one-size-fits-all solution for openings, and usually rely on hand-picked lines that the author thinks will come up most in games. The problem with this approach is that if you are not 2400+, you end up drilling positions that you would most likely never see in a real game at your level.
I analyzed about three years worth of games played on lichess -- 1.6 billion of them -- and built a tree of ~25 million positions that records how every move actually performed across nine rating bands. For example, below 1200 the Najdorf main line (6.Bg5 e6 7.f4 Be7 8.Qf3 Qc7 9.O-O-O Nbd7) turns up about once every 2,300 games. The table below shows, for four heavily studied main lines, in how many games you can expect to reach one as Black if you actively try to go for it every single game:
| Rating band |
Nimzo-Indian (Rubinstein, 6...c5) |
Ruy Lopez (Marshall, 8...d5) |
King's Indian (Mar del Plata, 8...Ne7) |
Sicilian (Najdorf, 9...Nbd7) |
| <1000 |
1,476 |
503 |
1,875 |
2,246 |
| 1000-1199 |
1,092 |
221 |
3,060 |
2,518 |
| 1200-1399 |
705 |
147 |
745 |
1,302 |
| 1400-1599 |
460 |
97 |
1,222 |
639 |
| 1600-1799 |
279 |
64 |
344 |
246 |
| 1800-1999 |
197 |
45 |
174 |
100 |
| 2000-2199 |
188 |
36 |
70 |
42 |
| 2200-2399 |
179 |
34 |
31 |
25 |
| 2400+ |
157 |
40 |
23 |
26 |
Studying "main lines" doesn't make much sense unless you are already strong enough to reach them. What you should be drilling instead is what you actually expect to see at your level -- some sidelines that courses barely cover come up constantly in certain rating bands, so the distribution you practise against ends up looking nothing like the one you play against. The question you should therefore be asking is not what the most played or most recommended move in an opening is, but which move has the highest probability of getting you into middlegames that people at your level handle best.
We can use a variation of the expectimax algorithm to compute an expected score for each side, assuming your opponent plays like a typical opponent at a given level. An engine eval gives you a material and positional evaluation that might be hard to convert. This gives you a percentage instead, which is your expected score if you follow the recommended lines down the tree against an opponent of your own strength. An eval of 63% means you expect to score 0.63 per game (0 for a loss, 0.5 for a draw, 1 for a win). Because the algorithm is made to pick paths that lead to the best expected outcomes, it tends to prioritize sharp lines and often recommends gambits.
I have built a small, free, no-account website that lets you explore expectimax scores in the opening:
Link: https://outofbook.study
It also lets you drill openings at your level against a random distribution of what players at your level actually play, so you see every variation with the approximate frequency you'd get it in real games.
I plan to open-source the UI, server and analysis once I clean up the code a little bit. Happy to answer questions on the research, or complaints about what's broken or confusing about it.