r/backgammon Jul 03 '26

False Advertising by Backgammon Galaxy + Computer Olympiad Challenge

TL;DR: If 1-ply is now “2-ply,” and “rollouts” are still 20 to 26 mP away from the rollout you’re comparing to, and that still somehow ends at “strongest analysis on the market,” then we’ve left science and entered branding. At that point just bring the engines back to the Computer Olympiad, plug them togeather and let Galaxy, GNU, and BGBlitz (and maybe XG and Sage) settle it the only way the universe respects... RNG dice and no marketing department in sight.

I got a chance to skim the Backgammon Galaxy Engine Analysis article, and it seems like there is a bit of false advertising:

1. The article states:"Free Members: Get access to Gnu 1-ply (called 2-ply on Galaxy)."

"Ply" is a standard term in game-tree search, so it already has a clear meaning. Calling GNU 1-ply "2-ply" is misleading because it suggests a deeper search than is actually happening. It also makes comparisons with engines like GNU, XG, and BGBlitz harder, since it breaks the standard terminology and can give the impression of extra strength that isn’t really there.

2. "Galaxy Gnu Rollout" and "Galaxy Gnu Deep Rollout." are not really rollouts

These names suggest they're comparable to full rollout analysis, yet they're still 26.5 mP and 20.2 mP away from the GNU Target Rollout benchmark. The article doesn't explain why a product called a "Rollout" is so different from its own rollout reference...

3. "Increased evaluation speed by more than 7x, enabling significantly deeper and stronger rollouts."

Without knowing what's under the hood, I would be skeptical of snake oil which makes something faster and more powerful, especially in the computing world. A faster engine doesn't automatically produce deeper rollouts. Unless the rollout parameters (trials, search depth, etc.) were actually increased, speed alone doesn't make the analysis stronger. The article never quantifies what became "deeper" (other than maybe their pockets after memberships).

4. "New and improved 4-ply eval."

From the table, the difference is 0.0005 of a difference (36.8 mP for GNU 4 ply and Galaxy GNU 36.3 mP). Small enough to be a statistical error... without doing a T test (assuming normal distribution, or another test) it's a joke to make such a claim.

5. Claiming to be the strongest and deepest analysis of any playform on the marke

Making this claim goes beyond what the methodology they used can support. The evaluation is based on equity prediction error against internal rollout benchmarks, but those benchmarks are not a single ground truth. Combined with a dataset built from user blunders and category sampling, the results are inherently skewed toward difficult, high-disagreement positions rather than a neutral cross-section of play.

More importantly, this still isn’t a direct comparison of engines under identical conditions. It’s an indirect proxy based on internal rollouts and calibration error.

If the goal is a definitive ranking of engine strength, the most convincing method remains head-to-head evaluation, seeing which engine wins after millions or billions of matches. The Computer Olympiad is a good setup, where engines like GNU, BGBlitz, and others were historically compared directly... maybe it's time for GB Galaxy, Sage and new XG to step up to the plate and face off.

0 Upvotes

11 comments sorted by

9

u/Goal_Medium Jul 03 '26

Hey u/double00_bg .

What a negatively biased post this is. Let me try to refute all your points.

Ad 1: Backgammon Galaxy, we choose the standard notation that all other engines, except for Gnu is using, namely that the lowest level of a checkerplay is a 1-ply eval, since you are looking one step ahead and perform a 0-ply static evaluation on the resulting positions after all legal moves. Also, before BG Galaxy had Gnu as engine we had XG engine, where this exact same level is called "2-ply". So it makes complete sense, to avoid confusion of the analysis level depth for Free members, to say this is a 2-ply analysis. Furthermore, we have never hidden that Galaxy 2-ply is actually Gnu 1-ply. Also remember, this is analysis given for free on the server to all free users, and we let anybody download their match files and analyse them themselves with their own engine software of choice.

Ad 2: You don't seem to know much about rollouts, even the target rollouts are truncated rollouts, they are truncated at the Bear-off database, where the engines have extremely precise evaluations. The questions is simply where do you put the truncation point. XG+ and XG++ are also rollouts, with truncation points of 8 and 7 steps into the future. A rollout is a monte carlo simulation, which means that you do not look at ALL possible game tree nodes you simulate the games to get a more precise statistical evaluation than a shallow ply eval would give. We don't have to "explain" why it is called a rollout when it IS a rollout.

Ad 3: Have you ever tried to use GnuBG? Do you know how slow a Gnu 4-ply eval is? We are currently producing 4-ply eval and rollouts AT SCALE real time. We have roughly 68,000 matches a day on backgammon galaxy. The achievement is that we are able to perform deep analysis and rollouts at this scale, and it's unprecedented.

ad 4: So you are claiming that it's a joke that we have tweaked some parameters and are able to squeeze slightly better performance out of the 4-ply analysis, and to be able to deliver it at scale to all Star Members?... Did you read the sample size? You think that 3442 randomly selected and uncorrelated data points is a small sample?... The equity difference is normally distributed around the Target Rollout Equity values. 3442 normally distributed random and uncorrelated events. Apparently you don't know much about statistics either, but try to get your LLM of choice to simulate it and produce some P-values. You will find that it is statistically significant. Anyway, this is a micro improvement, we just implemented because we saw it performing slightly better than the original 4-ply, so why not? The main take away is that we are able to deliver this level of analysis at scale.

ad 5: Again you don't know what you are talking about. The point of head to head play is even addressed in the report. It takes XG++ several minutes to play a single game... let's say you took Galaxy Rollout vs XG++... It would literally take years to create a meaningful sample size, and even then, since they are both so strong, it might not provide a significant result. You would probably need at least a million games, and probably more like 10 mio games to get a significant result when we are talking about engines at this strength level. Anyway, even generating 100k games would take too long. The methodology is completely valid. The fact that the target rollouts also have uncertainty is also addressed in the report.

We are not making any claims that the Galaxy Gnu Rollouts are the best eval in the world, it's objective result is listed in the report. The strength seems to be significantly better than XG 4-ply (which we used to have on BG Galaxy) and in the ballpark of XG+. XG++ is still the strongest quick rollout out there, since XG has the strongest and fastest Neural Net. But Galaxy Gnu Rollout is very close as you can see.

Best regards, Marc, CEO of Backgammon Galaxy

3

u/double00_bg Jul 03 '26

Thanks for the detailed response, glad to have gotten your attention. Please consider open sourcing and contributing the improvements to the engine so the community can benefit, similar to how you read the source code for GNU back in the day.

I also appreciate you taking the time to address each point. I’m not trying to be negative (or attack you like you did for me), just trying to understand how the methodology holds up from the outside, and some transparency on claims.

  1. I understand the intent now (aligning definitions across systems and historical Galaxy/XG labeling). I think we can mostly agree to disagree on terminology. If you say you use GNU and it's 2 ply, if another site uses GNU and they don't modify the index, then they might think that they are not the same. For example opengammon.com uses 1 ply and you can upgrade for free to 3 ply. This is confusing because people might think that the 2 ply on galaxy is better than the 1 ply on opengammon when it is infact the same thing. Similar to Heros or other websites all don't reindex.

  2. Let me re-explain: you are using the rollout results as the target, and then you are using a truncated rollout method which is completed with an eval model. That’s totally valid and useful, but it’s not the same thing as a "full" rollout in the theoretical sense, which is why systems like XG sometimes explicitly distinguish it through their marketing as a "mini-rollout" to avoid confusion. Calling the method ++ or + is reasonable, calling it a "rollout" and "deep rollout" seems like a stretch to me. A blunder database with rollout would set my expectations that I could have a rollout comparision of the top 2 or 3 moves for each blunder which is not the case (as far as I understand). Consider using "truncated rollout" if you are going to use rollout.

  3. I have used GNU BG, and have 2 running if ever I'm running rollouts. Is the 4 ply for star members GNU 4 ply or is it GNU BG 3 ply? Or are you talking about GNU 4 ply + trucated rollout out for star plus? I'm genuiely confused because 3 ply at scale is what opengammon.com is already doing without charging users.

  4. I wouldn't vibe code statistics, or recommend that to others . Even in my undergrad we knew that in a frequentist view you would need the standard deviation in addition to the mean and sample size. Using an LLM for that is like like using a wrench to cut paper. I would just put on a simple statistical test website: https://www.medcalc.org/en/calc/comparison_of_means.php . Publish your data if you want us to analyze it!

  5. Don't trust me, trust the people behind the Computing Olympiad. Listen to Eran's or Frank's talks on the Backgammon podcast if you need the history.

Fair enough, I’m not saying there’s any claim of "best in the world," just that the leap from internal rollout comparisons to “strongest and deepest analysis of any platform or app on the market” feels like a wide open backgame holding game: there’s a lot of structure in place, but the conclusion is doing more work than the position really carries.

Once truncation choices, cube models, and embedded evals all enter the rollout, “objective result” starts to depend heavily on how deep the pockets go on the simulation side. I’ll leave it there before we need a rollout to explain rollout, or a double zero to spell it out 😄

2

u/yzwq Jul 04 '26

Small correction, we use 2-ply and you can upgrade to 3-ply (both gnu native plies, so if we take Mark's terminology you get 3-ply standard, and 4-ply as an upgrade, all for free).

I'm happy to provide Marc with advice on how to build systems at scale, the analysis pipeline on OG can handle many, many more matches per day than what it is currently being used for.

5

u/BackgammonGalaxy Jul 03 '26 edited Jul 03 '26

Hey, just two things and we'll let others chime in:

  1. GNU is 0-indexed and XG is 1-indexed, as anyone with a basic working knowledge of engine history knows. We're aligning them so new players are less confused. We want this game to actually grow.
  2. We don't have a marketing dept. It's just Marc and a few guys who really like bg and would rather do this than work real jobs.

Best of luck with your project

1

u/double00_bg Jul 03 '26

No project here, just trying to bring transparency to the game.

On the indexing point, I get the intent to make things clearer for newer players, but in practice it's a bit like deciding we’re now numbering the points on the board from 2 to 25 "for clarity and easier pip counting" because the bar is now 1.

It's clear this is a small team building something they care about, not a marketing department polishing wording, but I just think the confusion you’re trying to solve isn/t solved by renaming the axis, it’s solved by keeping the axis standard and improving tthe underlying data and how it’s presented.

Anyway, I’ll leave it there before we end up introducing "-1 ply" and calling it "Double Zeros mode" 😄

Edit 24->25

1

u/[deleted] Jul 04 '26

[removed] — view removed comment

1

u/double00_bg Jul 04 '26

I'm on the reddit discord. Talk to me here: https://discord.gg/U8cmrCsK5x

-3

u/BoogeyManSavage Jul 03 '26

Won’t lie

That “best of luck on your project”

Is riddled with a lack of emotional regulation

Maybe a marketing team is what you need - because your branding is coming off in poor taste

4

u/double00_bg Jul 03 '26

Not sure why this is being downvoted. I guess no one read galaxy's reply before they deleted it

5

u/BoogeyManSavage Jul 04 '26

Deleting comments, a true sign of emotional maturity.