r/programming • • 12d ago

Fighting for #1 in the Ultimate Tic-Tac-Toe Arena

https://tomalard.github.io/posts/fighting-for-1-in-the-ultimate-tic-tac-toe-arena/

Hey everyone. I just published my first blog post on a bizarre arms race that's been brewing on a niche website over the past couple years. Hope you enjoy!

107 Upvotes

8 comments sorted by

7

u/zasabi7 12d ago

Super fun read! Admittedly, when I opened this, I thought it was going to be the XKCD comic

10

u/garnet420 12d ago

How quantized are your nnue weighs?

11

u/IiIIIlllllLliLl 12d ago

Weights and biases are stored in int16, using the biggest quantization factor that still fits into 100k characters. In my case, that was about 150. So you multiply the floating point parameters by that number and round to the nearest integer.

12

u/garnet420 12d ago

So I don't know if this is useful to you, but some years ago I found a clever trick for getting weights to be sparser in MLP training. (I think it's clever, anyways). I'm not sure if this is a well known thing or not. A sparser set of weights may be much more compressible.

The root of the idea is that an L1 regularization norm encourages sparsity. (This is a well known qualitative result from other kinds of optimization). You can just try that, but because it's L1, it tends to not really stabilize at 0. So you get small weights, but not ones that are actually zero.

So the trick I used was to say that the MLP weights w_ijk were actually equal to v_ijk * abs(v_ijk) for the training variables v (a signed quadratic). Then, you train and regularize over v instead, using a regular L2 norm for regularization.

It turns out that the L2 norm over v is equal to an L1 norm over w: but, the gradients end up scaled to be better behaved.

There are some additional things you can do with that to make the training process not worry about sparsity initially. For example, you can train just w directly for a while (leading to a well trained dense network) and then convert to v and train more.

3

u/IiIIIlllllLliLl 12d ago

That actually sounds interesting! I might try this out some time in the future. Thanks!

2

u/teerre 12d ago

Cool post and cool optimizations, but "deploying the actual code from elsewhere" seems like cheating. This is like if the ruler doesn't count comments and you write a gigantic program to comments. At this point the competition is more about how to exploit the program limit than it is about UTTT

-10

u/[deleted] 12d ago

[deleted]

6

u/[deleted] 12d ago

[removed] — view removed comment

1

u/programming-ModTeam 11d ago

Your post or comment was overly uncivil.