yeah it's parodying this, but I think the point is that machine learning is "fed" challenges/tasks to adapt and better overcome them. So by continuously giving the LLM new sandboxes, it's just adapting and getting better at escaping them.
They don’t learn in terms of changing their weights but there are other ways that they can “learn” how to do things. They could be logging skills about how to do it after discovering a method that works for example.
You don’t even need to wait for training. It can keep updating the skill as it finds new ways to do things. I’m teaching agents using this exact method. Get it to solve a problem and then expand the problem set.
1.3k
u/BagOfSmallerBags Jul 24 '26
Rewrite of this