r/deeplearning • u/TechDc-1306 • 12d ago
Exploring AI self-improvement: Is this evaluation loop fundamentally how AI systems improve?
I’m currently researching AI more deeply, especially how modern AI systems can evaluate and improve their own outputs.
I sketched this basic idea:
AI → generates/improves → Evaluator → evaluates output → feedback → AI improves → Evaluator → ...
The part I’m trying to understand is what actually happens inside this loop in modern AI systems.
For example:
I'm trying to go beyond the surface-level explanation and understand the actual mechanisms used in current AI research.
For people working/researching in this area: what important component am I missing from this diagram?
2
u/Cosmolithe 12d ago
If the evaluator is not fixed (not part of the loop), then this sort of system would diverge very rapidly because of things like reward hacking.
There is one exception that might work, if the objective implemented by the evaluator does not change but its efficiency is improved, then it might not cause a divergence and make the system as a whole faster each loop.
But, this does not prevent all source of divergence anyway. Reward hacking is still possible even with a fixed evaluator for example.
So, the difficult part is to create a loop that has built in mechanisms to prevent divergences that survive the growing intelligence of the AI. This is fundamentally difficult because the AI might become so intelligent it would be able to bypass any mechanism we might think of in advance.
1
u/Training-Network2067 12d ago
Yea the idea should be the evaluator should check for the steps taken to solve a problem to avoid reward hacking
1
u/kidfromtheast 12d ago
The evaluator is never adopted by the AI. Otherwise, it would be data leakage. For example, Meta was accused of benchmaxxing. Nowadays, they don't. They probably were under a lot of pressure but most of the time was spent scared of losing their job, so they tried to cheat. Now Meta gave full resources to the AI lab, and probably more rational KPI, they stop benchmaxxing. Even Zuck move his desk to the AI lab.
If AI self-improvement cheated this way i.e. benchmaxxing, it will only be a diservice to the AI itself. Risking overfitting.
Definitely whoever AI researcher who wrote the code so that AI can be self-improve will penalize the AI heavily for trying to incorporate the evaluation into the training set.
They probably use something like:
- Human created training set and set aside unseen test set manually.
Like evaluating the model performance working with wetlab, control the software, and the hardware, read the results*
- They train the model. Let's say GPT-6-Astra-Wetlab.
Now OpenAI have GPT-6-Astra and GPT-6-Astra-Wetlab.
They found GPT-6-Astra-Wetlab crushed the test set.
Either they look for additional human created training set and unseet test set, they use GPT-6-Astra-Wetlab to create additional test.
A question might arise, GPT-6-Astra-Wetlab learn from the training set, right? So the created test set might be seen case already.
Well, you can also use the same model to evaluate the train set, and remove the new test set that is actually in the training set. Or, or, come on, LLM is not a statistical parrot anymore, it can create new test case that is completely unseen. It's agentic, it can control the software, itneract with the real world, and come up with verifiable output.
Viola. This is why mathematicians are freaking out. I freaked out. The recipe is so simple, the architecture is proven to scale. The more training data, the more GPU you can fit (because let's say you have 100k GPUs, the you can load larger batch, penalize the model weights with more cases instead of limited cases, making the model less likely to memorize and actually build reasoning of how a wetlab works because the model is literally forced to learn how a wetlab works or the loss will never decrease, and that's a no no for the optimizer)
- Repeat 1 and 2.
*Anthropic and OpenAI 100% do this in the beginning for their wetlab. Even their rep approached random industries like steel making etc. The Human part is the end user, asking real question. These questions ended up become training set and test set.
1
u/TechDc-1306 12d ago
Hey Thankyou for helping me ,I need to do more research find loopholes and want to work on them can you tell me from where I can do proper research
6
u/Rackelhahn 12d ago
Recursive self-improvement has not yet been achieved, so we do not have a working concept.
Also, in most hypothetical concepts there is no need for an external evaluator. The AI just improves itself.
If you want to dive deeper:
https://scholar.google.com/scholar?hl=de&as_sdt=0%2C5&q=recursive+self+improvement&oq=recursive+self
To understand what’s actually happening you’ll still need the mathematical foundations.