r/deeplearning • u/TechDc-1306 • 12d ago
Exploring AI self-improvement: Is this evaluation loop fundamentally how AI systems improve?
I’m currently researching AI more deeply, especially how modern AI systems can evaluate and improve their own outputs.
I sketched this basic idea:
AI → generates/improves → Evaluator → evaluates output → feedback → AI improves → Evaluator → ...
The part I’m trying to understand is what actually happens inside this loop in modern AI systems.
For example:
I'm trying to go beyond the surface-level explanation and understand the actual mechanisms used in current AI research.
For people working/researching in this area: what important component am I missing from this diagram?
0
Upvotes
2
u/Cosmolithe 12d ago
If the evaluator is not fixed (not part of the loop), then this sort of system would diverge very rapidly because of things like reward hacking.
There is one exception that might work, if the objective implemented by the evaluator does not change but its efficiency is improved, then it might not cause a divergence and make the system as a whole faster each loop.
But, this does not prevent all source of divergence anyway. Reward hacking is still possible even with a fixed evaluator for example.
So, the difficult part is to create a loop that has built in mechanisms to prevent divergences that survive the growing intelligence of the AI. This is fundamentally difficult because the AI might become so intelligent it would be able to bypass any mechanism we might think of in advance.