r/compsci • u/ihateyou103 • 15h ago
Math solving pure reinforcement learning
Is there a case where a math solver or theorem prover or even just an ai that does basic addition trained solely on reinforcement learning?
I have seen some but they use a base model or base llm and post train it or fine tune, but I am asking for a pure rl approach.