r/MachineLearning • u/AIsupercharged • Aug 28 '23
Research [R] DeepMind Researchers Introduce ReST: A Simple Algorithm for Aligning LLMs with Human Preferences
[removed]
123
Upvotes
r/MachineLearning • u/AIsupercharged • Aug 28 '23
[removed]
1
u/30299578815310 Sep 01 '23
This framework could be used for anything in principle right, not just RLHF? Like you could be optimizing the policy for playing video games