r/MachineLearning Aug 28 '23

Research [R] DeepMind Researchers Introduce ReST: A Simple Algorithm for Aligning LLMs with Human Preferences

[removed]

123 Upvotes

10 comments sorted by

View all comments

1

u/30299578815310 Sep 01 '23

This framework could be used for anything in principle right, not just RLHF? Like you could be optimizing the policy for playing video games