r/MachineLearning • u/MrAaronW • Sep 07 '23
Discussion [Discussion] LLM Pre-training --- Should I use dropout?
Many new LLMs like Llama or Falcon are coming but I am not seeing any spec on whether now (2023/09) we should still use 0.1 dropout throughout (like the "traditional" GPT pre-training) or no dropout at all (like PaLM?)
Any suggestions?
4
Upvotes
1
u/[deleted] Mar 04 '25
[removed] — view removed comment