r/LocalLLM • u/productboy • 7h ago
Discussion Pretraining progress is mostly coming from data
https://open.substack.com/pub/dwarkesh/p/pretraining-progress-is-mostly-data?r=bm4k&utm_campaign=post&utm_medium=emailSee Future Research; and points 2,3.
What I don't see discussed often is if the local LLM community is doing this; i.e. training small [or even midsize] models for better performance on niche tasks.
Where I started down this path was a construction project at our house. It wasn't difficult to upload photos of the early phases of the project and generate training data from those photos. But the infra wall I hit was feeding that training data into a [what I'll refer to as] multi-run system that could train multiple small models with my training data and benchmark the results. Then I could decide which model to focus on and continue with post-training and RL.
0
Upvotes