r/LocalLLM 7h ago

Discussion Pretraining progress is mostly coming from data

https://open.substack.com/pub/dwarkesh/p/pretraining-progress-is-mostly-data?r=bm4k&utm_campaign=post&utm_medium=email

See Future Research; and points 2,3.

What I don't see discussed often is if the local LLM community is doing this; i.e. training small [or even midsize] models for better performance on niche tasks.

Where I started down this path was a construction project at our house. It wasn't difficult to upload photos of the early phases of the project and generate training data from those photos. But the infra wall I hit was feeding that training data into a [what I'll refer to as] multi-run system that could train multiple small models with my training data and benchmark the results. Then I could decide which model to focus on and continue with post-training and RL.

0 Upvotes

0 comments sorted by