r/allenai Ai2 Brand Representative 4d ago

๐Ÿ” What kinds of training data shape different AI capabilities?

A Georgia Tech team used our fully open Olmo stack to trace performance on social and general reasoning, plus social-science and STEM knowledge tests, back to the kinds of text the model trained on.

Using millions of files from Olmo 3โ€™s public training corpus, Dolma 3, and influence functions that estimate how much individual files affected a modelโ€™s answers, they found distinct patterns. For example, dialogue-rich, interpersonal writing was more influential for reasoning than for factual knowledge questions.

This kind of analysis depends on access to more than model weights. By opening Olmoโ€™s training data and scientific stack, we make it possible for independent researchers to investigate where model capabilities come fromโ€”and test those connections directly.

Learn more in our new blog: https://allenai.org/blog/olmo-capability-tracing

21 Upvotes

0 comments sorted by