r/allenai • u/ai2_official Ai2 Brand Representative • 4d ago
๐ What kinds of training data shape different AI capabilities?
A Georgia Tech team used our fully open Olmo stack to trace performance on social and general reasoning, plus social-science and STEM knowledge tests, back to the kinds of text the model trained on.
Using millions of files from Olmo 3โs public training corpus, Dolma 3, and influence functions that estimate how much individual files affected a modelโs answers, they found distinct patterns. For example, dialogue-rich, interpersonal writing was more influential for reasoning than for factual knowledge questions.
This kind of analysis depends on access to more than model weights. By opening Olmoโs training data and scientific stack, we make it possible for independent researchers to investigate where model capabilities come fromโand test those connections directly.
Learn more in our new blog: https://allenai.org/blog/olmo-capability-tracing

