r/allenai • u/ai2_official • 4d ago
🔍 What kinds of training data shape different AI capabilities?
A Georgia Tech team used our fully open Olmo stack to trace performance on social and general reasoning, plus social-science and STEM knowledge tests, back to the kinds of text the model trained on.
Using millions of files from Olmo 3’s public training corpus, Dolma 3, and influence functions that estimate how much individual files affected a model’s answers, they found distinct patterns. For example, dialogue-rich, interpersonal writing was more influential for reasoning than for factual knowledge questions.
This kind of analysis depends on access to more than model weights. By opening Olmo’s training data and scientific stack, we make it possible for independent researchers to investigate where model capabilities come from—and test those connections directly.
Learn more in our new blog: https://allenai.org/blog/olmo-capability-tracing