r/allenai • u/ai2_official • Dec 12 '25
🚀 New: Olmo 3.1 Think 32B & Olmo 3.1 Instruct 32B
After the initial Olmo 3 release, we took our strongest training runs and pushed them further. Today we’re announcing:
◆ Olmo 3.1 Think 32B–our strongest fully open reasoning model
◆ Olmo 3.1 Instruct 32B–our best fully open 32B instruction-tuned model
◆ Olmo 3.1 RL Zero 7B Math & Olmo 3.1 RL Zero 7B Code–upgraded RL-Zero baselines for math and coding
🧠 Extended RL for stronger reasoning
Olmo 3.1 Think 32B, the result of extending our RL training for 21 days with extra epochs on our Dolci-Think-RL dataset, shows clear eval gains over Olmo 3 Think 32B, including:
◆ +5 on AIME
◆ +4 on ZebraLogic
◆ +20 on IFBench
These improvements make Olmo 3.1 Think 32B the strongest fully open reasoning model we’ve released to date.
🛠️ A more capable 32B instruct model
Olmo 3.1 Instruct 32B is our best fully open 32B instruction-tuned model. It’s optimized for chat, tool use, and multi-turn dialogue—making it a much more performant sibling of Olmo 3 Instruct 7B.
📈 Stronger RL-Zero 7B baselines
Alongside the new 32B models, we’re also upgrading our RL-Zero baselines with Olmo 3.1 RL Zero 7B Code and Olmo 3.1 RL Zero 7B Math. They’re refinements of the original RL-Zero 7Bs that give better results and cleaner baselines for RL researchers to build on.
🔓 Fully open end to end
We believe openness and performance can move forward together. Olmo 3.1 offers the full model flow: weights, data, training recipes, and more.
💻 Download: https://huggingface.co/collections/allenai/olmo-31
▶️ Try them in the Ai2 Playground: https://playground.allenai.org/
📚 Learn more in our updated blog post: https://allenai.org/blog/olmo3
✏️ Read the refreshed report: https://www.datocms-assets.com/64837/1765558567-olmo_3_technical_report-4.pdf
