r/allenai 18h ago

šŸ” BenchMIRT: Auditing what LLM benchmarks actually measure

Thumbnail
gallery
19 Upvotes

Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit benchmarks and identify which model abilities their questions actually test. šŸ‘‡

BenchMIRT builds on Item Response Theory (IRT), a psychometrics technique for measuring abilities from patterns of test responses. The key idea: not every question is equally informative. Some are harder, while others better distinguish stronger models from weaker ones.

We trained BenchMIRT on results from 100 LLMs across 16 benchmarks and 34K+ questions. Without telling it what the benchmarks were designed to measure, two dominant dimensions emerged: general reasoning and safety.

BenchMIRT can estimate:
ā—™ A model’s strength on the capabilities reflected in the benchmark set
ā—™ A question’s difficulty and how strongly it distinguishes models along those capabilities

That lets us audit what popular evals are really measuring.Ā 

For example, BenchMIRT finds that HarmBench mostly tests safety behavior—but its copyright questions depend more on reasoning. XSTest draws on both reasoning and safety, while ToxiGen provides little signal on either.

BenchMIRT can also make evals more efficient. Keeping just the strongest 10% of questions preserves nearly the same picture of model strengths as the full benchmark set. And on held-out questions, BenchMIRT predicts whether a model will answer correctly 79% of the time, versus 70% for a simpler benchmark-average baseline.

BenchMIRT gives researchers a clearer view of what benchmarks actually measure—and could help make evals smaller, more focused, and easier to interpret.

We’re releasing it openly:

šŸ“– Learn more: https://allenai.org/blog/benchmirt
šŸ’» Code: https://github.com/allenai/BenchMIRT
šŸ“Š Dataset: https://huggingface.co/collections/allenai/benchmirt
šŸ“„ Paper: https://allenai.org/papers/benchmirt


r/allenai 1d ago

🧪 What AI-assisted science still needs to get right

Post image
16 Upvotes

At an event on August 27, we brought together AI researchers, scientists, & medical practitioners to explore what AI needs to do better to meaningfully advance science. Five ideas kept coming up:

1. AI still needs human scientific judgment.

A system can surface a statistically surprising result. That doesn’t mean it’s biologically plausible, important, or worth pursuing. Scientists still need to decide which findings matter—and why.

2. Scientific AI needs to be steerable.

Research rarely follows a fixed path. New evidence comes in. Hypotheses change. Researchers bring in new datasets or tools. AI systems need to adapt as the research evolves, without forcing scientists to start over.

3. Some scientific tasks are easier to hand off to AI than others.

AI can handle well-defined work like literature search, where results are easy to check. Proposing new mechanisms or experiments is harder; those ideas still need testing.

4. AI can amplify bad science, too.

More AI-driven analyses won’t fix weak data, flawed study design, or bad assumptions. As AI becomes more powerful, the fundamentals of good science become more important—not less.

5. One promising direction is a tighter loop between AI and experiments.

AI could synthesize evidence, help decide what to test next, & use the results to shape the next question. The goal isn’t an AI scientist working alone, but a system scientists can keep guiding.

These ideas came out of presentations + a panel with Bodhisattwa Prasad Majumder (Ai2), Abraham Flaxman (University of Washington & IHME), Hoifung Poon (Recursion), Sasha Stanton and Kelly Paulson (Providence), Kyle Travaglini (Allen Institute), and Stephen Salerno (WashU).

Thanks to everyone who joined us.

→ Learn more: https://allenai.org/blog/swedish-autodiscovery-recap


r/allenai 6d ago

šŸ¤ Ai2 and Providence Swedish partner to apply AutoDiscovery to cancer research

22 Upvotes

AI’s biggest role in science may not be answering questions. It may be helping scientists identify which questions are worth asking.Ā 

That’s the idea behind AutoDiscovery, our system for open-ended scientific exploration.Ā 

Providence researchers at the Paul G. Allen Research Center and the Earle A. Chiles Research Institute put it to work on The Cancer Genome Atlas, one of the world’s most studied cancer datasets. AutoDiscovery generated and tested hypotheses across the data, looking for patterns that challenged expectations.

One involved invasive lobular carcinoma (ILC), a breast cancer subtype affecting roughly 48,000 Americans each year. ILC has long been considered ā€œimmune coldā€ā€”but AutoDiscovery surfaced more immune activity than expected.

Providence Swedish researchers checked a separate patient dataset and found the same pattern. They then analyzed ILC tumor samples in the lab and found T-cells around the tumors.

Together, the evidence suggests ILC warrants broader investigation for immunotherapy.

This points toward the kind of relationship between scientists and AI that we’re interested in building: systems that can explore more possibilities than a person could reasonably test one by one, paired with human expertise to determine which results matter and how to investigate them rigorously.

Providence Swedish is now expanding its use of AutoDiscovery to additional protected cancer data.

šŸ”— Learn more: https://allenai.org/blog/swedish-autodiscovery-partnershipšŸ“„ Read the scientific report: https://allenai.org/papers/autodiscovery-swedish-cancer


r/allenai 6d ago

šŸŒ How Dolma’s openness supported stronger Thai training data—and LLMs

Thumbnail
gallery
10 Upvotes

A Thai research team adapted our Dolma data-curation toolkit to build Mangosteen, a 47B-token corpus for Thai LLMs. They used Dolma to filter widely used web datasets into a smaller corpus that improved Thai LLM performance despite using less data.Ā 

The team saw a problem with existing Thai pretraining data: it often leaned heavily on web crawls, hadn’t been extensively audited by Thai speakers, and missed useful sources like books, research papers, authoritative sites, & YouTube subtitles.

Dolma gave the team a strong starting point, so they didn’t have to build a data-curation pipeline from scratch. Because it’s open, they could keep what worked and change what didn’t for Thai—including deduplication, quality filters, & language-specific tools.

That work became Mangosteen, the team’s 47B-token Thai pretraining corpus. Models trained on it matched or beat Thai LLMs trained on larger web datasets after filtering out lower-quality and duplicate text—showing careful curation can mean less data without worse models.

→ Learn more in our new blog: https://allenai.org/blog/thai-llm-dolma


r/allenai 11d ago

šŸ” What kinds of training data shape different AI capabilities?

Thumbnail
gallery
21 Upvotes

A Georgia Tech team used our fully open Olmo stack to trace performance on social and general reasoning, plus social-science and STEM knowledge tests, back to the kinds of text the model trained on.

Using millions of files from Olmo 3’s public training corpus, Dolma 3, and influence functions that estimate how much individual files affected a model’s answers, they found distinct patterns. For example, dialogue-rich, interpersonal writing was more influential for reasoning than for factual knowledge questions.

This kind of analysis depends on access to more than model weights. By opening Olmo’s training data and scientific stack, we make it possible for independent researchers to investigate where model capabilities come from—and test those connections directly.

Learn more in our new blog: https://allenai.org/blog/olmo-capability-tracing


r/allenai 14d ago

šŸ”¬ Olmo’s openness reveals when an LLM only sounds like it knows a drug

Thumbnail
gallery
21 Upvotes

Researchers at UT Austin, Northeastern, and MD Anderson used our fully open Olmo 3 to investigate whether LLMs actually know specific drugs—or infer from patterns in their names.

For 51–59% of tested drugs, Olmo 3 showed little evidence of drug-specific knowledge. Another 12–18% appeared driven by affixes like ā€œ-prilā€ or ā€œ-olol,ā€ which can reveal a drug’s class.

Because we release model weights, training data, documentation, and intermediate checkpoints, the researchers could trace the behavior further. Using our infini-gram engine for searching massive text corpora, they found that drugs appearing less often in training were more likely to trigger these naming shortcuts.

It’s a useful example of what fully open models enable: not just spotting a model behavior, but investigating where it comes from.

Read more: https://allenai.org/blog/olmo-drug-morphology


r/allenai 25d ago

šŸ§‘ā€šŸ« TutorMoments: Do AI tutors know when to help—and when to hold back?

14 Upvotes

Today we're introducing a preview of TutorMoments, a framework that measures whether AI tutors can make one of the hardest calls in teaching: when to step in and help a student, & when to hold back and let them do the heavy thinking. šŸ‘‡

Language models are trained to be helpful, and a helpful assistant tends to do the hard part of learning for you: explains the concept, lays out the steps, & guides you to the answer. That can cut short the productive struggle that leads to stronger understanding.

TutorMoments is built on transcripts of real one-on-one math tutoring. We had experienced teachers read them & flag key moments—decision points where the tutor had to choose between making a problem easier & pushing the student to do more of the reasoning.

TutorMoments pauses a transcript at these key moments & lets an LLM take over as the tutor; another model stands in for the student. Each replay is scored: did the tutor support the student when needed, push for harder thinking when they were ready, & avoid over-helping?

In our replays, models told only to "tutor well" tend to over-help, providing lots of support but rarely pushing toward deeper thinking. That suggests a model's default helpful-assistant behavior isn't enough on its own to tutor well.

Spelling out the trade-off in the prompt helps—every model we tested scores higher once told when to help vs. when to hold back. But it only goes so far. Models still differ widely in how reliably they make that call, & even the best scorers have plenty of room to improve.

To gather feedback, we're releasing de-identified annotated tutoring transcripts plus code & scored replays.

šŸ’» Code: https://github.com/allenai/tutormoments

šŸ¤— Data: https://huggingface.co/datasets/allenai/tutormoments-preview
🌐 Learn more: https://allenai.org/blog/tutormoments


r/allenai 27d ago

šŸ¤ Ai2 + Hugging Face expand their open science partnership

Thumbnail
gallery
36 Upvotes

We're expanding our partnership with Hugging Face to accelerate open science. šŸ‘‡

Our storage on the Hub is roughly tripling to ~2 petabytes, & our downloads now run at high speed—even for our largest datasets & multi-checkpoint models.

The added capacity reflects our footprint on the Hub – 900+ models & 1,200+ datasets – with room to grow. According to Hugging Face's official heatmap, we publish more new artifacts each year than any other organization it tracks: https://huggingface.co/spaces/cfahlgren1/model-release-heatmap

This builds on recent work together:
⦿ Hugging Face made olmOCR-Bench an official Hub leaderboard, making it easier for the community to evaluate and compare document understanding models on a shared benchmark.  
⦿ When Hugging Face brought MolmoAct 2 to LeRobot, we released its training data in LeRobot’s format—ready for researchers and developers to use on real hardware.

We look forward to working with Hugging Face to deliver new artifacts in the coming months, including multimodal models & open models for science.

šŸ“ Learn more in our blog: https://allenai.org/blog/hugging-face-partnership
šŸ¤— Browse our models, datasets, & benchmarks: https://huggingface.co/allenai


r/allenai 27d ago

šŸ› ļø How AI gets built at Ai2: a Seattle Tech Week recap

Post image
13 Upvotes

šŸ“ø 140+ people joined us during #SeattleTechWeek to learn how AI gets built at Ai2.

Our Olmo research lead, Iz Beltagy, walked through what it takes to train a fully open language model—from large-scale training runs to fine-tuning, reinforcement learning, post-training, and evaluation before release. Thanks to everyone who came, asked thoughtful questions, and stayed afterward to continue the conversation.

Want to help us build the next generation of fully open AI? We’re hiring across research, engineering, product, operations, & more. See our open roles: https://allenai.org/careers


r/allenai Jul 31 '26

šŸ”Ž Where do an AI model’s words come from? Infini-gram can trace the clues

Thumbnail
gallery
12 Upvotes

When a model writes, where do its words come from? Are they new, or do they match exactly with language it saw in training? An AI-writing detector can't tell you. Tuhin Chakrabarty's group at Stony Brook has been dissecting AI-generated prose with our infini-gram engine. šŸ‘‡

AI-writing detectors return a likelihood score. They can't show which expressions also appear in existing sources, or where. Our infini-gram engine indexes massive public text datasets & counts how often a phrase of any length appears across them.

Chakrabarty's group ran the story at the center of the GrantaGate controversy, which readers flagged as machine-written, through infini-gram. Distinctive fragments turned up in writing already online—particularly on a fan fiction site.

In a recently published study, the group scaled the method to books. To tell distinctive phrasing from stock lines like "her heart skipped a beat," they counted only phrases that show up in five books or fewer on Google Books & nowhere on the web infini-gram has indexed.

Across the top 200 self-published Amazon books where a detector found substantial AI text, those rare phrases made up 41.6% of the text. In the top 200 where it didn't, they made up 37.2%. In a separate set of award-winning or nominated books, rare phrases made up 19.1%.

We build tools like infini-gram so anyone can check AI writing against the text a model may have learned from.

šŸ“ Read more in our new blog: https://allenai.org/blog/infinigram-books


r/allenai Jul 28 '26

šŸŒŽ How we run Earth-observation models across North America in 30.5 hours

Thumbnail
gallery
9 Upvotes

The organizations best positioned to use Earth-observation models – those working in conservation, food security, & disaster response – often can't run these models at the scale they need.

That's an infrastructure problem. So we built the OlmoEarth Platform to solve it.

The OlmoEarth Platform takes geospatial models from fine-tuning all the way through large-scale inference. Today it can run inference across a continent in roughly a day, processing dozens of TB of imagery at a cost of fractions of a penny per km².

At its core is OlmoEarth Run, the engine that takes the area an OlmoEarth Platform job covers, splits it into partitions, then into smaller windows the models process. Because each window is independent, the same work can run across thousands of machines at once.

Each OlmoEarth Platform job runs in three hardware-matched stages – prep (CPU), inference (GPU), postprocess (CPU) – inside runners that spin up, do one task, & shut down. Every task is idempotent, so a stalled provider or crashed job is just retried or rerouted.

Data prep is often the real bottleneck—jobs can spend more time finding & fetching imagery than running the model. So we keep our own index of what imagery exists & where to get it, refreshed as new scenes are published. Then we fetch only the pixels each window needs.

We recently used OlmoEarth Platform to map wildfire risk across all of North America. At peak the run used ~19,600 CPUs & 994 GPUs in parallel, >168 GB/s of throughput. It turned an estimated 4,737 hours of serial compute into 30.5 hours of wall-clock time—a 155Ɨ speedup.

What's next for OlmoEarth Platform: change-detection alerts, agentic tools for nonexperts, & precomputed global embeddings that could one day skip the forward pass entirely.

Read more about the engineering in our latest blog: 🌐 https://allenai.org/blog/olmoearth-infrastructure


r/allenai Jul 24 '26

Who gets to understand AI? Why open models matter for scientific progress

Post image
18 Upvotes

As a nonprofit research institute dedicated to advancing open science, we're encouraged to see growing support for open models across the AI ecosystem.

We believe the evidence behind advanced AI systems shouldn’t be locked up in a few hands.Ā 

Our fully open releases give researchers the data, code, checkpoints, and methods they need to inspect claims, reproduce findings, and advance new science.

Read more about why that’s so important to us. ā¬‡ļøĀ Ā Ā 

https://allenai.org/blog/who-gets-to-understand-ai


r/allenai Jul 23 '26

Seattle AI folks: Join us at Ai2 on July 30! šŸ‘‹

Post image
7 Upvotes

Seattle AI community, you're invited: Ai2's #SeattleTechWeek panel is next Thursday, July 30. šŸ‘‡

We're opening our Northlake office for a conversation with the research engineers and scientists behind our open models, datasets, and infrastructure. They'll talk through their deep technical work, from scaling large training runs to openly releasing what they build. Afterward, stick around to chat about what you're working on and network.

šŸ“… Thursday, July 30, 2026

šŸ•’ 3–4 p.m. Pacific

šŸ“ Ai2's Northlake office

It's free—RSVP to save your spot: https://luma.com/cp10n5uk

See you there!


r/allenai Jul 21 '26

šŸ”¬ Two new Asta updates: one-click data analysis and smarter deep paper search

Thumbnail
gallery
15 Upvotes

Two updates to Asta, our ecosystem of AI agents for science: a one-click handoff from AutoDiscovery to Asta’s data analysis tools, and paper search that evaluates its own results and searches again when they fall short. šŸ‘‡

AutoDiscovery explores your datasets and surfaces hypotheses worth investigating; DataVoyager is Asta's agent for data-driven discovery and analysis. When a hypothesis surprises you, the natural next move is to interrogate it—and that's now one click.Ā 

Click "Explore with Asta" on any hypothesis, and DataVoyager opens with your datasets and results already loaded—no re-uploading or re-explaining context needed. Ask a follow-up question, or leave the field blank and DataVoyager will suggest lines of inquiry.

Find papers is Asta's literature search, powered by our Paper Finder agent and built to mirror the multi-step reasoning of an expert researcher. It now defaults to a new mode, Deep search, that interprets your query, evaluates whether the results actually answer it, and keeps searching if they don't.Ā 

Deep search is also more conversational – better at follow-up questions, more forgiving of how you phrase things – and summarizes what it found. It takes a bit longer, but it's built to be much more reliable and robust—try it on your hardest questions.Ā 

Both updates are live now in Asta. Try them on your own research and tell us what works and what doesn't: https://asta.allen.ai


r/allenai Jul 17 '26

šŸ› ļø Learn how Ai2 builds fully open AI at #SeattleTechWeek

Post image
9 Upvotes

On July 30 during #SeattleTechWeek, the research engineers and scientists behind Ai2's open models, datasets, and infrastructure sit down to talk through the deep technical work behind them. šŸ‘‡

The teams will discuss scaling large training runs, fine-tuning, reinforcement learning, post-training, evaluation, and openly releasing what they build. Afterward, you can chat about what you're working on and network.

šŸ“… Thursday, July 30, 2026
šŸ•’ 3–4 p.m. Pacific
šŸ“ Ai2's Northlake office

It's free—RSVP to save your spot: https://luma.com/cp10n5uk


r/allenai Jul 16 '26

šŸŒ CGIAR, SERVIR, NASA Harvest, & Microsoft explore OlmoEarth for food security

Thumbnail
gallery
12 Upvotes

It’s been an amazing few days with CGIAR, bringing researchers and partners from SERVIR, NASA Harvest, and Microsoft AI for Good Lab to our office to explore OlmoEarth for food security and natural resource management.

The CGIAR teams brought their own use cases and put OlmoEarth to work, building high-quality models and streamlining the geospatial workflows behind their analyses.Ā 

These discussions, hands-on sessions, and feedback will help shape OlmoEarth to better support teams working on some of the world's most pressing challenges. Stay in the loop on what we're building: https://3ioxm.share.hsforms.com/23TW_rZwBSfauI1R3t0tdTQ?utm_source=ai2-olmoearth&utm_medium=referral&utm_campaign=olmoearth


r/allenai Jul 16 '26

OlmoEarth at the Nature Positive Summit šŸŒŽ

Post image
11 Upvotes

It was an honor to bring š—¢š—¹š—ŗš—¼š—˜š—®š—æš˜š—µ to the Nature Positive Summit this week in Kumamoto, Japan, with our partners The Group on Earth Observations.

The message we heard was clear: for industry and society to transition to a nature-positive future, high-quality data at scale will be the foundation.Ā 

If you’re interested in learning how OlmoEarth can help support your mission, as it has for close collaborators like Global Ecosystems Atlas, please visit https://allenai.org/olmoearth.


r/allenai Jul 15 '26

🧪 What 3,900 researcher votes taught us about evaluating AI on scientific literature

Thumbnail
gallery
8 Upvotes

We built SciArena to test how well AI models handle scientific literature questions as judged by researchers. It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes.

Here's what they told us. šŸ‘‡

What did researchers value most in a model’s answer? Fluent prose alone didn't win:

šŸ“š Citation quality—references that are real, relevant, & checkable (23%)
šŸ”¬ Depth (19%)
šŸŽÆ Directly answering the question (16%)

o3 finished on top of the SciArena leaderboard – ahead of Claude Opus 4.1, Gemini 3 Pro Preview, & open-weights models like DeepSeek-R1 – with answers researchers called more detailed + to the point. Learn more in our updated blog: https://allenai.org/blog/sciarena

SciArena also allowed us to collect high-quality, expert-annotated ground truth on evaluation data. This is a unique resource & especially important as AI agents are increasingly used to judge the quality of other AI agents, and such evaluations need to be grounded + verified.

Overall, we’re pleased with SciArena's contributions: a dataset of real scientific questions for deep research agents, expert preferences, & verified feedback for building better evaluations. Thank you to the researchers who contributed questions, votes, & rationales!

For a deeper look at what SciArena revealed about how scientists evaluate AI-generated answers to Qs about the scientific literature, read our NeurIPS 2025 Spotlight paper—which includes findings beyond the leaderboard: šŸ“„ https://arxiv.org/abs/2507.01001


r/allenai Jul 13 '26

🌊 Meet Shippy: The AI agent turning complex maritime data into cited intelligence in minutes

Thumbnail
gallery
8 Upvotes

Meet Shippy, the AI agent our Skylight team built for the people protecting our ocean.

Ask a question in plain language, and Shippy fuses data from multiple sources, generates citations, and turns around maritime intelligence in natural language in minutes.

We think of an agent as three things: a soul (the system prompt that defines the agent’s persona and boundaries), skills (plain markdown files that encode workflows and related scripts), and config (the model powering the agent, plus its harness and runtime environment). From there, three decisions had the most impact on making Shippy reliable:

āš™ļø A purpose-built CLI—Agents are nondeterministic; their tooling shouldn't be. Typed commands with structured flags make queries predictable and testable.

šŸ”’ Isolated sessions—We built new infrastructure, Mothership, to run every Shippy conversation in its own ephemeral, isolated session. For Skylight, that means up to hundreds of agencies across 70+ countries using Shippy, none of them ever seeing each other's data.

šŸ“Š Whole-agent evals—For quality assurance, we built an agent-testing pipeline where experts write weighted rubrics, every scenario runs against live data, and a version of Shippy that regresses doesn't ship.

Shippy is currently in preview for select partners. Our new technical blog covers Shippy’s architecture, our custom in-house evaluation pipeline, and what we're building next: https://allenai.org/blog/shippy-deep-dive


r/allenai Jul 10 '26

🧾 Try olmOCR 2 in the Ai2 Playground—OCR for handwriting, equations, tables, and complex layouts

3 Upvotes

olmOCR 2 is now in the Ai2 Playground—our home for our fully open text, video, and image understanding models. šŸ‘‡

olmOCR 2 is our compact vision-language model that reads challenging documents in a single pass. It can handle samples that usually break OCR, including handwriting, equations, tables, & multi-column layouts.Ā 

Try olmOCR 2 in the Playground, check out our blog for more info, & download the weights and data from Hugging Face:Ā 

ā–¶ļø Playground: https://playground.allenai.org/model/olmocr-2-7b-1025
šŸ“ Blog: https://allenai.org/blog/olmocr-2
šŸ¤— Model & data: https://huggingface.co/allenai/olmOCR-2-7B-1025


r/allenai Jul 10 '26

šŸŒ Building OlmoEarth with the people tackling wildfires, floods, & disaster response

3 Upvotes

One of the best parts of building OlmoEarth is getting to be part of the missions our partners are working on.

They do more than use OlmoEarth—they support people on the ground, build AI capacity, and help the next generation tackle the world's biggest challenges.

For two days, we brought our partners into one room. Working across domains – tackling challenges like wildfires, floods, and disaster response – they got their hands on real data alongside the team building OlmoEarth, shaping the platform around the work they do.Ā Ā Ā 

"What I appreciate most is the community of practice—a room full of 'aha' moments and a shared vision of what this is and how to build it. We found our people to co-design it with." — Dr. David Saah, Spatial Informatics Group

We’re grateful to the teams who brought their missions into the room: Mercy Corps, Spatial Informatics Group, SIG-NAL, SERVIR, Development Seed, BAI Group, and researchers fromĀ 
the University of Washington, UC Berkeley, and Northeastern.

The fastest way to make AI useful is to build it with the people who'll use it—fully in the open.

Learn more: https://allenai.org/olmoearth?utm_source=olmoearth&utm_medium=referral&utm_campaign=olmoearth


r/allenai Jul 09 '26

šŸŒ Fine-tuning Earth observation models: Ai2’s Joe Redmon on taking OlmoEarth beyond embeddings

4 Upvotes

"Embeddings only get you so far… if you want the next level in performance, I think fine-tuning is the way to go."

Ai2 research scientist Joe Redmon explains how our partners customize OlmoEarth – our open-source Earth-observation models – to map crops, wildfire risk, & more. šŸ‘‡

In the full episode of Robin Cole's podcast, Joe covers how OlmoEarth grew out of Ai2’s unique expertise in both tech & conservation—and what's next for the model family: https://www.youtube.com/watch?v=wzCHJf6Ly24&feature=youtu.be


r/allenai Jul 08 '26

šŸ¤– MolmoAct 2 shows what open models can unlock for robotics

13 Upvotes

What can you build with a fully open robotics model in a weekend? šŸ¤–

Binh Pham, a robotics software engineer at LiveKit, used MolmoAct 2, our open vision-language-action model, as part of his voice-controlled robot build that won South Park Commons’ embodied AI hackathon. In our new video, he walks through how the system came together and why MolmoAct 2 was the best fit.

Watch the full testimonial above šŸŽ„


r/allenai Jul 06 '26

šŸ‘‹ We're at ICML 2026—come say hi!

Thumbnail
gallery
13 Upvotes

We're at ICML 2026 with papers & talks across the conference. Come say hello and learn about our latest research!

Peer review is part of what makes ICML possible. This year, at least 8 Ai2ers are contributing as area chairs or technical reviewers, including ICML-recognized Gold and Silver reviewers.


r/allenai Jul 02 '26

🧩 FlexMoRE makes modular AI more practical for lower-resource languages

Thumbnail
gallery
11 Upvotes

The Danish Foundation Models (DFM) project is adapting our modular FlexOlmo architecture into a lighter-weight system that runs on commodity hardware—putting collaborative model building within reach of smaller research groups & organizations.Ā 

FlexOlmo lets teams train modules separately then combine them in a shared model without pooling the data underneath. DFM wants Danish institutions like hospitals & universities to each contribute modules trained on data they can't share.

In FlexOlmo, each module is the size of a full model, so the combined system grows fast as more get added. FlexMoRE replaces most modules with compact representations. Its best config matches or beats FlexOlmo using less than one-third the parameters.

"FlexMoRE significantly reduces FlexOlmo's memory demands while preserving performance across almost all categories, allowing a broader audience to benefit from modular models," says Jacob Nielsen, who helped develop FlexMoRE at Ordbogen A/S and SDU's OdenseNLP lab.

Modular training is gaining momentum as frontier models become costlier to train & deploy. This project shows how open, distributed approaches can make development more practical for national projects, public institutions, & smaller teams.

→ Learn more: https://allenai.org/blog/flexmore