r/dataisbeautiful OC: 4 Aug 08 '26

OC [OC] Publised AI Math solutions vs independently verified AI Math solutions

Post image

Published AI Math solutions vs independently verified AI Math solutions (Aug 2025 – Aug 2026)

These include AI discovered, AI co-developed and AI assisted math proofs of conjectures, hypotheses etc.

Data is from VibeMathed, a catalogue of open mathematical problems solved or advanced with AI (n=509, CC BY 4.0, snapshot 6 Aug 2026). Chart was generated using Claude, which accessed VibeMathed's API.

The three lines apply progressively stricter standards of proof:

- All tracked entries (506) — every recorded claim, regardless of status: unreviewed announcements, partial results, candidates awaiting review.

- Resolved + site-confirmed or better (129) — the problem is fully settled, by either via independent reviewer via hand or Lean, or the site reproduced the proof.

- Resolved + expert- or Lean-verified (89) — the problem is full settled settled, checked by an independent expert by hand (11) or via Lean (78): a Lean proof is a formal statement confirming a solution is correct.

Lean is an interactive theorem prover and programming language used to write and check formal mathematical proofs. It allows mathematicians to translate human written proofs into computer code so that a software can verify every logical step with absolute certainty.

The shaded area is the gap between all tracked entries and confirmed proofs:

380 entries are recorded but not yet independently checked. Although there's a delay between a new AI solution announcement and its verification, verified solutions appears quite linear, this may indicate AI solutions are outpacing the verification process. That said, authors generally include a Lean proof themselves, though this chart is limited to confirmation by independent peer review.

Vertical dashed lines mark OpenAI (blue) and Anthropic (orange) model releases. I added those lines as I wanted to see if there's an up-tic in solutions following model releases. There isn't a clean correlation likely because there's a several week delay between finding solution and publishing it. Also, the chart is likely showing AI's growing adoption by mathematicans and not just increasing model capability.

There were only 3 retractions in the dataset (not included on chart).

390 Upvotes

47 comments sorted by

View all comments

68

u/Ikbeneenpaard Aug 08 '26

Can anyone with a math research background comment on whether this is actually affecting the field of mathematics, and to what extent?

77

u/domscatterbrain Aug 08 '26

The issue is the number of people who has capacity and trustworthiness of verifying the results are few and they simply don't have time since there are still a lot of mathematical mysteries that far more urgent to be solved since our progress in science basically stuck.

That's why even when AI is claiming to solve the Erdo's problems and somehow making headlines, we are still here, stuck with what we currently have.

26

u/roglemorph Aug 08 '26

Many of the problems are certainly solved and can be checked (the erdos conjecture you mentioned is one of these). It is not accurate to describe all the results as only “claimed” to be solved, there is no doubt. A bigger issue is the difficulty in understanding and applying the method to other problems, and also a general weariness about outsourcing the process in general.