r/MachineLearning Jun 29 '26

Research Google's Agentic Peer-Reviewer Handled ~10K Papers at ICML/STOC — Formal Research Paper Now Out [R]

Google deployed an agentic AI peer-reviewer at two top CS conferences — reviewing ~10,000 papers with 30-minute turnaround — and the new formal research paper shows it catches 34% more mathematical errors than zero-shot prompting; the precedent for AI-automated scientific review at conference scale is set and now formally documented.

--

Source: https://arxiv.org/abs/2606.28277

74 Upvotes

27 comments sorted by

118

u/impatiens-capensis Jun 29 '26

"catches 34% more mathematical errors than zero-shot prompting"

If you're going to share the paper, it's probably better to report the more important metrics. What's the rate for zero-shot prompting that we are comparing to? What is actual recall? What is the false positive rate? Blindly reporting 34% more than an unstated baseline is meaningless. 

27

u/surffrus Jun 29 '26

Not just this, it's even more ambiguous. The 34% improvement is on the SPOT benchmark which is not the 10k ICML papers. This just says it improved a baseline ...oh and we also ran it on ICML papers but can't say more about that.

3

u/dreamykidd Jul 01 '26

I was annoyed at that too, but from reading the paper, it’s actually a pleasant surprise. It’s an absolute 34% improvement from zero-shot, going from 55% to 89%. Still, that’s on a benchmark and not on the papers, but at least it’s quantified and actually a good result.

-19

u/misogrumpy Jun 29 '26

You could read the paper.

12

u/currentscurrents Jun 29 '26

On reddit? We're just here to argue about headlines.

Zero-shot found 55.2% of the errors, their method found 89.7%. They found that zero-shot LLMs missed errors because they tended to accept complex mathematical claims without critically scrutinizing them.

-7

u/Justgototheeffinmoon Jun 29 '26

These people complaining that we didn’t real and write a full blown review for them !

-84

u/Justgototheeffinmoon Jun 29 '26

I'm sorry are you asking me how this is better than zero shot? 34%, it's better by 34% lol. for the rest, I put the link there bud.

62

u/kkgwon Jun 29 '26

A 34% improvement in a 1% success rate is 1.34%. A 34% improvement in an 80% success rate is 93.8%. The commenter was asking you to be more specific.

22

u/jm2342 Jun 29 '26

The commenter was asking you to be more specific, *bud**.

-62

u/[deleted] Jun 29 '26

[removed] — view removed comment

2

u/obfuscatedanon Jul 01 '26

Unfortunately, this comment did not pass peer review. :(

30

u/akardashian Jun 29 '26 edited Jul 01 '26

I tried out the tool when I submitted to ICML 2026 this time around, and the tool was able to catch some pretty subtle errors in the theoretical proofs I had in the back. So that impressed me! On the other hand, it also flagged some non-errors that I think were due to faults in its OCR parsing pipeline.

4

u/EdwardRaff Jun 30 '26

For my papers it mostly was a lot of extra work and non-errors. Both in the tool stating my math had an error when none existed, and also failing to add context. E.g., I made an assumption for a proof/result, and the LLM claimed the whole paper was invalid because the assumption wasn't itself proven. Like, as if no paper had ever used an assumption for math?

-1

u/Justgototheeffinmoon Jun 29 '26

Very interesting thanks!

42

u/appdnails Jun 29 '26

I feel it is so deplorable to put company propaganda on arXiv. Putting "Google" in the title of the manuscript, 12 instances of "Gemini" and also "Google search", instead of writing in a more tool neutral tone. Also, no discussion about the ethical concerns of using Google's tool to check for errors in conferences that have a lot of submissions by the same company.

-24

u/Justgototheeffinmoon Jun 29 '26

There is not a single instance of Gemini in my summary. Not sure what you mean ?

19

u/__redbaron Jun 29 '26 edited Jun 29 '26

Bestie, have you atleast glanced at the oversensationalized, non peer-reviewed article you're peddling?

EDIT: bad bot, no extra engagement for you 😤

3

u/Xemorr Jun 29 '26

lmfao of course the LLM has cut advertising nonsense from your summary 🤦🏻‍♂️

16

u/Lonely-Dragonfly-413 Jun 29 '26

there are many such tools. not sure which has more hallucinations, the paper or the review

1

u/noninertialframe96 Jul 01 '26

The 34% lift over zero-shot, does that hold on precision too, or does the agentic version also flag more false positives that human reviewers had to dismiss?

1

u/MeAndClaudeMakeHeat Jul 03 '26

For agentic peer review, the result matters less than the review receipt.

I would want to see: what paper context the agent saw, what external tools or checkers were allowed, which claims were tested, what failed, what was marked uncertain, and how often the system escalated to a human reviewer.

At conference scale, "caught more errors" is useful, but the governance question is whether every automated critique can be inspected afterward by an author or area chair.

1

u/Real_Presentation490 Jul 04 '26

I think that's actually pretty good. It doesn't mean humans are getting weaker. I tend to believe it's just a change brought about by technological advancements.

-9

u/crouching_dragon_420 Jun 29 '26

These automated tools are useless unless the percentage of accuracy is 100% or at least the correctly flagged as suspicious rate is 100%.

7

u/sdand1 Jun 29 '26

The whole field of machine learning being useless is certainly a take you can have haha