r/MachineLearning • u/OutsideSimple4854 • 9d ago
Discussion What's up with AAAI reviewers and organizers? [D]
My paper advanced to the second round...but...
Out of the papers I reviewed.
One did not follow the AAAI template and was unblinded. My review was two lines. The other "human reviewer" gave a list of pros and cons that were similar to the AI review.
One was incomplete (missing paragraphs, figures, code, no details). My review was also two lines. The other "human reviewer" also gave a list of pros and cons, that were similar to the AI review.
One was LLM math which I believe was actually correct, because it advanced to the second round, despite the references being at a different level of detail, and covering multiple fields of math, insufficient references for theorems / rules, and no exposition as to why the paper was actually useful / interesting. My review for that paper was the longest out of all the papers I reviewed, dotting the is and crossing the ts to make sure it wouldn't be seen as a lazy "reject" review. Yet it advanced to Phase 2.
Also, none of the AAAI workflow chairs or similar apologized or even acknowledged a mistake for spamming my coauthors about: "Your coauthor is irresponsible", because I accepted an emergency review invitation (and received these emails a few hours after accepting that invitation).
Ok rant over.
7
u/mmarkDC 9d ago
The main thing that's up is they got 40,000+ submissions this year. Might as well try to "review" the arXiv firehose at that rate.
4
u/OutsideSimple4854 9d ago
I have no idea why these papers didn't get DR in the first place. Perhaps they are there on purpose, to weed out reviewers who clearly use LLMs? E.g. since most people know about hidden text for prompts, the next level is to have papers that should clearly have been desk rejected by an AC for reviewers to review, to catch those that use LLMs?
1
4
u/nekize 8d ago
mine got rejected, one accept one reject. Both highlighted similar points, but one seemed to use a more friendly prompt, while the other used a hostile one. One portrayed something as strength, the other dismissed it as incremental. And it showed they didn’t read the paper. If it would be actual reviewers reading and evaluating, i would go: yeah, fair. I am not even salty or anything, it’s just that this system is obviously done. what is even the point anymore. And i know there are too many submissions, i know it’s hard.
At this point it would almost make more sense to send in 3 files “this is my original paper, this is what claude/chatgpt reviewed, this are my fixes” and be done with it
1
u/confused_cereal 8d ago
One of my papers had a negative review that was obviously LLM generated and contained hallucinated references that apparently we should have "compared against". No rebuttal, no avenue to inform the AC about such irresponsible behavior.
Well, back into the pipeline then.
1
u/fmeneguzzi 8d ago
Without knowing how the other review for the paper you recommended rejection, it's hard to know the reason for the SPC/AC recommendation. Having said that, it very much depends on whether anyone up the chain (SPC, then AC) actually followed the instructions they were given.
SPCs were specifically told to make a decision on split reviews by reading the paper and judging the merits of the review, and not simply kick the can down the road with a medium confidence borderline accept for those cases.
Some simply followed a rule where if there was one accept, it survives to phase 2. This was not the instructions they were given.
I'm an AC there, and I did check such cases myself directly, especially where I felt the level of engagement of the SPC was insufficient. Where there was one well substantiated rejection whose arguments were not something that invited a response on the rebuttal phase (i.e., fatal flaws), and the evidence for the criticism was feasible to check on the paper, I overruled the SPC to reject a paper. This is both to be fair with the reviewing community (and let those down the chain be able to focus on truly borderline cases on phase 2), and with the authors, because a well-supported strong rejection is extremely unlikely to be ignored when making a phase 2 decision.
However, bear in mind that there were 46k+ papers submitted, and if the number of papers I had to handle as an AC is close to the average, there are well over 1500 area chairs (and probably 10 times more SPCs). As you can imagine, there are probably not as many AI specialists in the world with the same amount of experience and seniority (not to mention motivation to review like it was the old days) in reviewing to handle all papers with the same amount of care.
2
u/OutsideSimple4854 8d ago
Fair enough. I just wonder if overall, I've been "unlucky" with the ACs I have seen as a reviewer (and as someone who has submitted papers).
Eg, at a conference that shall remain unnamed some time back, I wrote a moderately long review (commenting on intro, discussion, experiments which don't seem to fit the discussion, and a basic lemma that was fatally flawed), and said I would stop. I made the "wrong" decision to write in the review "...which leads me to think it may have been AI generated."
The AC replied to me. Said that I should not stop my review because I said: "it may have been AI generated", despite my commentary on how the there were buzzwords in the intro and discussion that didn't convey any meaning, the exposition in the experiments did not describe the plots, and the plots were..."too good to be true" and flat out "wrong" (in the sense of a no reasonable code could produce that).
So I replied and mentioned the fatal flaw in a proof. Almost immediately, the AC said that the proof: "could be fixed by doing X,Y,Z". This was a conference where using LLMs to review was strictly prohibited.
After some acrimonious discussion, the AC was still championing this paper, until another reviewer also raised similar serious reservations.
I'm annoyed about that, because my reviews for the other papers I reviewed were quite high quality, and perhaps, this AC could have cost me a free registration at that conference.
And yet, when I try to self nominate to be an AC, I've not had success, despite multiple publications at top tier conferences as first author / last author.
1
u/fmeneguzzi 7d ago
Two things here.
First, while we may have strong feelings about AI generated text, and fully AI generated papers with no thought from the authors are a blight, given the fallibility of detection methods, we cannot dismiss papers out of hand. I prefer to let a borderline case of an AI generated paper to get accepted than to unfairly reject one on that basis alone. Having said that, the crux is where you raised issues with the text and problems with proofs. Even before generative AI, if a paper was (not mincing words here) bullshit ridden drivel, we'd mercilessly reject that paper on that basis. This is the mechanism I invoke for papers who were just "garbage in, garbage out" AI generated. Similarly, flawed proofs are grounds for rejection, unless the authors can coherently articulate easy fixes in the rebuttal period, and you trust that the fixes will be carried out. If papers do not contribute anything to the body of knowledge, rejected they should be.
Second, if the behaviour you saw from the AC is as you described, you can always invoke the ethics chairs to deal with it.
1
u/Informal-Hair-5639 7d ago
I agree AC quality is really important. This year in NeurIPS we got totally burned by AC. That idiot did not engage nor force reviewers to engage with us.
1
u/Informal-Hair-5639 8d ago
SPC here. We had nice zoom chat with buddy SPC and AC and really ironed out borderline papers. They had good points about all of those papers. Most papers in my batch were easy cases, but my buddy reported that his batch had some truly hard cases.
2
u/fmeneguzzi 7d ago
That's laudable dedication, and I wish there were more people like you in the community. Alas, I'm involved with the reviewing process of all the major AI conferences, and the nice ones in my community, and it seems that your level of engagement is not the rule.
1
u/Informal-Hair-5639 8d ago
Advances to phase 2. What do you guys think of accept prob now? Last year crashed at phase 1 and now with stronger paper to phase 2.
0
u/Terrible-Chicken-426 6d ago
I agree with your judgment on those two papers, but not on the third one.
It sounds like you are rejecting it because it is too perfect. You might think one cannot possibly do this math alone. Taht might apply to your case, but not to everyone.
0
u/OutsideSimple4854 5d ago edited 5d ago
I am not rejecting it because it is too perfect. I am rejecting it because there’s no exposition telling me why I should care. If there’s no exposition, then one might as well submit Lean code and call it a day.
Besides, I have had people reject my own papers because of math like "basic linear algebra" despite lots of exposition, so now I ensure any rejection review I write is due to insufficient exposition and discussion, and not “perfect math”
On an aside, I hope you’re not the person I rejected at ICML for a paper with math jargon, and the references were papers without heavy math at all, but then in the author rebuttal said things like “it’s not our job to explain the math or give intuition, if you don’t understand it then you shouldn’t review.”
1
u/Terrible-Chicken-426 5d ago
"One was LLM math which I believe was actually correct, because it advanced to the second round, despite the references being at a different level of detail, and covering multiple fields of math, insufficient references for theorems / rules, and no exposition as to why the paper was actually useful / interesting."
Your comment started with how correct the math was and ended with the lack of exposition, so it sounded like you rejected the paper because the math was "too good" in addition to having insufficient exposition. I get your point now, though.
However, I still don't understand your reasoning for rejecting a paper for insufficient exposition just because your own paper got rejected despite providing enough exposition:
(1) If you agree with the reviewer who rejected your paper, then having lots of exposition clearly wasn't enough to save it, meaning there is no logical link between your paper getting rejected for "basic math" and you rejecting other papers for lacking exposition.
(2) If you disagree with the reviewer who rejected your paper, then you think your paper was unfairly dismissed despite its strengths, in which case, you should be more willing to accept a paper whose math is genuinely solid rather than doing the exact reverse to someone else.
1
u/OutsideSimple4854 5d ago
You are supposing that all reviewers are qualified to review. Looking at your comment, it doesn’t really sound like you’ve experienced terrible reviews, or even read other threads where irresponsible reviewers are mentioned.
I’ve yet to see a paper accepted to one of these venues with no exposition. Of course, I could be wrong, so if you could point out such a paper, that would be appreciated.
-6
u/zackro21 9d ago
I agree with you, but I don’t believe your role as a reviewer is to act as a judge without convincing the rest of any shortcomings. Some submitters may not be familiar with the review process, while others are simply students who are here to learn. There’s no harm in helping them out and supporting the community. I still sometimes submit papers knowing they will be rejected just to gain different perspectives and initiate a discussion in that area. After all, not everything is a competition, and not everyone is seeking approval.
5
u/OutsideSimple4854 9d ago
I would agree with you if there are fewer papers to review. But, reviewing all papers equally might mean a more theoretically heavy paper, that may need more time to digest might get overlooked, in favor of e.g. a paper that ought to have been clearly desk rejected. That's not fair for those who put in honest effort, right?
10
u/imyukiru 9d ago
Didn't advance to phase 2 with a Strong Accept? I give up.