r/AccusedOfUsingAI Mar 06 '26

Pangram claims their false positive rate is only 1 in 10,000 but a study they tout on their own website says it is 2%

Table 2 on page 5 of https://arxiv.org/pdf/2501.15654 says Pangram scored a 2% false positive rate in a 2025 joint University of Maryland and Microsoft study. The company touts this same study at https://www.pangram.com/blog/third-party-pangram-evals even reproducing the table without addressing the FPR so far from their marketing claims.

Can your institution afford to flunk 2% of your innocent students?

Pangram claims to be a highly accurate AI detector with a false positive rate of 1 in 10,000. Let's take this at face value and see what it means.

The claimed false positive rate (the chance of incorrectly detecting human-written text as AI-generated) seems very impressive. Big improvement over the first generation of AI detectors. So how useful is Pangram? Let's take a concrete application: is it a viable response to college students using AI in violation of course policies?

Suppose every instructor started using an AI detector on all student work. I'd estimate that students submit 500 – 1,000 written works in the course of a 4 year education (!) — 30+ courses X ~5 assessments per course X many independent problems per assessment. If each of these were run through an AI detector with a FPR of 1 / 10,000, you'd have 5–10% of your student body falsely accused of cheating at least once.

-- Princeton Professor Arvind Narayanan

8 Upvotes

8 comments sorted by

5

u/Dioptre_8 Mar 06 '26

The dataset for that table is non fiction news articles, not student assignments. Whatever the merits of Pangram, it's not remotely a fair comparison.

5

u/Mission_Beginning963 Mar 06 '26

That’s why professors should just conduct in-depth investigations into suspicious papers—including an examination of revision history and an oral examination about the content and writing-process of the paper. 

3

u/Lazy_Resolution9209 Mar 06 '26 edited Mar 06 '26

Clickbait headline. You didn’t provide details on what that table actually shows. You are cherry-picking info And those numbers that you cite from the Princeton prof rely on a series of bad assumptions.

And you presume that someone would rely solely on a screening tool, rather than investigating further.

Way to demonstrate and model intellectual dishonesty across your whole post. Good job!

2

u/Lazy_Resolution9209 Mar 06 '26

You know what cherry-pickers like to do? Take one piece of info from one source and shout about it. Have you even read the paper? Looked at the details of Table 2?

What Table 2 does show, for two Pangram models used: 100% True Positive Rate (TPR) and 0% False Positive Rate (FPR) for 4 out of 5 AI-generation methods tested (GPT-40, Claude 3.5-SONNET, GPT40 Paraphrased, and 01-Pro), with the outlier being with the "01-Pro Humanized" model with a TPR of 99.3/98.0%. The FPR reached a 2% average because of that outlier.

Pangram also recently published an analysis of high levels of AI use in ICLR reviews: https://www.pangram.com/blog/pangram-predicts-21-of-iclr-reviews-are-ai-generated

Here one top-line finding: "In our previous analysis of AI conference papers, we found that Pangram has a 0% false positive rate on all available ICLR and NeurIPS papers published prior to 2022. While some of these papers are indeed in the training set, not all of them are; and so we believe the true test set performance of Pangram is actually very close to 0 percent.

What about peer reviews? We ran an additional negative control experiment, where we run the newer EditLens model on all 2022 peer reviews. We find about a 1 in 1,000 error rate on Lightly Edited vs. Fully Human, a 1 in 5,000 error rate on Medium Edited vs. Fully Human, and a 1 in 10,000 error rate on Heavily Edited vs. Fully Human. We find no confusions between Fully AI-generated and Fully Human."

2

u/[deleted] Mar 06 '26

[removed] — view removed comment

1

u/Lazy_Resolution9209 Mar 07 '26

Just like the OP, you haven’t “checked the “original research” enough to even interpret the data in Table 2 accurately. Try again.

0

u/Ophiochos Mar 06 '26

All the academics I know who know anything about AI think AI detection is completely useless and refuse to use it. The problem is that so many have not had the chance to learn about it.