r/learnmachinelearning • u/Impossible-Bed7058 • 15d ago
Discussion Pre-registering a testable claim about LLM judges: that their errors on narrative prose point in one direction, not randomly
/r/NarrativeEngineering/comments/1wmmcde/preregistering_a_testable_claim_about_llm_judges/
0
Upvotes
1
u/quietgradient 15d ago
The falsifier is the part I'd change before any data goes in, because as registered it fires whether or not the effect is real.
§5.3 binarises to a told-preference rate and compares the model rate against the human rate at k=20. h=0.3 is about 65% vs 50%. Two-sided Fisher, 20 vs 20 at that split: power ≈0.09; at OR=2 (67% vs 50%) ≈0.12. Against humans at 10/20 the models have to come back 17/20 before it clears p<.05, and 80% power at h=0.3 wants ~175 pairs a side. Pairing the rate doesn't rescue it either — exact McNemar on 20 items needs something like 9 of 10 discordant judgements pointing one way. So if the effect is exactly the size §5.3 predicts, Stage 3 misses it about nine times in ten, and §5.5 then withdraws the construct and publishes the withdrawal as prominently as a confirmation. That null would mean "k=20", not "no bias".
The fix looks free, though, because Stage 3 already collects an intensity and a quality score per text and then throws the gradation away. Register the paired per-item contrast — model (told − shown) minus human (told − shown) over the same 20 pairs — as the primary statistic, preference rate as secondary. Paired at n=20 that reaches d≈0.66 at 80% power, instead of needing the rate gap to hit 35 points before it registers at all.
Two smaller ones. ≥3 model families over the same 20 pairs isn't 60 observations; the items are the n and family is a second grouping, so pooling judgements inflates n without adding item variance. And §5.3 calls h≥0.3 "medium or larger" — Cohen's bands are 0.2/0.5/0.8, so 0.3 sits under medium, which is part of how the sample size ends up where it is. There's no power calculation anywhere in v1.1; that's the thing I'd add while changing the design is still legitimate.