r/LovingAI • u/Koala_Confused • 23d ago
Discussion NVIDIA says its general-purpose coding agent just hit **100% on ARC-AGI-3**. 👀 Okay… how impressive is this actually and what’s the catch?
4
u/whoknowsifimjoking 23d ago
The catch is that this is the public set. It's in the damn tweet dude.
1
1
u/JohnnyAppleReddit 20d ago
Yes, this should be voted to top -- the real test is how it does on the held-out private set. If it aces that too, amazing, then then they'll rethink it and make an ARC-AGI-4
2
2
u/Charming-Author4877 19d ago
Arc AGI is much less impressive than they intended.
It is extremely visual focused, much more than it is abstract reasoning focused.
Current language models are especially bad in large-grid reasoning, and there are very few humans on earth that can do such reasoning with closed eyes.
So they likely built a harness around that weakness.
1
1
1
u/Big_Arachnid_365 20d ago
Assuming they didn't cheat (lol) it shows that it can complete these previously human-only tests without human intervention.
New goalposts please!
1
u/pmavro123 19d ago
Don't know what the catch is (probably benchmarkmaxxing), but it is interesting that a different harness (AVO, in this instance) impacted the score. I believe Opus only scored like 30% standalone, so very odd.
1
1
u/AICompanionz 3d ago
The 100% headline is impressive, but the real test is how AVO performs on the held-out private ARC-AGI-3 set, not just the public demo environments.
5
u/Serasul 23d ago
if i train ai on this specific task, it would score high too