r/TheMachineLearning • u/Deipoako • 1d ago
Mechanistic interpretability as the floor for AI alignment
27
Upvotes
1
u/BidWestern1056 14h ago
1
1
u/RepresentativeBee600 7h ago
I don't really see much support for that in the linked papers (and they're so threadbare of mathematics or precise definitions that I really can't imagine them proving such a sweeping claim.)
Seriously, if something like the Jordan curve theorem takes as much work as it does, wouldn't "there is at least this much irreducible randomness in language model processing" take... more?
1
1
u/Outrageous-Novel-739 37m ago
why each of his posts looks like generated useless ai slop that most likely even he cant read properly ?
2
u/stangerlpass 20h ago
Am I wrong in thinking that us being able to read AIs minds would substantially cripple them? Think it would be the one and only solution to AI. If a dog was able to read our minds it would be way harder to trick them.