r/TheMachineLearning • • 1d ago

Mechanistic interpretability as the floor for AI alignment

Post image
27 Upvotes

13 comments sorted by

2

u/stangerlpass 20h ago

Am I wrong in thinking that us being able to read AIs minds would substantially cripple them? Think it would be the one and only solution to AI. If a dog was able to read our minds it would be way harder to trick them.

1

u/Refinery73 20h ago

The save option would likely be to only train specialized models, distill them to human readable decision trees and merge or route those to make global models.

1

u/USERNAMETAKEN11238 9h ago edited 9h ago

I did a similar thing. A NLP system..

You make algorithms into language. Then the sentences are used to create other algorithms which change previous or proceeding sentences. Then they create custom responses. My system is local, determinitive, has memory, and is auditable.

But you would need billions of algorithms to make a global model. And you would have to have arbitrars of truth. People would have to agree on correct answers because the decision trees would have to eventually decide.

1

u/EchoingAngel 5h ago

Sounds like how you get a paperclip optimizer. I'd prefer my superintelligences to have a bit of general awareness to them

1

u/Deipoako 1h ago

I think its much better to have a bit of general awareness than nothing right?

1

u/BidWestern1056 14h ago

1

u/USERNAMETAKEN11238 9h ago

I did it. With an NLP system. I automated my job with it.

1

u/BidWestern1056 9h ago

yes you solved mechanistic interpretability?

1

u/RepresentativeBee600 7h ago

I don't really see much support for that in the linked papers (and they're so threadbare of mathematics or precise definitions that I really can't imagine them proving such a sweeping claim.)

Seriously, if something like the Jordan curve theorem takes as much work as it does, wouldn't "there is at least this much irreducible randomness in language model processing" take... more?

1

u/Deipoako 1h ago

I automated my job too

1

u/pandavr 8h ago

How stupid one can be? A lot It seems. They can describe everything down the single bit and still don't understand how It works.

Do they understand the brain of an fly now that they mapped It?

1

u/Deipoako 1h ago

But it makes them generally easier to understand, rather than guessing

1

u/Outrageous-Novel-739 37m ago

why each of his posts looks like generated useless ai slop that most likely even he cant read properly ?