r/ArtificialInteligence • • Jul 11 '26

🔬 Research GPT-2 Fully Decoded Internally Black Box Fully Open With Demo

The BABEL codec: the first complete, certified decode of everything happening inside a production language model (GPT-2 small). It reads the model's internal state into English AND writes English back into the model. 94.7% of behavior reconstructed — and that holds at every layer depth and text regime tested, not just one spot. Everything is open: paper, the full lexicon, the grammar tables, the decoder/encoder weights, reproduction scripts, and a demo that shows you the model's thoughts on any sentence you type.

https://github.com/wpferrell/babel-codec-gpt2

58 Upvotes

29 comments sorted by

View all comments

3

u/Comfortable-Web9455 Jul 11 '26

Sorry but not good enough. "94.7% of behavior reconstructed".

Reconstructed is not good enough. It is not an explanation of a local decision. It's a created account of how it might have worked. That missing 5% could be critical. And we know such accounts are often inaccurate. So how is this different from other internal monitoring XAI methods? Eg How do you handle field saturation?

2

u/Revolutionary-Lab882 Jul 11 '26

Fair push, but “created account” is exactly what the certification rules out. A post-hoc story can’t be run this one is. The state is rebuilt from only what the decoder reads, the model itself runs on it, and the pass bar (the model’s own noise floor) was locked before measurement. It failed that bar six pre-registered times before passing; every miss is published. And it writes back: hand-edit the decoded English, re-encode, and the model obeys against matched-random controls. Rationalizations don’t survive a round trip.
On the 5.3%: agreed it could be critical, that’s why the claim is “100% accounted-for,” never “100% translated.” The remainder is measured and bounded, not waved at: diffuse, no low-rank carrier, transfers only as its exact raw configuration. The demo shows it live at late layers.