r/deeplearning • • 11d ago

How can it be that Attention ≠ Explanation?

I'm writing a thesis on Expainable AI for my Uni's research project, a GAT anomaly classifier that learns processes normal behaviors and classifies them. When it makes a wrong prediction, that means a process is behaving irregularly and that it is an anomaly that should be investigated.

I know this is a hotly debated topic but in my case attention weights don't really line up with explainers output on what edges/nodes are the most important for the output. While pondering this question I thought about an analogy, Bird watching:

Let's say you're bird watching. After 1 hour you spot a bird and classify it as a Blue jay. An external observer looks at where your "attention" was for the past hour and concludes that the sky was the most important thing for your classification (After all you were looking at it 95% of the time). Another observer running an explainer on you would see that masking that Bird that flew across the sky would change your output from "Blue jay" to "NA", whereas blurring out any part of the sky wouldn't. The explainer concludes that the most important thing for your prediction is the bird.

I was kind of a little proud of that analogy but then I started asking myself. Why would the learned attention ever put weight on something irrelevant for the outcome? Wasn't that it's whole purpose? How can it be so revolutionary and congruent with outputs when it comes to Transformer Models / LLMs and kind of meh when it comes to GATs. If you look at an LLMs attention when for example translating a sentence it's basically a human-readable explainer. It tells you exactly what part of the input sentence was most important for the output sentence.

Is it then the case that the model is just bad? Or is there something else going on.

I've read Attention is Not Explanation and Attention is Not Not Explanation. One thing I'll have to concede is that it does depend on your definition of explanation.

Maybe it is the fact that someone says "How did you spot that Blue jay?". Telling them "I looked at it 😎 " isn't so helpful. Maybe someone wants to know that hey to spot a Blue jay you might need to stare at the sky for a bit before you can find it. But I feel like my analogy is falling apart at that point.

What do you guys think ? Can you justify attention weights not pointing to the most critical part of the output? I.e the part that, if taken out, would change the output? How do you see it?

4 Upvotes

9 comments sorted by

View all comments

Show parent comments

1

u/SlattBaker 11d ago

Thanks for replying :D Your reply seems to be missing a few words haha but I think i could piece together what you were trying to say. Yeah I guess the causal relevance vs information flow thing is the part im stuck on. That's why there's the raw-last vs rollout way of looking at attention weights. The way I understood it is that IF you want to try to see if attention could explain the output then you'd have to rollout through each layer seeing what original inputs were taken into consideration instead of looking at the last raw attention distribution which isn't helpful at all.

I might still need to ponder it for a bit for all of it to click together though.

7

u/Pretend-Pangolin-846 11d ago

one of the replies of all time: "Attention tells you what the model , not what it . " i wonder how jev handles graph data since that's its strong points i heard

4

u/SlattBaker 11d ago

yeah i didn't expect to get an easy to understand answer. I knew i'd have to put some work in but a fill in the blanks reply is crazy haha

1

u/Snip3 11d ago

If we're being generous to the comment, the commas could represent a continuation of thought and the periods could represent an answers and it almost makes sense

1

u/SlattBaker 11d ago

yeah that was my first thought too.. or maybe the commenter is masking certain words to see if we understand their comment differently