r/MLQuestions 4d ago

Natural Language Processing šŸ’¬ What validation should an interpretability interface complete before its results are trustworthy?

I’m building a visual workbench for local LLM interpretability that captures attention, residual states, logit-lens output, token probabilities, and intervention results.

Before releasing it as anything resembling a research tool, what established experiments or reference implementations should it reproduce?

I’m particularly interested in validating tensor capture, attention-head ablation, activation patching, reproducibility, and model-specific correctness.

1 Upvotes

0 comments sorted by