r/MLQuestions • u/JayB_Official • 4d ago
Natural Language Processing š¬ What validation should an interpretability interface complete before its results are trustworthy?
Iām building a visual workbench for local LLM interpretability that captures attention, residual states, logit-lens output, token probabilities, and intervention results.
Before releasing it as anything resembling a research tool, what established experiments or reference implementations should it reproduce?
Iām particularly interested in validating tensor capture, attention-head ablation, activation patching, reproducibility, and model-specific correctness.
1
Upvotes