Assessing skeptical views of interpretability research
Interpretability vs. Explainable AI (XAI) • A Look Inside the Toolbox: Attribution, Probes & Interventions • The Theory: Understanding Models Through Causal Abs
Video Chapters
- 1:45 Interpretability vs. Explainable AI (XAI)
- 5:16 A Look Inside the Toolbox: Attribution, Probes & Interventions
- 14:14 The Theory: Understanding Models Through Causal Abstraction
- 15:27 The 5 Skeptical Arguments Against Interpretability
- 18:32 Is It Just Analysis Without Real Improvements?
- 25:24 A Cautionary Tale for AI Researchers
- 27:33 The Counter-Argument: When Interpretability Actually Works
- 31:55 Case Study: Tackling the Sycophancy Problem
- 35:56 An Optimistic Future for Interpretability