04 Explainable Reinforcement Learning
What does it mean for an intelligent system to understand and explain its own decisions?
Area 04: Explainable Reinforcement Learning
Master Question
What does it mean for an intelligent system to understand and explain its own decisions?
What We Want to Discover
A policy that performs well is not the same as a policy whose behavior can be understood. This area studies what forms of explanation are meaningful for sequential decision making, whether an explanation is faithful to the actual computation the agent performed, and whether an explanation is useful to the person receiving it.
Why It Matters
Trust, debugging, certification, and human oversight all depend on some form of explanation. A system that cannot explain a harmful decision is difficult to correct and difficult to certify, regardless of how well it performs on average.
Core Concepts
- Policy interpretability
- Saliency and attention based explanation
- Reward decomposition
- Counterfactual explanation
- Faithfulness versus plausibility of an explanation
Relevant Disciplines
- Philosophy of explanation
- Cognitive science
- Human computer interaction
- Safety engineering
Potential Mimicry Sources
- How humans explain their own decisions after the fact
- Expert reasoning in high stakes professions
- Scientific explanation and causal reasoning
Projects
No projects yet.
Active Questions
No projects yet, so there are no active derived questions to report.
Key Findings Across Projects
Pending. No projects in this area have produced findings yet.
Unresolved Questions
Pending.
Connections to Other Areas
Pending. Connections will be identified as projects in this area develop.