04 Explainable Reinforcement Learning

What does it mean for an intelligent system to understand and explain its own decisions?

Area 04: Explainable Reinforcement Learning

Master Question

What does it mean for an intelligent system to understand and explain its own decisions?

What We Want to Discover

A policy that performs well is not the same as a policy whose behavior can be understood. This area studies what forms of explanation are meaningful for sequential decision making, whether an explanation is faithful to the actual computation the agent performed, and whether an explanation is useful to the person receiving it.

Why It Matters

Trust, debugging, certification, and human oversight all depend on some form of explanation. A system that cannot explain a harmful decision is difficult to correct and difficult to certify, regardless of how well it performs on average.

Core Concepts

  • Policy interpretability
  • Saliency and attention based explanation
  • Reward decomposition
  • Counterfactual explanation
  • Faithfulness versus plausibility of an explanation

Relevant Disciplines

  • Philosophy of explanation
  • Cognitive science
  • Human computer interaction
  • Safety engineering

Potential Mimicry Sources

  • How humans explain their own decisions after the fact
  • Expert reasoning in high stakes professions
  • Scientific explanation and causal reasoning

Projects

No projects yet.

Active Questions

No projects yet, so there are no active derived questions to report.

Key Findings Across Projects

Pending. No projects in this area have produced findings yet.

Unresolved Questions

Pending.

Connections to Other Areas

Pending. Connections will be identified as projects in this area develop.