19 Safe & Constrained Reinforcement Learning
How should an intelligent system pursue goals when some actions must never be acceptable?
Area 19: Safe & Constrained Reinforcement Learning
Master Question
How should an intelligent system pursue goals when some actions must never be acceptable?
What We Want to Discover
Maximizing reward is not the same as respecting a hard constraint. This area studies how safety constraints should be encoded so that they hold even during learning and exploration, not only at convergence, and how a system should behave when a constraint and the reward signal point in different directions.
Why It Matters
A system that is safe on average is not safe. Constraint violations during exploration, or in rare edge cases, can be catastrophic even if they are statistically infrequent. This area addresses the gap between average case performance and worst case guarantees.
Core Concepts
- Constrained Markov decision processes
- Safe exploration
- Shielding and runtime monitors
- Barrier functions
- Reward hacking and specification gaming
Relevant Disciplines
- Safety engineering
- Control theory
- Ethics and moral philosophy
- Law and regulation
Potential Mimicry Sources
- Regulatory constraints in aviation and medicine
- Instinctive avoidance behaviour in animals
- Engineered fail safes in industrial systems
Projects
No projects yet.
Active Questions
No projects yet, so there are no active derived questions to report.
Key Findings Across Projects
Pending. No projects in this area have produced findings yet.
Unresolved Questions
Pending.
Connections to Other Areas
Pending. Connections will be identified as projects in this area develop.