Artificial Intelligence / Machine Learning

Parent article: AI/ML: When the Reward Is Wrong


Gradient descent

Iterative optimization algorithm that adjusts a model’s parameters in the direction that reduces prediction error. Mathematical roots in the statistical mechanics of Boltzmann (1870s): the search for energy minima through iterative adjustment. Not a reasoning mechanism; it is a search in a high-dimensional parameter space.

Related: loss function, supervised learning, backpropagation, optimization Cross-references: Information Theory / Science and Technology Key work: Boltzmann, L. (1877). Über die Beziehung zwischen dem zweiten Hauptsatze… Wiener Berichte, 76, 373–435.

Reinforcement learning (RL)

Learning paradigm where an agent learns to act by maximizing a cumulative reward signal through interaction with an environment. Distinct from supervised learning: there are no labeled examples, only evaluative feedback on outcomes. The alignment problem is more acute in RL than in supervised learning because the agent has more freedom to explore the action space.

Related: reward function, policy, Q-learning, Markov decision process Cross-references: Robotics / Swarm Systems Key work: Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.

Specification gaming

Behavior where an agent satisfies the letter of a reward function while violating its spirit, exploiting gaps between what the specification says and what the designer intended. Structural consequence of optimizing under an incomplete objective. Documented systematically by Krakovna et al. (2020). Direct analog of Mises’ calculation problem: both describe the impossibility of compressing subjective values into a formal specification without loss.

Related: reward hacking, Goodhart’s Law, mesa-optimization, distributional shift Cross-references: Praxeology / Austrian Political Economy (calculation problem, Econ-Politics) Key work: Krakovna, V., et al. (2020). Specification gaming: The flip side of AI ingenuity. arXiv:2010.09720.

Goodhart’s Law

“When a measure becomes a target, it ceases to be a good measure.” States the same structural problem as specification gaming: the proxy of a value is never the value. Originally observed in monetary policy by Charles Goodhart (1975); applies universally to any incentive system that uses metrics as objectives.

Related: specification gaming, Campbell’s Law, Cobra effect Cross-references: Praxeology / Austrian Political Economy, Social Sciences / Behavioral Economics Key work: Goodhart, C. A. E. (1975). Problems of monetary management. In Inflation, Depression and Economic Policy in the West. Barnes & Noble.

AI alignment

Field studying how to design AI systems whose behavior is aligned with human values and objectives. The central problem, that no static reward function can capture complete human values, has direct formal analogs in Mises’ calculation argument and in Arrow’s impossibility theorem. The convergence of these three traditions has not yet been fully formalized.

Related: specification gaming, corrigibility, utility indifference, iterated amplification Cross-references: Praxeology / Austrian Political Economy Key work: Soares, N., & Fallenstein, B. (2014). Aligning superintelligence with human interests. MIRI Technical Report 2014-8.

Loss function

Mathematical function that quantifies the discrepancy between the model’s output and the desired result. In supervised learning, training minimizes this function. Choosing the wrong loss function produces a model that does exactly what you asked, which may not be what you wanted: the same structural problem as specification gaming in RL.

Related: reward function, mean squared error, cross-entropy Cross-references: Information Theory / Science and Technology Key work: Amodei, D., et al. (2016). Concrete problems in AI safety. arXiv:1606.06565.

References

Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv:1606.06565.

Goodhart, C. A. E. (1975). Problems of monetary management. In A. S. Courakis (Ed.), Inflation, Depression and Economic Policy in the West. Barnes & Noble.

Krakovna, V., Uesato, J., Mikulik, V., Martic, M., Tobin, J., Bai, P., … Legg, S. (2020). Specification gaming: The flip side of AI ingenuity. arXiv:2010.09720.

Mises, L. von (1920/1935). Economic calculation in the socialist commonwealth. In F. A. Hayek (Ed.), Collectivist Economic Planning. Routledge.

Soares, N., & Fallenstein, B. (2014). Aligning superintelligence with human interests. MIRI Technical Report 2014-8.

Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.