AI Alignment Lacks Human Compliance Mechanisms
TL;DR. A new framework suggests AI alignment efforts focus too heavily on enforcement, overlooking layered human compliance evolved over millennia. - Current AI safety models often mirror law enforcement, neglecting social and internalized moral systems inherent in human societies. - The research maps human compliance layers onto AI, identifying critical missing areas in present AI alignment strategies. - Integrating concepts like continuous moral reinforcement and narrative curriculum could build more resilient AI systems.
- Human compliance systems are multi-layered, combining internalization, social pressure, institutions, markets, and enforcement.
- AI alignment currently overinvests in 'law enforcement' and 'constitutional principles' equivalents.
- Critical missing areas in AI alignment include longitudinal moral memory, reputation mechanisms, and restorative correction loops.
- A continuous, guilt-based alignment approach is proposed over static, shame-based methods for more robust AI safety.