Bai, 2022 — Constitutional {AI}: Harmlessness from {AI} Feedback

Rule-based critiques and AI feedback for alignment training; shares RLHF pointing and legitimacy cruxes.

Publication links

Rule-based critiques and AI feedback for alignment training; shares RLHF pointing and legitimacy cruxes.

Chapter-grouped bibliography All reference cards