Bai, 2022 — Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Bootstraps alignment from smaller aligned models; inherits RLHF-style optimization-pressure limits.

Publication links

Bootstraps alignment from smaller aligned models; inherits RLHF-style optimization-pressure limits.

Chapter-grouped bibliography All reference cards