by Nathan Lambert
1 video mention tracked
This is the authoritative guide for Reinforcement learning from human feedback, alignment, and post-training LLMs. In this book, author Nathan Lambert blends diverse perspectives from fields like philosophy and economics with the core mathematics and computer science of RLHF to provide a practical guide you can use to apply RLHF to your models. Aligning AI models to human preferences helps them become safer, smarter, easier to use, and tuned to the exact style the creator desires. Reinforcement Learning from Human Feedback (RLHF) is the process for using human responses to a model’s output to shape its alignment, and therefore its behavior. In Reinforcement Learning from Human Feedback you’ll discover: • How today’s most advanced AI models are taught from human feedback • How large-scale preference data is collected and how to improve your data pipelines • A comprehensive overview with derivations and implementations for the core policy-gradient methods used to train AI models with reinforcement learning (RL) • Direct Preference Optimization (DPO), direct alignment algorithms, and simpler methods for preference finetuning • How RLHF methods led to the current reinforcement learning