AI Book Club: Reinforcement Learning from Human Feedback (RLHF)

AI Builders and Learners SF Novembers's book is "Reinforcement Learning from Human Feedback"! ​This is a casual-style event. Not a structured presentation on topics. Sometimes, the discussion even drifts away from the chapters, but feel free to grab the mic to help steer it back. ​Feel free to join the discussion even if you have not read the book chapters! :) ​Want to discuss the contents during the reading week? Join the Flyte MLOps Slack group. (https://flyte-org.slack.com/join/shared_invite/zt-47whsaemd-S4KFny5mk0uHCtRF2gU7ig?utm_source=luma#/shared-invite/email) ​------------------------------------------------- About the book: • ​Title: Reinforcement Learning from Human Feedback • ​Authors: Nathan Lambert • ​Published: August 2026 ​Manning (Promo code: AIBookClub should give you 45% off: https://www.manning.com/books/reinforcement-learning-from-human-feedback (https://www.manning.com/books/reinforcement-learning-from-human-feedback?utm_source=luma) ​O'rielly platform: https://learning.oreilly.com/library/view/reinforcement-learning-from/9781633434301/ (https://learning.oreilly.com/library/view/reinforcement-learning-from/9781633434301/?utm_source=luma) ​ ​Chapters: • ​1 Introduction • ​2 A tiny history of RLHF • ​3 Training overview • ​4 Instruction fine-tuning • ​5 Reward modeling • ​6 Reinforcement learning • ​7 Reasoning and inference-time scaling • ​8 Direct-alignment algorithms • ​9 Rejection sampling • ​10 The nature of preferences • ​11 Preference data • ​12 Synthetic data • ​13 Tool use and function calling • ​14 Over-optimization • ​15 Regularization • ​16 Evaluation • ​17 Crafting model character and products ​Book Description ​Reinforcement Learning from Human Feedback: LLM alignment and post-training helps you understand how modern AI models can be adapted to better match the needs and expectations of their users. Rather than surveying the vast field of reinforcement learning, elite AI researcher Nathan Lambert concentrates exclusively on RLHF and its immediate importance to post-training generative AI models. This compact book gets right to the point. Early chapters establish the training overview, explain instruction fine-tuning, and build reliable reward models. The middle chapters transition into the heart of alignment, exploring core policy gradient algorithms, Direct Preference Optimization (DPO), and inference-time scaling. Later chapters tackle the messy reality of data, guiding you through preference data collection, synthetic data generation, and the nuances of function calling.
Want more like this?

More in San Francisco

See all
8 Today October 2026
  1. 9 AM

    Oct 8 - MCP, Agents and Skills Meetup Meetup

    SF Machine Learning Meetup Join our virtual meetup to hear talks from experts on MCP, agents and skills. Date, Time and Location Oct 08, 2026 9:00 AM - 11:00 AM PST Online. Regi…

    Tech See website
  2. 3:29 PM

    Oakland Claude Code Hands-on Workshop

    Claude Code Enthusiasts Evening Social Club We are getting together in Oakland for a hands-on workshop focused on exploring Claude Code. Whether you have already started tinkering…

    Tech See website
  3. 4:30 PM

    AWS Partner Showcase (SV) - Elastic, MongoDB, Boomi, HiddenLayer

    SF/Bay AI Developers Group Important Note: register on the AICamp event website (https://www.aicamp.ai/event/eventdetails/W2026100817) is required for admission. RSVP on meetup is…

    Tech See website
  4. 6 PM

    ODSC AI Skills Accelerator | San Francisco

    ODSC AI San Francisco ODSC AI Skills Accelerator | San Francisco Event Details Date: October 8, 2026 Time: 6:00-8:00 PM PT Location: Creative Landing Spaces. SoMa Loft for Team …

    Tech See website