BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//events//Events//EN
CALSCALE:GREGORIAN
X-WR-CALNAME:AI Book Club: Reinforcement Learning from Human Feedback (RLHF
 )
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-PUBLISHED-TTL:PT1H
BEGIN:VEVENT
UID:event-eehd3325@ontown.app
DTSTAMP:20261008T212223Z
DTSTART:20261110T210000Z
SUMMARY:AI Book Club: Reinforcement Learning from Human Feedback (RLHF)
LOCATION:San Francisco\, United States
DESCRIPTION:AI Builders and Learners SF\nNovembers's book is "Reinforcement
  Learning from Human Feedback"!\n\n​This is a casual-style event. Not a 
 structured presentation on topics. Sometimes\, the discussion even drifts 
 away from the chapters\, but feel free to grab the mic to help steer it ba
 ck.\n\n​Feel free to join the discussion even if you have not read the b
 ook chapters! :)\n​Want to discuss the contents during the reading week?
  Join the Flyte MLOps Slack group. (https://flyte-org.slack.com/join/share
 d_invite/zt-47whsaemd-S4KFny5mk0uHCtRF2gU7ig?utm_source=luma#/shared-invit
 e/email)\n​-------------------------------------------------\nAbout the 
 book:\n\n• ​Title: Reinforcement Learning from Human Feedback\n• ​
 Authors: Nathan Lambert\n• ​Published: August 2026\n\n​Manning (Prom
 o code: AIBookClub should give you 45% off: https://www.manning.com/books/
 reinforcement-learning-from-human-feedback (https://www.manning.com/books/
 reinforcement-learning-from-human-feedback?utm_source=luma)\n​O'rielly p
 latform: https://learning.oreilly.com/library/view/reinforcement-learning-
 from/9781633434301/ (https://learning.oreilly.com/library/view/reinforceme
 nt-learning-from/9781633434301/?utm_source=luma)\n​\n​Chapters:\n\n•
  ​1 Introduction\n• ​2 A tiny history of RLHF\n• ​3 Training ove
 rview\n• ​4 Instruction fine-tuning\n• ​5 Reward modeling\n• ​
 6 Reinforcement learning\n• ​7 Reasoning and inference-time scaling\n
 • ​8 Direct-alignment algorithms\n• ​9 Rejection sampling\n• ​
 10 The nature of preferences\n• ​11 Preference data\n• ​12 Synthet
 ic data\n• ​13 Tool use and function calling\n• ​14 Over-optimizat
 ion\n• ​15 Regularization\n• ​16 Evaluation\n• ​17 Crafting mo
 del character and products\n\n​Book Description\n​Reinforcement Learni
 ng from Human Feedback: LLM alignment and post-training helps you understa
 nd how modern AI models can be adapted to better match the needs and expec
 tations of their users. Rather than surveying the vast field of reinforcem
 ent learning\, elite AI researcher Nathan Lambert concentrates exclusively
  on RLHF and its immediate importance to post-training generative AI model
 s.\n\nThis compact book gets right to the point. Early chapters establish 
 the training overview\, explain instruction fine-tuning\, and build reliab
 le reward models. The middle chapters transition into the heart of alignme
 nt\, exploring core policy gradient algorithms\, Direct Preference Optimiz
 ation (DPO)\, and inference-time scaling. Later chapters tackle the messy 
 reality of data\, guiding you through preference data collection\, synthet
 ic data generation\, and the nuances of function calling.
URL:https://ontown.app/e/eehd3325-ai-book-club-reinforcement-learning-from-
 human-feedback-rlhf/
END:VEVENT
END:VCALENDAR
