natolambert/rlhf-book

Textbook on reinforcement learning from human feedback

Latest release code/v0.4 — Code v0.4 - Instruction Tuning, On-Policy Distillation, Config-Driven RMs · 7 Aug 2026

These files are natolambert/rlhf-book's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

2,368Stars on the repository
8Files it configures its agents with
1,668Tokens loaded in every session
3Agents configured

Instructions

Skills