Textbook on reinforcement learning from human feedback
Latest release code/v0.4 — Code v0.4 - Instruction Tuning, On-Policy Distillation, Config-Driven RMs · 7 Aug 2026
These files are natolambert/rlhf-book's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 4 tok CLAUDE.md A 1,664 tok .claude/skills/gemini-feedback/SKILL.md A 18 tok .claude/skills/pre-submit-pr/SKILL.md A 9 tok .claude/skills/push-to-pr/SKILL.md A 20 tok .claude/skills/run-rlhf-code-experiment/SKILL.md A 22 tok .claude/skills/serve-course-lecture/SKILL.md A 43 tok .claude/skills/update-pr-body/SKILL.md A 31 tok