0bserver07/Study-Reinforcement-Learning

RL study guide — foundations through RLHF, DPO, GRPO, RLVR, agentic RL, and offline RL. Hand-written CS294 notes, 19 lecture drafts, 5 tested exercises, citations that resolve.

These files are 0bserver07/Study-Reinforcement-Learning's own configuration. They tell Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

163Stars on the repository
2Files it configures its agents with
2,220Tokens loaded in every session
3Agents configured

Instructions