Reward Modeling skills

6 tagged Reward Modeling, measured the same way as everything else here.

Browse within: DPO 6GRPO 6PPO 6Post-Training 6RLHF 6Reasoning 6Structured Output 6TRL 6reinforcement-learning 6

paper-scout

01

AaronCIH/Awesome-AutoSkill-AutoRubric

Skill Claude Code

Daily paper scout for Auto-Skill and Auto-Rubric research. Use when: searching for new papers on self-evolving agents, skill evolution, rubric learning, preference alignment, reward modeling, agentic evolution. Searches arxiv for latest papers, recommends noteworthy ones, and updates the Awesome-AutoSkill-AutoRubric…

not rated 7 3mo ago A 72 tokens original MIT