Preference Alignment skills

7 tagged Preference Alignment, measured the same way as everything else here.

Browse within: DPO 7GRPO 7HuggingFace 7PPO 7Post-Training 7RLHF 7SFT 7TRL 7fine-tuning 7reinforcement-learning 7

paper-scout

01

AaronCIH/Awesome-AutoSkill-AutoRubric

Skill Claude Code

Daily paper scout for Auto-Skill and Auto-Rubric research. Use when: searching for new papers on self-evolving agents, skill evolution, rubric learning, preference alignment, reward modeling, agentic evolution. Searches arxiv for latest papers, recommends noteworthy ones, and updates the Awesome-AutoSkill-AutoRubric…

not rated 7 3mo ago A 72 tokens original MIT