rlhf skills

38 tagged rlhf, measured the same way as everything else here.

Browse within: Alignment 26Evaluation 18reward 18reward-model 18GRPO 11Post-Training 11PPO 10DPO 9reinforcement-learning 8TRL 7Distributed Training 5

paper-page-figure

03

natolambert/rlhf-book

Skill Claude CodeCodex

Turn the first page of a paper (arXiv or any PDF) into a slide-ready PNG for a colloquium deck's 30-40% right column. Use when a slide cites a landmark paper and a screenshot of the paper itself is the best visual.

2.3k 11d ago A 60 tokens

qa-video-timestamps

04

natolambert/rlhf-book

Skill Claude CodeCodex

Extract per-question (MM:SS) timestamps from a recorded course video by OCR-ing the slide counter, then add YouTube deep links to the course page's Q&A question lists. Use after a Q&A or lecture recording is published, or when adding/refreshing "Show questions" timestamp links in book/templates/course.html.

2.3k 11d ago A 71 tokens

claude-authenticity

05

agentscope-ai/OpenJudge

Skill Claude CodeCodex

Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project. Also extracts injected system prompts from providers that override Claude's identity. Fully self-contained — copy the code below and run, no extra…

807 29d ago A 121 tokens original Apache-2.0

metric-design

06

agentscope-ai/OpenJudge

Skill Claude CodeCodex

Use when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining multiple metrics into a composite score, or building an automated evaluation pipeline. Also use when the user mentions grader selection, metric…

807 29d ago A 88 tokens original Apache-2.0

align-human

07

agentscope-ai/OpenJudge

Skill Claude CodeCodex

Use when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic evaluation can replace human review, or build a human-reduction roadmap. Also use when the user mentions calibration, TPR/TNR, judge validation…

807 29d ago A 100 tokens original Apache-2.0

grpo-rl-training

08

liortesta/ClawdAgent

Skill Claude CodeCodex

Expert guidance for GRPO/RL fine-tuning with TRL for reasoning and task-specific model training.

11 5d ago A 26 tokens copy · 100% Apache-2.0