reward model skills

18 tagged reward model, measured the same way as everything else here.

Browse within: Alignment 18Evaluation 18RLHF 18reward 18

claude-authenticity

01

agentscope-ai/OpenJudge

Skill Claude CodeCodex

Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project. Also extracts injected system prompts from providers that override Claude's identity. Fully self-contained — copy the code below and run, no extra…

809 +2 29d ago A 121 tokens original Apache-2.0

metric-design

02

agentscope-ai/OpenJudge

Skill Claude CodeCodex

Use when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining multiple metrics into a composite score, or building an automated evaluation pipeline. Also use when the user mentions grader selection, metric…

809 +2 29d ago A 88 tokens original Apache-2.0

redteam

03

agentscope-ai/OpenJudge

Skill Claude CodeCodex

Use when the user wants to test their LLM/agent application for safety and security vulnerabilities — jailbreaks, prompt injection, PII extraction, harmful content generation, or evaluator gaming. Also use when the user mentions security testing, adversarial testing, red teaming, safety evaluation, ASR (Attack Success…

809 +2 29d ago A 89 tokens original Apache-2.0