Use this skill when judging a single candidate response for instruction following and a scalar score is needed. It is especially useful when the sample contains exact constraints, a visible checklist, verifier resources, format requirements, word/count limits, language restrictions, or multiple sub-instructions that…
Use this skill when judging instruction following for IF-RewardBench-style samples. It is especially useful when the sample contains exact constraints, a visible checklist, format requirements, word/count limits, language restrictions, or multiple sub-instructions that should be decomposed before the final benchmark…
Use this Skill-RM reward judge to compare candidate responses for a visible user request with generic rubric, principles, bias controls, output contract, and Python sandbox checks over visible text.
Use this Skill-RM reward judge when response judging may benefit from resource-rich evidence: benchmark/task metadata, visible references or ground truth, checklists, verifier signals, code/math/factuality tool protocols, same-backbone judging pipeline signals, or bias-control resources. Load it when these resources…
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: