SkillsBench evaluates how well skills work and how effective agents are at using them.
SkillsBench is a benchmark for measuring how effectively AI agents use modular skills—folders containing instructions, scripts, and resources—to complete specialized tasks. It helps researchers and developers evaluate both skill quality and agent behavior, including tasks that require combining multiple skills. The catalogue’s skills and instructions are evaluated as part of this workflow.
Latest release v1.1 — SkillsBench v1.1 — native task.md package release · 14 Jun 2026
These files are benchflow-ai/skillsbench's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 692 tok CLAUDE.md A 17 tok .agents/skills/skill-creator/SKILL.md A 46 tok .agents/skills/skillsbench/SKILL.md A 35 tok .agents/skills/task-creator/SKILL.md B 104 tok .agents/skills/task-review/SKILL.md B 141 tok