systematic evaluation skills

12 tagged systematic evaluation, measured the same way as everything else here.

Browse within: annotations 12feedback-loop 12hypothesis-testing 12llmops 12regression-testing 12

docs-design-tokens

01

rhesis-ai/rhesis

Skill Claude CodeCodex

Colour, font and design-token rules for the docs site — when hex is banned, when it is correct, and which surfaces ignore the theme. Use when styling docs components or editing CSS under docs/src.

391 3d ago A 46 tokens

worktree

02

rhesis-ai/rhesis

Skill Claude CodeCodex

Manage git worktrees with their own dev ports, symlinked .env files, playground, and simulations. Use when the user asks to create, set up, list, enter, or remove a worktree.

391 3d ago A 46 tokens

rhesis

03

rhesis-ai/rhesis

Skill Claude CodeCodex

Design, run, and analyze AI test suites on Rhesis — explore endpoints, build test foundations from a spec, create requirements and metrics, execute tests, and analyze results. Use when testing an AI endpoint, pasting a Product Requirements Document (PRD) or product spec, or working with Rhesis via MCP.

391 3d ago A 67 tokens