regression testing skills

21 tagged regression testing, measured the same way as everything else here.

Browse within: annotations 12feedback-loop 12hypothesis-testing 12llmops 12systematic-evaluation 12

docs-design-tokens

01

rhesis-ai/rhesis

Skill Claude CodeCodex

Colour, font and design-token rules for the docs site — when hex is banned, when it is correct, and which surfaces ignore the theme. Use when styling docs components or editing CSS under docs/src.

391 4d ago A 46 tokens

worktree

02

rhesis-ai/rhesis

Skill Claude CodeCodex

Manage git worktrees with their own dev ports, symlinked .env files, playground, and simulations. Use when the user asks to create, set up, list, enter, or remove a worktree.

391 4d ago A 46 tokens

rhesis

03

rhesis-ai/rhesis

Skill Claude CodeCodex

Design, run, and analyze AI test suites on Rhesis — explore endpoints, build test foundations from a spec, create requirements and metrics, execute tests, and analyze results. Use when testing an AI endpoint, pasting a Product Requirements Document (PRD) or product spec, or working with Rhesis via MCP.

391 4d ago A 67 tokens

gclean

04

wookiya1364/scv-claude-code

Skill Claude CodeCodex

A procedure for removing local Git branches whose matching remote branches have been deleted, often after a change has been merged. Git branches are separate lines of work in a repository.

7 4d ago A 59 tokens original MIT

resilireplay

05

aliengineering-byte/resilireplay

Skill Claude CodeCodex

Capture a supported coding-agent tool failure as bounded, sanitized evidence and generate an executable deterministic regression. Use when a user asks to capture, explain, reproduce, or prevent recurrence of a Claude Code, Codex, Hermes, or MCP tool failure, or to validate a ResiliReplay adapter or campaign.

2 4d ago A 65 tokens original Apache-2.0

aoa-evals-skills

06

8Dionysus/aoa-evals

Skill Claude CodeCodex

Route the aoa-evals skill family for central proof selection, review, evolution, named results or verdicts, source-linked reports, Eval Forge owner review, and proof lifecycle. Hand repository-local eval selection, application, intake/design, or session-hit classification to aoa-eval. Candidates, readiness checks…

2 2d ago A 84 tokens original Apache-2.0

capture-failure

09

jiangkoumo/agenttape

Skill Claude CodeCodex

Turn a captured Codex tool failure into reviewed .tape evidence and an offline regression test. Use when the user asks to inspect a failed run, fork captured evidence, save a regression, or prepare AgentTape evidence for CI.

1 2d ago A 50 tokens original MIT

repoimmune

10

Alex0AI/RepoImmune

Skill Claude CodeCodex

Query evidence-backed historical bugs before risky code changes and verify patches before completion.

0 12d ago A 18 tokens original Apache-2.0