agent evals skills

30 tagged agent evals, measured the same way as everything else here.

Browse within: business-ontology 21ontology 21gemini 6

autoevolve

02

MrTsepa/autoevolve

Skill Claude CodeCodex

Iteratively improve code, strategies, or prompts through mutation, evaluation, and selection.

22 2mo ago A 20 tokens original MIT

beta-skill

04

bitwise-media-group/evolve

Skill Claude CodeCodex

Test skill beta-skill. Use when exercising the evolve test fixtures.

5 2d ago A 18 tokens original MIT

build-brain

06

Vladick-Pick/business-ontology

Skill Claude CodeCodex

Use after accepted ontology cards change. Compiles cards into the registry graph, checks links, and prepares the model for dashboard, MCP, or agent consumers.

2 1mo ago A 35 tokens original MIT

connect-source

07

Vladick-Pick/business-ontology

Skill Claude CodeCodex

Use before mining a new source. Registers a chat export, spreadsheet, PDF, repo, CRM, dashboard, or feed in 02-source-map.md with trust and read-only access policy.

2 1mo ago A 41 tokens original MIT

drift-flag

08

Vladick-Pick/business-ontology

Skill Claude CodeCodex

Use when an accepted card no longer matches reality. Records the mismatch in 08-drift-and-open-questions.md and routes fixes through propose-change.

2 1mo ago A 35 tokens original MIT

aoa-evals-skills

09

8Dionysus/aoa-evals

Skill Claude CodeCodex

Route the aoa-evals skill family for central proof selection, review, evolution, named results or verdicts, source-linked reports, Eval Forge owner review, and proof lifecycle. Hand repository-local eval selection, application, intake/design, or session-hit classification to aoa-eval. Candidates, readiness checks…

2 2d ago A 84 tokens original Apache-2.0