eval harness skills

5 tagged eval harness, measured the same way as everything else here.

Browse within: evals 5github-actions 5

HelloThisWorld/agent-skill-verification-template

Skill Claude CodeCodex

Answers questions about a codebase using source-grounded evidence. Every factual claim must cite a specific file and line; ambiguous or unsupported questions return insufficientevidence instead of a guess.

1 9d ago A 41 tokens original MIT

glossary

02

HelloThisWorld/agent-skill-verification-template

Skill Claude CodeCodex

Looks up "glossary " on English Wikipedia and returns a source-grounded definition rendered as a web page. Every claim cites the exact snapshot line carrying the term; unknown terms return insufficientevidence instead of a guess.

1 9d ago A 51 tokens original MIT

HelloThisWorld/agent-skill-verification-template

Skill Claude CodeCodex

Route any query to exactly one documented Open Mind capability (glossary / structure / search) using the deterministic if-else floor — the mode where routing never depends on a model. The decision is grounded by citing the capability's documented registry line. Runs Open Mind's real Python implementation through the…

1 9d ago A 68 tokens original MIT