prompt evaluation agents

7 tagged prompt evaluation, measured the same way as everything else here.

Browse within: prompt-engineering 7

skill-creator

01

shinpr/rashomon

Agent

Generates or modifies optimized skill files. In creation mode, builds from raw user knowledge. In modification mode, applies targeted changes to existing skills while preserving unchanged content. Use when creating new skills or updating existing ones.

18 2d ago A 47 tokens original MIT

skill-eval-reporter

02

shinpr/rashomon

Agent

Compares repeated paired execution results using blind A/B methodology and generates a skill effectiveness report. Use when valid skill-evaluation result pairs are available.

18 2d ago A 35 tokens original MIT

skill-reviewer

03

shinpr/rashomon

Agent

Evaluates skill file quality against optimization patterns and editing principles. Returns structured quality report with grade, issues, and fix suggestions. Use when reviewing created or modified skill content.

18 2d ago A 38 tokens original MIT