llm-evals skills

38 tagged llm-evals, measured the same way as everything else here.

Browse within: prompt-engineering 37

samarailly51-pixel/claimpilot-harness

Skill Claude CodeCodex

Review and structure auto insurance bodily injury claims, including intake triage, evidence completeness, injury causation, treatment chronology, medical necessity, wage-loss support, negotiation risks, and human escalation. Use when analyzing bodily injury claim files, designing claims-agent workflows, drafting…

117 9d ago A 76 tokens original MIT

Extract Approach

02

ralfyishere/rules-with-receipts

Skill Claude CodeCodex

After solving a non-trivial problem, capture the reusable approach as a short learning note in .claude/learnings/ before calling the work complete. Activate after: a hard bug is solved, a tricky architecture or strategy decision lands, a difficult prompt is fixed, a mistake occurs that must not repeat, an eval failure…

2 1mo ago A 112 tokens original MIT

Plan Gate

03

ralfyishere/rules-with-receipts

Skill Claude CodeCodex

Require a written plan before starting complex or multi-step work. Activate when the task involves building a feature, implementing something new, refactoring, migrating, restructuring a document, multi-file edits, research projects, or any work with more than 3 dependent steps, unclear requirements, or an expensive…

2 1mo ago A 108 tokens original MIT

ralfyishere/rules-with-receipts

Skill Claude CodeCodex

Learn from prior mistakes and repeated patterns - review what failed, extract a reusable lesson, apply it to the next attempt. Activate when the user corrects your output, when an approach fails and needs a retry, when you notice the same friction recurring across tasks, and at the end of significant multi-step work.…

2 1mo ago A 111 tokens original MIT