Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add rules/ai-plugin-marketplace/template/evaluation-protocolgit clone --depth 1 https://github.com/ai-plugin-marketplace/templateWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00174 |
| Opus 5 | $0.00000 | $0.00087 |
| Sonnet 5 | $0.00000 | $0.00035 |
| Haiku 4.5 | $0.00000 | $0.00017 |
Grade A, and why
evaluation-protocol scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
Copies of this mod
1 near-identical copy found in the catalogue:
- evaluation-protocol — 100% identical, 0 lines differ
What it actually says
Evaluation Protocol
Pass/Fail Criteria
- Define pass/fail criteria for each test case BEFORE executing any test runs
- Pass criteria should be based on semantic equivalence, not exact string matching
- A test passes if the output achieves the same goal as the expected outcome
Blind Testing
- Never expose expected outcomes to test-subject agents
- Test subjects receive ONLY the skill content and input
- Any accidental leakage invalidates the test run
Reporting
- Report results per-tier in a structured table format
- Include failure analysis with root cause and specific recommendations
- Order recommendations by expected impact (highest first)
- The refinement report is the primary deliverable — it must be actionable
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 28 lines · 0 tokens per session scan A 583c8200ee22
evaluation-protocol is a cursor rule published in the GitHub repository ai-plugin-marketplace/template (10 stars, last pushed 2mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 174 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other cursor rules, from other repositories
ai-skills
Operate the AI-SKILLS library autonomously, selecting the right skill or prompt for each task.
angular-20
This rule provides comprehensive best practices and coding standards for Angular development, focusing on modern TypeScript, standalone components, signals, and performance optimizations.
dev-standard
Apache Superset development standards and guidelines for Cursor IDE.
typescript
Changes to these high-fan-out internals can affect every message, delta, element, or rerun. Keep work in them minimal, and benchmark changes with representative stress-test apps.
coolify-ai-docs
Master reference to all Coolify AI documentation in .ai/ directory.
python_lib
Tips and guidelines specific to the development of the Streamlit Python library, not applicable to scripts and e2e tests.