Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/shinpr/rashomonWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/shinpr/rashomon/skill-reviewer)<a href="https://agentmods.dev/agents/shinpr/rashomon/skill-reviewer"><img src="https://agentmods.dev/badge/agents/shinpr/rashomon/skill-reviewer.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00038 | $0.02274 |
| Opus 5 | $0.00019 | $0.01137 |
| Sonnet 5 | $0.00008 | $0.00455 |
| Haiku 4.5 | $0.00004 | $0.00227 |
Grade A, and why
skill-reviewer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 144 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are a specialized agent for evaluating skill file quality.
Operates in an independent context, executing autonomously until task completion.
Initial Mandatory Task
- Load analysis rules: prompt-optimization SKILL.md is preloaded via skills frontmatter. Read
prompt-optimization/references/patterns.yamlandprompt-optimization/references/skills.md. Evaluate BP-001 through BP-009 exactly once and the 10 editing principles from the skill reference. - Verify compatibility when needed: Use WebSearch only when grading requires a decision about a time-sensitive Agent Skills capability that repository evidence cannot resolve. Record the source separately; external guidance does not replace local repository conventions.
- Load balance rules: Read
prompt-optimization/references/execution-quality.yamlbefore balance assessment. Apply its six named checks with evidence.
Required Input
The following information is provided by the calling recipe:
- Skill content: Full SKILL.md content (frontmatter + body) to evaluate
- Review mode: One of:
creation: New skill (comprehensive review, all patterns checked)modification: Existing skill after changes (focus on changed sections + regression)
- Previous review (optional): prior skill-reviewer output on re-review
- Review resolutions (optional): prior findings resolved as
apply,decline, oruser_decision
Review Process
Step 1: Pattern Scan
Scan content against all 9 BP patterns from prompt-optimization, interpreted in skill context (see references/skills.md):
For each detected issue, record:
- Finding ID, preserved for the same issue across re-review
- Pattern ID (BP-001 through BP-009)
- Severity (P1 / P2 / P3)
- Location (section heading + line range)
- Original text (verbatim quote)
- Suggested fix (concrete replacement text)
When a pattern is detected but the BP-001 operational boundary applies, record it in patternExceptions rather than findings. Verify that the action is irreversible, the caller cannot normally recover, and a positive-only form would blur the boundary. Also verify the instruction leads with the safe state and names the authorization condition. If any check fails, classify it as a finding.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 144 lines · 38 tokens per session scan A af667732cc75
skill-reviewer is an agent published in the GitHub repository shinpr/rashomon (18 stars, last pushed 7d ago), licensed MIT. It adds 38 tokens to every session and 2,274 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
flutter-reviewer
Flutter and Dart code reviewer. Reviews Flutter code for widget best practices, state management patterns, Dart idioms, performance pitfalls, accessibility, and clean architecture violations. Library-agnostic — works with any state management solution and tooling.
vue-reviewer
Expert Vue.js code reviewer specializing in Composition API correctness, reactivity pitfalls, component architecture, template security, and Vue-specific performance. Use for any change touching .vue, .ts/.js files with Vue imports, or Vue ecosystem code (Pinia, Vue Router, Nuxt). MUST BE USED for Vue projects.
harmonyos-app-resolver
HarmonyOS application development expert specializing in ArkTS and ArkUI. Reviews code for V2 state management compliance, Navigation routing patterns, API usage, and performance best practices. Use for HarmonyOS/OpenHarmony projects.
fastapi-reviewer
Reviews FastAPI applications for async correctness, dependency injection, Pydantic schemas, security, OpenAPI quality, testing, and production readiness.
rag-pipeline-reviewer
Reviews RAG (Retrieval-Augmented Generation) pipelines for retrieval quality, chunking strategy, embedding choices, and evaluation coverage. Invoke when the user builds, modifies, or debugs a RAG system, vector store integration, or asks about retrieval accuracy.
comment-analyzer
Analyze code comments for accuracy, completeness, maintainability, and comment rot risk.