Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add CassioRoos/godfly-skills --skill red-blue-reviewgit clone --depth 1 https://github.com/CassioRoos/godfly-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/cassioroos/godfly-skills/red-blue-review)<a href="https://agentmods.dev/skills/cassioroos/godfly-skills/red-blue-review"><img src="https://agentmods.dev/badge/skills/cassioroos/godfly-skills/red-blue-review/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/cassioroos/godfly-skills/red-blue-review"><img src="https://agentmods.dev/badge/skills/cassioroos/godfly-skills/red-blue-review.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00056 | $0.00507 |
| Opus 5 | $0.00028 | $0.00253 |
| Sonnet 5 | $0.00011 | $0.00101 |
| Haiku 4.5 | $0.00006 | $0.00051 |
Grade A, and why
red-blue-review scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Red/Blue Review
Use opposing lenses to find what breaks and what defends it. Do not list generic risks. Produce concrete attack paths, failure paths, and mitigations.
Inputs To Gather
- The artifact: code, diff, design, ADR, spec, migration, runbook, or launch plan.
- The asset at risk: money, data integrity, availability, privacy, developer time, customer trust.
- The threat or failure model: malicious actor, operator error, dependency failure, scale, bad data, ambiguous requirements.
- The deployment context: environments, permissions, observability, rollback path.
Red Team
Act like the system will be misused, overloaded, misconfigured, or attacked.
Check:
- Trust boundaries and privilege escalation.
- Input validation and malformed data.
- Authentication, authorization, and tenant isolation.
- Secrets handling and accidental leakage.
- Retry storms, duplicate effects, race conditions, and idempotency gaps.
- Silent data corruption and partial writes.
- Missing timeouts, circuit breakers, and backpressure.
- Operational mistakes: bad deploy, bad config, bad rollback, stale docs.
Blue Team
Defend with practical controls, not wishful thinking.
For each credible issue:
- Prevention: code/design change that stops it.
- Detection: logs, metrics, traces, alerts, invariants, tests.
- Containment: rate limits, circuit breakers, feature flags, kill switches.
- Recovery: rollback, replay, reconciliation, data repair.
- Ownership: who notices and who acts.
Output
## Red Team Findings
### Critical
- Attack/failure path:
Evidence:
Impact:
Why current controls fail:
### High
- Attack/failure path:
Evidence:
Impact:
Why current controls fail:
## Blue Team Plan
- Control:
Covers:
Implementation:
Validation:
## Launch Gate
- Ship:
- Block:
- Spike:
Rule
If there is no evidence for a control, say "control not demonstrated." A design promise is not a control.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago Changed · -3 lines 117015df1640
- 11d ago First seen · 88 lines · 56 tokens per session scan A fea893195631
red-blue-review is a skill published in the GitHub repository CassioRoos/godfly-skills (1 stars, last pushed 2d ago), licensed MIT. It adds 56 tokens to every session and 507 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
auto-test-code
A structured process for critically reviewing and testing software code. It records review findings, test plans, commands, results, and supporting files in a project workspace.
git-pr-review
A read-only reviewer for GitHub pull requests, which are proposed code changes submitted for review. It produces an evidence-based report about whether a pull request should be merged.
code-remediate
Apply selected review fixes; bare PR targets use current online items, while PR +review adds the latest matching artifact.
code-review
Close PRs at an evidence gate or review local diffs/PRs with specialists and JSON artifacts.
assess
Analyze issue/PR/problem before implementation; produce source-backed findings and measurable gates.
code-reviewer
A code-review workflow that checks completed work against its requirements or plan before merging. It groups findings by severity and gives each issue a fix and a way to verify it.