Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/hypnguyen1209/offensive-claudeWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/hypnguyen1209/offensive-claude/engage.scorecard)<a href="https://agentmods.dev/commands/hypnguyen1209/offensive-claude/engage.scorecard"><img src="https://agentmods.dev/badge/commands/hypnguyen1209/offensive-claude/engage.scorecard/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/commands/hypnguyen1209/offensive-claude/engage.scorecard"><img src="https://agentmods.dev/badge/commands/hypnguyen1209/offensive-claude/engage.scorecard.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00018 | $0.00395 |
| Opus 5 | $0.00009 | $0.00198 |
| Sonnet 5 | $0.00004 | $0.00079 |
| Haiku 4.5 | $0.00002 | $0.00040 |
Grade A, and why
engage.scorecard scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
/engage.scorecard
Tracks how often a model's verdicts (finding-validator / finding-checker PASS/KILL/DOWNGRADE) were
later overturned, and decides — fail-closed — when a (model, decision_class) cell is trustworthy
enough to skip an expensive re-validation in the autopilot loop. Backed by engine/model_scorecard.py.
Usage
/engage.scorecard {record|rate|trusted|stats} ...
Process
- Record outcomes as you confirm them:
model_scorecard.py record --model opus --class finding-validator:PASS --outcome correct|overturned(a "miss" = a verdict later overturned: a PASS that was a false positive, a KILL that was real). - Consult before skipping a re-check (autopilot):
model_scorecard.py trusted --model opus --class finding-validator:PASS→ exit 0 trusted, 3 not. A cell is trusted ONLY when the Wilson 95% upper bound on its miss-rate ≤ 5% and there are enough samples (≈73+ clean) — fail-closed: a new/rarely-seen model or one recent miss is NOT trusted, so the autopilot keeps re-validating. - Review —
model_scorecard.py statsshows per-(model, class) n / overturned / upper-bound / trusted.
Notes
- Separate sqlite store (
$MODEL_SCORECARD_DB), distinct from the JSONL engagement-memory pattern store. - This only ever adds a fast-path for well-proven cells; it never relaxes the proof bar — an untrusted cell simply gets the normal finding-validator + finding-checker treatment.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 32 lines · 18 tokens per session scan A f56d15576d22
engage.scorecard is a command published in the GitHub repository hypnguyen1209/offensive-claude (357 stars, last pushed 24d ago), licensed MIT. It adds 18 tokens to every session and 395 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
discover
Run a full product discovery cycle — from outcome definition through opportunity mapping, prioritisation, and experiment design. Use when the team isn't sure what to build next, or before writing a PRD for a complex feature space.
voice-compliance
Voice/telephony compliance check — invokes voice-ai-reviewer to produce TM-voice-{slug}.md with TCPA, STIR/SHAKEN, state recording-consent, EU AI Act Art. 50, and synth-voice deepfake-law gaps.
claude-tracker
List recent Claude Code sessions with live status.
qa
Smoke or browser-walk a running app. Report only. Do not implement. Do not merge.
esp-harden
Harden and inspect ESP32 firmware for field failures, crashes, memory, and security.
replication-package
Scaffold or audit a social-science replication package at a target directory, and audit the manuscript and its archived research objects against FAIR principles.