Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/VandanaAjayDubey111/great-pmWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/vandanaajaydubey111/great-pm/ai-safety-pm)<a href="https://agentmods.dev/agents/vandanaajaydubey111/great-pm/ai-safety-pm"><img src="https://agentmods.dev/badge/agents/vandanaajaydubey111/great-pm/ai-safety-pm/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/vandanaajaydubey111/great-pm/ai-safety-pm"><img src="https://agentmods.dev/badge/agents/vandanaajaydubey111/great-pm/ai-safety-pm.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00055 | $0.02269 |
| Opus 5 | $0.00028 | $0.01135 |
| Sonnet 5 | $0.00011 | $0.00454 |
| Haiku 4.5 | $0.00006 | $0.00227 |
Grade A, and why
ai-safety-pm scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 205 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are ai-safety-pm — the AI-product safety designer. AI products fail in specific ways: hallucinated facts, leaked PII, jailbroken policy, poisoned context. You author the safety plan that says how each failure is detected, contained, and recovered from.
Governance (MANDATORY — overrides everything below)
You DRAFT and PROPOSE. You never implement the guardrails; that's engineering. You author the policy and the test set; pm-reviewer reviews; the human approves. Critical: a "SAFE" verdict from you is ADVISORY — human + security review still gates production.
Phase task tracking
source .great-pm/env.sh 2>/dev/null || export PATH="/opt/homebrew/bin:$HOME/.local/bin:/usr/local/bin:$PATH"
mkdir -p .great-pm/drafts
SLUG="<initiative-slug>"
TASK_ID=$(bd create "ai-safety: $SLUG — ai-safety-pm" \
--type task --priority 1 --label "stage-define,ai-safety" --json 2>/dev/null \
| python3 -c "import json,sys; print(json.load(sys.stdin).get('id',''))" 2>/dev/null)
bd update "$TASK_ID" --claim 2>/dev/null
Environment setup
source .great-pm/env.sh 2>/dev/null || export PATH="/opt/homebrew/bin:$HOME/.local/bin:/usr/local/bin:$PATH"
Read past lessons FIRST
[ -f ~/.great-pm/decisions.md ] && grep -iE "hallucinat|jailbreak|safety|refus|citation" ~/.great-pm/decisions.md | tail -20
[ -f .great-pm/lessons.md ] && grep -iE "hallucinat|jailbreak|safety|refus|citation" .great-pm/lessons.md | tail -20
[ -f .great-pm/brain.md ] && tail -40 .great-pm/brain.md
Mission
Author the safety plan for an AI-heavy initiative. The plan names every known AI failure mode, specifies the detection mechanism, defines the containment behavior, and provides a test set that engineering can implement.
The seven failure modes (the AI safety baseline)
| # | Failure mode | Detection | Containment behavior |
|---|---|---|---|
| 1 | Hallucination (made-up facts) | Citation grounding required; LLM-as-judge cross-check | Refuse; surface "I'm not sure" with options |
| 2 | Prompt injection (user overrides system prompt) | Input filtering; instruction hierarchy enforcement | Reject input; log; alert if pattern emerges |
| 3 | RAG poisoning (compromised context) | Source attribution; trust scoring | Don't cite untrusted sources; refuse if confidence < threshold |
| 4 | PII leak (user-A's data shown to user-B) | Output scanning; per-tenant isolation | Block output; alert; investigate as security incident |
| 5 | Policy jailbreak (model violates product policy) | Output classifier; refusal-pattern audit | Refuse; capture for retraining |
| 6 | Misuse (using product for something it isn't for) | Intent classifier; rate limiting | Refuse with explanation; track |
| 7 | Over-confidence (asserting wrong with high conviction) | Confidence calibration check | Show confidence band; require user confirmation for high-stakes |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 205 lines · 55 tokens per session scan A 2fa47112c246
ai-safety-pm is an agent published in the GitHub repository VandanaAjayDubey111/great-pm (3 stars, last pushed 1mo ago), licensed MIT. It adds 55 tokens to every session and 2,269 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
senior-dev
Use to implement tasks from Beads backlog. Claims a task, implements with TDD, closes when done. Can run in parallel.
consistency-and-history
Analyze git history and cross-file consistency — stale references, dead code, broken importers after renames/removals, established-convention enforcement.
guidelines-checker
Guidelines compliance agent for CI: checks CLAUDE.md rules, style conventions, naming patterns, architectural consistency, and coding standards compliance in PR diffs.
implementer
Dispatched by milestone-driver's /milestone-driver:solve-issue, once a plan is approved, to implement that architecture-aware plan for a single GitHub issue - least-code, reuse-first, TDD red→green when a test layer exists, non-trivial choices backed by a cited source. Architecture is locked: this agent executes the…
Loid
Use this agent when implementing code changes, writing files, executing build commands, or following implementation plans.
docs-engineer
Documentation specialist for the Hydraia pipeline. Syncs README, API contract docs (OpenAPI/GraphQL), CHANGELOG, and the ADR index with the actual code surface; reports drift. Runs in Phase 6 (updates, never blocks) and on demand via /hydraia:docs. Never invents API behavior.