Use this agent to predict how users will creatively misuse and weaponize features for social manipulation, fraud, or unintended behavioral cascades. This agent understands that every feature becomes a tool for gaming the system and unintended consequences.
Use this agent when you need brutally honest, unfiltered feedback on complex problems, especially when the primary model is being too agreeable or when an implementation is going off track. Deploy when you need a high-IQ contrarian perspective that cuts through politeness to deliver maximum informational density.
Use this agent to file a GitHub issue from a short problem report — it validates labels, searches open AND closed issues for duplicates, enhances an existing issue (comment/reopen-as-regression) instead of filing a duplicate, and writes a well-formed issue with actionable repro data. Invoke it (ideally with…
Use this agent when you need exhaustive, obsessive-level analysis of a specific technical problem or implementation detail. Perfect for deep dives into performance bottlenecks, architectural decisions, algorithm optimizations, or any scenario where you need every possible angle examined, every edge case documented…
Use this agent for catastrophic failure analysis of proposed solutions. This agent channels Murphy's Law to identify everything that WILL go wrong - race conditions, data corruption, cascade failures, and 3 AM production disasters. Essential devil's advocate for production readiness.
Review project documentation for consistency with recent code changes. Docs are part of the product — if a feature was added but not documented, or an old pattern was removed but still referenced, that's a bug.
You are the QA controller for the mcp-cli project. Your job is to verify that a GitHub issue's described functionality is implemented, tested, and working — then apply qa:pass or qa:fail to the PR with evidence. Both labels are equally valid outcomes and both are valuable; your job is to be accurate about which one…
Orchestrate implementation of one or more GitHub issues in parallel using worktree-isolated agents. Each issue gets its own git worktree, its own agent, and is independently implemented, tested, and verified.
Analyze open GitHub issues, group them into thematic arcs, and write a sprint file for orchestrator context. Use when deciding what to work on next, assessing board health, or updating strategic context.
Build an autonomous sprint skill for a target project. Explores the project's workflows, constraints, and tooling, then produces a tailored sprint skill that lets Claude orchestrate parallel implementation sessions end-to-end. Use when setting up auto-sprint in a new repo, or when adapting sprint orchestration to a…
Review recent PRs in a feature area for architectural consistency, missed integration points, duplicated patterns, and wrong directions. File issues for problems found. Use after 3-4 related issues merge or when a session's cost exceeds $15.
Backfill diary entries from Claude Code session transcripts. Extracts insights, learnings, and patterns from past sessions and writes structured diary entries to .claude/diary/. Use when the user says "write diary", "backfill diary", "/diary", "what did we do today", or wants to capture session learnings. Also use…
Post-implementation triage: measure actual diff metrics to determine review depth. Run after implement, before review. Usage: /estimate (in a worktree with changes).
Full-text search of Claude Code session histories to find and resume past conversations. Use when user says "find session", "search sessions", "find that conversation where", "what were my last sessions", or wants to locate a previous Claude session.
Mine flaky test failures from Claude Code session transcripts. Scans JSONL session files for test failure patterns in tool results, aggregates by test file and name, and reports which tests fail across multiple sessions. Use when investigating test reliability, after a sprint to check for new flaky tests, or when the…
Audit and garbage-collect a committed Claude-memory store (MEMORY.md + per-fact files). Evidence-based, parallel, then a gated apply. Use when MEMORY.md has grown large, after several sprints, when memory has drifted from CLAUDE.md, or when the user says "/memory-gc", "clean up memory", "audit memory", "the memory is…
Author, refine, or migrate a doing-it-wrong architectural rule, and harvest recurring mistakes from merged PRs. Routes on its argument — harvest (mine PR review comments for recurring author mistakes), init (bootstrap the rule engine into a repo), add (author a rule), or no arg (overview + "is this a good rule?"…
Extract actionable lessons from sprint history. Buckets every Claude Code session into the sprint it ran in, builds a deterministic per-session digest (tokens, duration, errors, orchestration signals, evidence), then runs a multi-agent Workflow that classifies each session (orchestrator vs worker), scans for tool…
Sprint lifecycle: plan, run, review, retro. The top-level orchestrator for autonomous issue resolution. Use for "run a sprint", "plan the sprint", "sprint review", "sprint retro", "/sprint", "/sprint plan", "/sprint review", "/sprint retro", or any variant.
Step back and interrogate the frame instead of executing it. /why GENERATES the sharp questions that should have been asked about a request, plan, artifact, or your own intended action — it does NOT answer them. Use when handed an assignment that smells mis-scoped, when about to loosen a constraint to silence a signal…