DYAI2025/Plumbline
Command Claude Code
Launch or attach to a Plumbline benchmark slice, poll it to completion, and emit the canonical anti-Goodhart per-arm-model summary (catch-rate AND cry-wolf). For the run-slice / wait-for-it / show-the-per-arm-model-summary loop.
DYAI2025/Plumbline
Command Claude Code
Verify the Plumbline version is current and self-consistent (VERSION equals the latest GitHub release equals the CLI), dogfood plumbline update --check, and confirm README claims are derived not stale. For is-the-github-version-up-to-date and are-all-claims-correct checks.
DYAI2025/Plumbline
Command Claude Code
Triage the Plumbline backlog and recommend the next move by impact times (safe times fast) — the biggest authenticity win that is also low-risk and bounded. For biggest-impact-from-the-backlog and identify-the-most-effective-levers asks.
DYAI2025/Plumbline
Command
Run the drift-vs-precision comparison for /agileteam — the frozen main process vs. the evolving agileteam-improved process over a fixed task corpus, with pinned agent versions, then analyse process health.
DYAI2025/Plumbline
Command
Orchestrate an autonomous, defense-in-depth TDD multi-agent team (requirements → spec-sanity gate → planner → coder/reviewer loop → verification/security/validation/judgment gates → human acceptance → retrospective) to build a feature end-to-end against fully verified, independently validated requirements.
DYAI2025/Plumbline
Command
Empirically measure an agent/prompt/process change instead of asserting it works — build a task corpus, run a deterministic mutation oracle (sabotage the code, see which tests catch it), and write an honest report including negative results. Use to compare two agent variants (e.g. baseline vs. a DNA/prompt change) or…
DYAI2025/Plumbline
Command
Run the Concilium — a four-body council (Market Realist · Tech Arbiter · Skeptic · Distribution Realist) that critically stress-tests a product idea AND its team/agent constellation, generates real friction, then iterates to an emergent shared pattern and a single evidence-grounded recommendation (proceed / sharpen /…
DYAI2025/Plumbline
Command
Give an honest status of the current work — what was actually done, hoped-for vs. real result, and what was missed or remains unproven. The plumb line for your own progress. Use when the user asks "what did you do / how well did it go / what's the status", before continuing a long task, or whenever a claim of…
DYAI2025/Plumbline
Command
Gate a PR/branch merge on Plumbline's TRUE-green standard — never on passing tests alone. Here "green" means the work hangs true (Reality-Ledger real-boundary, wired-in-prod, independence, confirmed customer value, no silently-downgraded RED), the CI conclusion is success (not merely mergeable), and runall.sh is green…
DYAI2025/Plumbline
Command
Run an honest, credit-careful, key-safe LIVE real-boundary smoke against OpenRouter (a single model via councilinference, a council via deepseekreview preset, or the GUI proxy). Probe reachability first, gate the live call, leak-check the key, capture honest attrition, and record real-boundary-smoke in the reality…
DYAI2025/Plumbline
Command
Persist a human-APPROVED retro/learning at the narrowest fitting level (CLAUDE.md / a global agent / a new skill) and land it safely — preview the exact diff before writing, branch, runall, PR, watch CI by the pushed HEAD sha, merge normal. Use after /reflect has surfaced a learning and the human has approved which…
DYAI2025/Plumbline
Command
Check or apply Plumbline updates with verified-or-revert semantics. Use for /plumbline-update check, /plumbline-update apply, or rollback planning.
DYAI2025/Plumbline
Command
Run the local claude-reflect skill-discovery fallback and propose skill/rule changes without applying them.
DYAI2025/Plumbline
Command
Run the local claude-reflect retrospective discovery fallback over the current session/workspace evidence.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: