DYAI2025/Plumbline

Plumbline — a self-learning, customer-value-governed agile AI agent team for Claude Code. 87 subagents + skills, TDD defense-in-depth gates, Kaizen retros, a four-body adversarial council, and an empirically benchmarked QA harness. Does it hang true?

6Stars on the repository
44Mods indexed here, across every type
11d agoLast push, which is what freshness is scored on
noneNo LICENSE: all rights reserved, so bodies are not copied

bench-watch

01

DYAI2025/Plumbline

Command Claude Code

Launch or attach to a Plumbline benchmark slice, poll it to completion, and emit the canonical anti-Goodhart per-arm-model summary (catch-rate AND cry-wolf). For the run-slice / wait-for-it / show-the-per-arm-model-summary loop.

6 11d ago A 54 tokens

release-doctor

02

DYAI2025/Plumbline

Command Claude Code

Verify the Plumbline version is current and self-consistent (VERSION equals the latest GitHub release equals the CLI), dogfood plumbline update --check, and confirm README claims are derived not stale. For is-the-github-version-up-to-date and are-all-claims-correct checks.

6 11d ago A 60 tokens

triage-impact

03

DYAI2025/Plumbline

Command Claude Code

Triage the Plumbline backlog and recommend the next move by impact times (safe times fast) — the biggest authenticity win that is also low-risk and bounded. For biggest-impact-from-the-backlog and identify-the-most-effective-levers asks.

6 11d ago A 49 tokens

agileteam-bench

04

DYAI2025/Plumbline

Command

Run the drift-vs-precision comparison for /agileteam — the frozen main process vs. the evolving agileteam-improved process over a fixed task corpus, with pinned agent versions, then analyse process health.

6 11d ago A 43 tokens

agileteam

05

DYAI2025/Plumbline

Command

Orchestrate an autonomous, defense-in-depth TDD multi-agent team (requirements → spec-sanity gate → planner → coder/reviewer loop → verification/security/validation/judgment gates → human acceptance → retrospective) to build a feature end-to-end against fully verified, independently validated requirements.

6 11d ago A 58 tokens

bench-oracle

06

DYAI2025/Plumbline

Command

Empirically measure an agent/prompt/process change instead of asserting it works — build a task corpus, run a deterministic mutation oracle (sabotage the code, see which tests catch it), and write an honest report including negative results. Use to compare two agent variants (e.g. baseline vs. a DNA/prompt change) or…

6 11d ago A 76 tokens

concilium

07

DYAI2025/Plumbline

Command

Run the Concilium — a four-body council (Market Realist · Tech Arbiter · Skeptic · Distribution Realist) that critically stress-tests a product idea AND its team/agent constellation, generates real friction, then iterates to an emergent shared pattern and a single evidence-grounded recommendation (proceed / sharpen /…

6 11d ago A 104 tokens

honest-status

08

DYAI2025/Plumbline

Command

Give an honest status of the current work — what was actually done, hoped-for vs. real result, and what was missed or remains unproven. The plumb line for your own progress. Use when the user asks "what did you do / how well did it go / what's the status", before continuing a long task, or whenever a claim of…

6 11d ago A 80 tokens

merge-when-true

09

DYAI2025/Plumbline

Command

Gate a PR/branch merge on Plumbline's TRUE-green standard — never on passing tests alone. Here "green" means the work hangs true (Reality-Ledger real-boundary, wired-in-prod, independence, confirmed customer value, no silently-downgraded RED), the CI conclusion is success (not merely mergeable), and runall.sh is green…

6 11d ago A 107 tokens

DYAI2025/Plumbline

Command

Run an honest, credit-careful, key-safe LIVE real-boundary smoke against OpenRouter (a single model via councilinference, a council via deepseekreview preset, or the GUI proxy). Probe reachability first, gate the live call, leak-check the key, capture honest attrition, and record real-boundary-smoke in the reality…

6 11d ago A 92 tokens

persist-learning

11

DYAI2025/Plumbline

Command

Persist a human-APPROVED retro/learning at the narrowest fitting level (CLAUDE.md / a global agent / a new skill) and land it safely — preview the exact diff before writing, branch, runall, PR, watch CI by the pushed HEAD sha, merge normal. Use after /reflect has surfaced a learning and the human has approved which…

6 11d ago A 84 tokens

plumbline-update

12

DYAI2025/Plumbline

Command

Check or apply Plumbline updates with verified-or-revert semantics. Use for /plumbline-update check, /plumbline-update apply, or rollback planning.

6 11d ago A 34 tokens

reflect-skills

13

DYAI2025/Plumbline

Command

Run the local claude-reflect skill-discovery fallback and propose skill/rule changes without applying them.

6 11d ago A 21 tokens

reflect

14

DYAI2025/Plumbline

Command

Run the local claude-reflect retrospective discovery fallback over the current session/workspace evidence.

6 11d ago A 18 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: