DYAI2025

44 mods across 1 repository, 6 stars between them.

bench-watch

01

DYAI2025/Plumbline

Command Claude Code

Launch or attach to a Plumbline benchmark slice, poll it to completion, and emit the canonical anti-Goodhart per-arm-model summary (catch-rate AND cry-wolf). For the run-slice / wait-for-it / show-the-per-arm-model-summary loop.

6 9d ago A 54 tokens

release-doctor

02

DYAI2025/Plumbline

Command Claude Code

Verify the Plumbline version is current and self-consistent (VERSION equals the latest GitHub release equals the CLI), dogfood plumbline update --check, and confirm README claims are derived not stale. For is-the-github-version-up-to-date and are-all-claims-correct checks.

6 9d ago A 60 tokens

triage-impact

03

DYAI2025/Plumbline

Command Claude Code

Triage the Plumbline backlog and recommend the next move by impact times (safe times fast) — the biggest authenticity win that is also low-risk and bounded. For biggest-impact-from-the-backlog and identify-the-most-effective-levers asks.

6 9d ago A 49 tokens

SessionStart

04

DYAI2025/Plumbline

Hook Claude Code

Runs when a session starts, executing session-start.sh. From DYAI2025/Plumbline.

6 9d ago A tokens not measured

Plumbline

05

DYAI2025/Plumbline

Settings file Claude Code

Agent settings declaring 1 hook event (SessionStart).

6 9d ago A tokens not measured

DYAI2025/Plumbline

Skill Claude CodeCodex

Commands and hard-won honesty rules for Plumbline measurement, benchmark, council A/B and Council-GUI slices. Load before running the metrics harness (emitrun, processhealth, councilreviewscorer, armareviewrunner, councilmeasurementrun, councilfreediversityprobe), before the Council GUI composition root, or before…

6 9d ago A 86 tokens

Plumbline CLAUDE.md

07

DYAI2025/Plumbline

Instructions file

Claude Code instructions for DYAI2025/Plumbline, covering claude.md, what this repo is, where things live (the four-way coupling), common commands and run a single test module (each is a standalone bash script).

6 9d ago B 12,434 tokens

der-minimalist

08

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Der Minimalist direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Reduktionsdruck: den kleinsten starken Kern einer Idee freilegen, damit ein Produkt nicht an Ballast, Overengineering…

6 9d ago A 129 tokens

der-nutzeranwalt

09

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Der Nutzeranwalt direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Nutzerschutz: das echte menschliche Erleben vor dem Bildschirm gegen Entwicklerlogik, Feature-Verliebtheit und…

6 9d ago A 150 tokens

der-provokateur

10

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Der Provokateur direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Stoerung: einen zu glatten Konsens aufbrechen, damit eine Entscheidung nicht aus Gruppendruck und bequemen…

6 9d ago A 134 tokens

der-pruefer

11

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Der Pruefer direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Qualitaetsdruck: schwache Konzepte nicht hoeflich durchwinken, sondern belastbarer machen.

6 9d ago A 115 tokens

der-systemdenker

12

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Der Systemdenker direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Kohaerenzdruck: aus einer Idee die Wechselwirkungen, Rueckwirkungen und Langzeitfolgen im Gesamtsystem sichtbar…

6 9d ago A 150 tokens

die-macherin

13

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Die Macherin direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Umsetzungsdruck: aus Diskussionen konkrete naechste Schritte, MVPs, Experimente, Tickets und Entscheidungen machen…

6 9d ago A 143 tokens

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Die Marktschaerferin direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Marktdruck: ein unklares Nutzenversprechen nicht hoeflich durchwinken, sondern auf reale Aussenwirkung hin…

6 9d ago A 169 tokens

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Die Risiko-Waechterin direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Schutzdruck: Vertrauens-, Sicherheits- und Missbrauchsrisiken sichtbar machen, damit ein Produkt nicht in…

6 9d ago A 155 tokens

die-uebersetzerin

16

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Die Uebersetzerin direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Verstaendlichkeitsdruck: aus Komplexitaet klare Sprache, einfache Modelle, gute Namen und verstaendliche…

6 9d ago A 164 tokens

die-visionaerin

17

DYAI2025/Plumbline

Skill Claude CodeCodex

Verkoerpere den Charakter Die Visionaerin direkt in der Antwort. Erzeuge keinen neuen Skill, keinen Meta-Prompt und kein Prompt-Paket, sofern der Nutzer das nicht ausdruecklich verlangt. Die Aufgabe ist Moeglichkeitsdruck: aus einer Idee ihr groesseres Potenzial herausarbeiten, damit ein Produkt nicht nur aus…

6 9d ago A 146 tokens

agileteam-bench

18

DYAI2025/Plumbline

Command

Run the drift-vs-precision comparison for /agileteam — the frozen main process vs. the evolving agileteam-improved process over a fixed task corpus, with pinned agent versions, then analyse process health.

6 9d ago A 43 tokens

agileteam

19

DYAI2025/Plumbline

Command

Orchestrate an autonomous, defense-in-depth TDD multi-agent team (requirements → spec-sanity gate → planner → coder/reviewer loop → verification/security/validation/judgment gates → human acceptance → retrospective) to build a feature end-to-end against fully verified, independently validated requirements.

6 9d ago A 58 tokens

bench-oracle

20

DYAI2025/Plumbline

Command

Empirically measure an agent/prompt/process change instead of asserting it works — build a task corpus, run a deterministic mutation oracle (sabotage the code, see which tests catch it), and write an honest report including negative results. Use to compare two agent variants (e.g. baseline vs. a DNA/prompt change) or…

6 9d ago A 76 tokens

concilium

21

DYAI2025/Plumbline

Command

Run the Concilium — a four-body council (Market Realist · Tech Arbiter · Skeptic · Distribution Realist) that critically stress-tests a product idea AND its team/agent constellation, generates real friction, then iterates to an emergent shared pattern and a single evidence-grounded recommendation (proceed / sharpen /…

6 9d ago A 104 tokens

honest-status

22

DYAI2025/Plumbline

Command

Give an honest status of the current work — what was actually done, hoped-for vs. real result, and what was missed or remains unproven. The plumb line for your own progress. Use when the user asks "what did you do / how well did it go / what's the status", before continuing a long task, or whenever a claim of…

6 9d ago A 80 tokens

merge-when-true

23

DYAI2025/Plumbline

Command

Gate a PR/branch merge on Plumbline's TRUE-green standard — never on passing tests alone. Here "green" means the work hangs true (Reality-Ledger real-boundary, wired-in-prod, independence, confirmed customer value, no silently-downgraded RED), the CI conclusion is success (not merely mergeable), and runall.sh is green…

6 9d ago A 107 tokens

DYAI2025/Plumbline

Command

Run an honest, credit-careful, key-safe LIVE real-boundary smoke against OpenRouter (a single model via councilinference, a council via deepseekreview preset, or the GUI proxy). Probe reachability first, gate the live call, leak-check the key, capture honest attrition, and record real-boundary-smoke in the reality…

6 9d ago A 92 tokens