warden-bench
01Command
Run the golden benchmark suite for one token-warden agent (or all) and compare results against the frozen run1 and best baselines.
Claude Code plugin that makes coding agents measurably cheaper over time: collect token costs, distill candidate rules, benchmark them on a frozen golden suite, and keep only rules that earn their context rent.
Command
Run the golden benchmark suite for one token-warden agent (or all) and compare results against the frozen run1 and best baselines.
Command
Dollar accounting — what does each active rule actually save, in money? Translates the token-measured verdict into dollars using a price table and the agent's own token-type mix, with a per-session net and a break-even. Read-only; spends no tokens.
Command
Zero-token power planner — from the agent's own recorded run-to-run variance, report the minimum detectable saving (MDS) at each run count and how many runs per side a target saving needs, so a benchmark burn is provably adequately powered before it starts.
Command
Show token-warden rule receipts — the per-rule verdict card with token savings vs. rent, per-task pass/fail, the tool-call/file-reread quality profile, and provenance (model + golden-suite hash).
Command
Measure pending token-warden candidate rules for an agent on the golden suite, evict or activate them, and recompile the agent's memory.
Command
Show token-warden status — per-agent runs and rules, golden-suite totals vs frozen baselines, learning curves, and the rule ledger with recent evictions.