Agent Claude Code
Adversarial Loop 3 evaluator for the vertical bench. Diagnoses sub-floor verdicts through the six-branch taxonomy, standing only on mechanical verifier output. WIN confirmation lives in bench-win-confirm.
4 tagged codebase understanding, measured the same way as everything else here.
Agent Claude Code
Adversarial Loop 3 evaluator for the vertical bench. Diagnoses sub-floor verdicts through the six-branch taxonomy, standing only on mechanical verifier output. WIN confirmation lives in bench-win-confirm.
Agent Claude Code
Reads any bench run (validation or paid, win or loss) and returns the material for a better scenario - where the baseline struggled and what Sense reached that it did not. Never issues a verdict; never diagnoses a loss (that is bench-evaluator).
Agent Claude Code
WIN-confirmation vertex for the vertical bench. Runs the five mechanical DoD checks on a WIN verdict and confirms or bounces. Never diagnoses a sub-floor verdict; never fault-finds a clean win.