Turn "my app + an agent that demos" into "benchmarked, browser-verified, evidence-backed, prod-proven, and looping." Use when a (solo) founder wants to prove an AI agent works IN their real app — across all its UI surfaces — without cheating. Triggers: "set up proofloop", "proof-loop my app", "benchmark my agent's…
Every UI change MUST be visually verified in the running app before declaring done. Code-level changes without visual confirmation are unverified assumptions. But verification is not just "does it render?" — it's "does it belong?".
Cursor rule "forecasting_os" from HomenShum/NodeBenchAI, covering forecasting os, architecture, three surfaces, one data spine, cross-reference engine (deterministic, no llm) and trace wrapping.
Every element must earn its place. The default answer to "should we add this?" is no. Reduction is not simplification — it's the discipline to remove everything that doesn't serve the user's immediate task.
Master multi-agent orchestration using Claude Code's TeammateTool and Task system. Use when coordinating multiple agents, parallel code reviews, pipeline workflows, or self-organizing task queues.
Turn messy input into a sourced report. 15 default tools: workflow, bidirectional sync, canonical NodeBench research bridge, discovertools, and loadtoolset. Runs locally from the nodebench-mcp npm package.