Proofpunk — execution-first delivery plugin for Claude Code, oh-my-pi (OMP), and OpenCode. 18 skills where done means proven by end-user testing, plus a 20-theme flat-black cyberpunk pack.
Structured solution brainstorming with trade-off analysis and brutal honesty: scans the codebase BEFORE asking questions, pins exact requirements (expected output, acceptance criteria, scope boundary constraints, touchpoints), presents 2-3 approaches with pros/cons in visible text, and writes no code until the user…
Conduct an evidence-backed, end-to-end audit and safe remediation of any software repository. Reconstructs change intent from history, verifies code, configuration, documentation, runtime behavior, dependencies, and production readiness; identifies drift, dead code, stale docs, cleanup risks, and remediation options…
The proof standard for end-user testing: every completion claim is proven by driving the real system as the end user, with run-scoped fresh evidence (timestamped, sequential, non-empty, never reused across runs), full-path citations describing what is SEEN, personally examined proof before any task is marked done…
App-wide functional audit of a live site or app that inventories EVERY user interaction — every screen, page, button, form, link, endpoint, and flow — then clicks through everything and validates each one against the real running system, remediates failures immediately, and revalidates until clean. Five phases…
The single write path: scouts the real codebase (mandatory, subagents) distills TRUE success criteria, mines past sessions (session-intent) forges the build prompt (prompt-forge), decomposes into a task graph where every task carries a proof obligation, then runs the execution loop — implement a task, drive the…
Second-pass plan strengthening: red-teams a draft plan through multiple adversarial lenses, scores confidence gaps, researches weak sections injects or strengthens proof obligations, and surgically remediates findings while preserving the original intent. Also converts arbitrary prompts or plans into proof-carrying…
Take a codebase from 'works on my machine' to shippable — systematic 8-phase production-readiness audit (risk-based cleanup waves, dead code documentation drift, zero-regression enforcement), spec-vs-implementation compliance audits that find COVERED/INCOMPLETE/MISSING gaps, and dependency supply-chain health (CVEs…
Prompt engineering, rating, and pipeline design: authors high-quality prompts on the canonical XML tag skeleton, rates any prompt against a quantitative rubric with test cases and metrics, optimizes weak prompts against real failure evidence, and builds multi-stage meta-prompt pipelines (.prompts/ directories…
Router and entry point for the proofpunk plugin's 17 delivery skills — reads a request, names the single best-fit skill (or short ordered chain for compound asks), and hands off without repeating that skill's own doctrine. Covers the full arc: brainstorm a design, plan or harden a multi-phase build, implement it end…
Attack your own plans, prompts, and outputs before reality does — 4-lens hostile review (security, scope-creep, evidence-rigor, failure-modes) against plans/prompts/artifacts, formal eval-driven development (EDD) scoring agent sessions against rubrics, QA cycling loops (test, verify fix, repeat until goal met), and…
Find and fix the real cause of bugs, never the symptom — disciplined reproduce-minimize-hypothesize-instrument loops for hard bugs and performance regressions, backward call-stack tracing to the original trigger, test-pollution bisection with a find-polluter script, expert investigation protocols, and…
Reconstruct what was actually ASKED from the sessions themselves — parses Claude Code JSONL transcripts into a per-session intent matrix (first user prompt = stated intent, subsequent prompts = steering, tool calls, files touched, commits made), aligns sessions to git history, and builds intent-vs-implementation…
Real-system test discipline per stack — pytest/Go/C++/Django/Spring Boot gotchas that cause flaky CI, FastAPI HTTP/SSE testing with curl Playwright browser automation with server lifecycle management, and condition-based waiting to kill timing flakes. Use when writing debugging, or deflaking test suites in Python, Go…
End-user proof for terminal UIs (Ink, blessed, textual, ratatui, curses): drive the real TUI in a real PTY as the end user with observe-then-act discipline, matched-assertion waits, three-facet evidence (screen + disk + logs), and pixel proof for visual claims. Codifies the measured lessons of driving agent-tty…
Deep end-to-end audit of any UI screen across four dimensions — visual defects, interactive elements, content quality, and UX heuristics — ending in one severity-classified report. Inventories every action item (button link, field, gesture), verifies each is discoverable and reachable, audits prose / code-block /…
Authors multi-phase project plans where every phase carries blocking cumulative proof obligations — BRIEF, ROADMAP, per-phase PLAN, SUMMARY + VALIDATION with run-scoped evidence. Proofs are cumulative: phase N's validation re-verifies phases 1..N-1, so a regression in earlier work blocks advancement. Use when asked to…
Mandatory visual QA protocol for UI screenshots — iOS (Apple HIG), web (WCAG 2.2), and cross-platform. Evaluates layout, overflow, spacing typography, contrast, touch targets, dark mode, and visual hierarchy against universal and platform-specific checklists, with severity classification and a defect-pattern database.…
★not rated 4 yesterdayA139 tokens
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: