benchmarking plugins

15 tagged benchmarking, measured the same way as everything else here.

tonyblu331/research-proof

Plugin Claude Code

Falsifiable research planning for AI agents: pressure-test claims, freeze evaluators, use AI-lab and medical evidence patterns, build proof ladders, and maintain proof ledgers.

45 3mo ago A tokens not measured original MIT

business-consulting

04

abinauv/business-consulting

Plugin Claude Code

Full-stack management consulting toolkit — 16 skills covering market research, competitive analysis, financial modeling, strategy frameworks, operations, benchmarking, data analysis, deliverables, change management, pricing, M&A, customer insights, digital transformation, risk management, talent strategy, and…

25 6mo ago A tokens not measured original MIT

vukkt-plugins

05

vukkt/token-warden

Plugin Claude Code

Plugins by Vuk Topalovic — currently token-warden, a measurement and selection layer that makes Claude Code subagents measurably cheaper over time.

13 4d ago A tokens not measured original MIT

token-warden

06

vukkt/token-warden

Plugin Claude Code

Agent memory, charged rent. A Stop hook records what each session costs; expensive runs are distilled into candidate efficiency rules; each candidate is benchmarked on a frozen golden suite and kept only if it saves at least twice its context rent. Six commands. Two theorems survived measurement and three were deleted.

13 4d ago A tokens not measured original MIT

proyecto26/autoresearch-ai-plugin

Plugin Claude Code

Autonomous experiment loop for Claude Code. Run an endless optimize-measure-keep/discard cycle on any optimization target: LLM training loss, test speed, bundle size, build time, and more.

12 1mo ago A tokens not measured original MIT

mcpbr

09

greynewell/mcpbr

Plugin Claude Code

mcpbr - MCP Benchmark Runner plugin marketplace.

10 4mo ago A tokens not measured original MIT

mcpbr

10

greynewell/mcpbr

Plugin Claude Code

Expert benchmark runner for MCP servers using mcpbr. Handles Docker checks, config generation, and result parsing.

10 4mo ago A tokens not measured original MIT

widget-fixture

12

yaniv-golan/skill-creator-plus

Plugin Claude Code

Fixture plugin exercising the script-invocation stanzas in assets/. Not shipped; harness only.

4 2d ago A tokens not measured original MIT

skill-creator-plus

13

yaniv-golan/skill-creator-plus

Plugin Claude Code

Create, test, evaluate, and iteratively improve Claude Code skills with built-in benchmarking, blind A/B comparison, and description optimization.

4 2d ago A tokens not measured original MIT

performance-deity

14

v0idOS/performance-deity

Plugin Claude Code

A marketplace of benchmark-driven performance engineering skills for Claude Code.

2 4mo ago A tokens not measured original MIT

performance-deity

15

v0idOS/performance-deity

Plugin Claude Code

Nine benchmark-driven performance engineering skills for Claude Code. Adds CPU optimization, memory leak detection, race condition testing, database query analysis, chaos engineering, network payload reduction, CI/CD acceleration, telemetry instrumentation, and UI frame-rate enforcement.

2 4mo ago A tokens not measured original MIT