benchmark plugins

23 tagged benchmark, measured the same way as everything else here.

codspeed

02

CodSpeedHQ/codspeed

Plugin Claude Code

CodSpeed plugin for Claude Code helping with performance measurement and optimization.

280 4d ago A tokens not measured original Apache-2.0

spmind

03

tomtommyyuan/spmind

Plugin Claude Code

Plugin marketplace listing 1 plugin: spmind.

277 1mo ago A tokens not measured original Apache-2.0

spmind-plugin

04

tomtommyyuan/spmind

Plugin Claude Code

Spatial proteomics analysis tools (illumination correction, registration, TMA dearraying, segmentation, quantification, clustering) exposed as MCP tools.

277 1mo ago A tokens not measured original Apache-2.0

slopkit

05

ehmo/slopkit

Plugin Claude Code

Plugin marketplace listing 2 plugins: slopbeth, slopgent.

94 1mo ago A tokens not measured original MIT

slopbeth

06

ehmo/slopkit

Plugin Claude Code

Remove AI-writing tells while preserving meaning, voice, and density.

94 1mo ago A tokens not measured original MIT

slopgent

07

ehmo/slopkit

Plugin Claude Code

Shape the agent's own replies so they are honest about what ran, action-first, and plain — without dropping load-bearing precision.

94 1mo ago A tokens not measured original MIT

skill-optimizer

08

fastxyz/skill-optimizer

Plugin Claude Code

Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs.

77 3mo ago A tokens not measured original MIT

skill-optimizer

09

fastxyz/skill-optimizer

Plugin Claude Code

Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs.

77 3mo ago A tokens not measured original MIT

modelharness

12

vitaliikapliuk/modelharness

Plugin Claude Code

Fable 5 behavioral patterns as a harness for Opus 4.8: grounded progress, self-verification loops, delegation triggers, file-based memory, autonomy calibration. Zero-config: installs a SessionStart hook; skills auto-trigger.

38 2mo ago A tokens not measured original MIT

many-ppt-skills

14

brycewang-stanford/many-ppt-skills

Plugin Claude Code

Routes to the right AI slide-deck skill and names a style id you can ask for, backed by sample imagery read from the projects' own repositories.

37 2d ago A tokens not measured

folio-ops

15

jfrog/agent-belt

Plugin Claude Code

Folio operations pack: bundles a plugin command (lookup) and a plugin skill (inventory-audit), both reaching the workspace's Folio MCP server.

18 1mo ago A tokens not measured original Apache-2.0

orders-pack

16

jfrog/agent-belt

Plugin Claude Code

Bundles a slash command (report-orders) and a skill (processing-watch) that both reach the workspace's ordersdb MCP server.

18 1mo ago A tokens not measured original Apache-2.0

metrillm

17

MetriLLM/metrillm

Plugin Claude Code

Benchmark local LLM models — performance, quality & hardware fitness verdict.

5 3mo ago A tokens not measured original Apache-2.0

agent-ste

18

abryfs/agent-ste

Plugin Claude Code

Plugin marketplace listing 1 plugin: agent-ste.

4 21d ago A tokens not measured original MIT

agent-ste

19

abryfs/agent-ste

Plugin Claude Code

Write technical and agent-facing text in ASD-STE100 Simplified Technical English. Benchmarked: 95.5% fewer STE violations than baseline across 12 models.

4 21d ago A tokens not measured original MIT

ade-bench

21

typedef-ai/ade-bench-plugin

Plugin Claude Code

Generate ADE-Bench benchmark tasks from your own dbt project. Scans your models, proposes realistic bug-injection scenarios, and writes the task scaffolding (config, patches, scripts, custom assertion tests) ready to run against AI agents.

3 3mo ago A tokens not measured original MIT

distil

22

munhq/distil

Plugin Claude Code

Context optimization for LLM agents: measure where a session's tokens actually go, and compress context only where compression pays for the prompt cache it invalidates.

1 3d ago A tokens not measured

distil

23

munhq/distil

Plugin Claude Code

Skills and MCP server for context optimization — decide whether compressing an agent's context pays for the prompt cache it invalidates, and compress it when the answer is yes. Bundles the distil MCP server so install wires everything in one step.

1 3d ago A tokens not measured