benchmarking agents

14 tagged benchmarking, measured the same way as everything else here.

Browse within: Apple Silicon 6agent-skill 6machine-learning 6mlx 6

backend

01

vukkt/token-warden

Agent

Backend specialist — Node.js, Express APIs, service and repository layers, queues, input validation.

13 4d ago A 21 tokens original MIT

frontend

02

vukkt/token-warden

Agent

Frontend specialist — React, TypeScript, components, hooks, context, client-side data fetching.

13 4d ago A 21 tokens original MIT

sql

03

vukkt/token-warden

Agent

SQL specialist — schema design, indexes, query optimization, migrations, eliminating table scans and N+1 patterns.

13 4d ago A 24 tokens original MIT

proyecto26/autoresearch-ai-plugin

Agent

Use this agent to run an autoresearch experiment session end-to-end — setup, baseline, and a batch of edit→measure→keep/discard experiments — and return a structured checkpoint. Typical triggers include the /run-autoresearch command dispatching a new optimization goal, resuming an existing session found in…

12 1mo ago A 115 tokens original MIT

coverage-skeptic

05

Amal-David/mlx-porting-skill

Agent

Look for blind spots, unsupported architecture families, missing validation gates, and overclaimed optimizations.

5 3d ago A 0 tokens original Apache-2.0

analyzer

08

yaniv-golan/skill-creator-plus

Agent

Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.

4 2d ago A 0 tokens copy · 97% MIT