benchmark commands

72 tagged benchmark, measured the same way as everything else here.

Browse within: Apple Silicon 59inference 59lm-studio 59

integrate-pipeline

01

run-llama/ParseBench

Command Claude Code

Integrate a new document parsing pipeline into ParseBench: $ARGUMENTS.

554 3d ago A 0 tokens original Apache-2.0

goal

02

vitaliikapliuk/modelharness

Command

Extract a full task specification up front (goal, constraints, definition of done), then execute — the primary documented lever for Opus 4.8 long-horizon quality.

38 2mo ago A 35 tokens original MIT

retro

03

vitaliikapliuk/modelharness

Command

Write this session's lessons into the project's lessons/ directory using the memory-discipline format.

38 2mo ago A 18 tokens original MIT

verify

04

vitaliikapliuk/modelharness

Command

Spawn a fresh-context verifier subagent to check completed work against its specification before trusting it.

38 2mo ago A 18 tokens original MIT

refund

05

jfrog/agent-belt

Command Claude Code

Issue a refund for a Folio order by id with a reason. Calls the Folio MCP server's refundorder tool directly and emits a strict pipe-separated reply that downstream parsers can consume.

18 1mo ago A 39 tokens original Apache-2.0

lookup

06

jfrog/agent-belt

Command Claude Code

Look up a Folio customer by id and report their order count, in a strict pipe-separated format that downstream parsers can consume.

18 1mo ago A 27 tokens original Apache-2.0

report-orders

07

jfrog/agent-belt

Command Claude Code

Generate a one-line aggregate report of every order in ordersdb, in a strict colon-separated format that downstream parsers can consume.

18 1mo ago A 26 tokens original Apache-2.0

config.fr

08

druide67/asiai

Command

Comment configurer asiai : gérer les URL des moteurs, les ports et les paramètres persistants pour votre configuration de benchmark LLM sur Mac.

11 2d ago A 29 tokens original Apache-2.0

config.ja

09

druide67/asiai

Command

A saved settings manager for the local programs that run language models, including their addresses, ports, versions and labels.

11 2d ago A 33 tokens original Apache-2.0

daemon.de

10

druide67/asiai

Command

Hintergrunddienste über macOS launchd LaunchAgents verwalten.

11 2d ago A 28 tokens original Apache-2.0

adr

11

PavelTkachenk0/ContextTax

Command Claude Code

Scaffold a new Architecture Decision Record in the memory bank.

5 2mo ago A 11 tokens original MIT

create-task

13

typedef-ai/ade-bench-plugin

Command

Automatically scan a dbt project and generate ADE-Bench benchmark tasks via pattern matching.

3 3mo ago A 16 tokens original MIT

plan-tasks

14

typedef-ai/ade-bench-plugin

Command

Interactively plan and generate ADE-Bench benchmark tasks from a dbt project (recommended).

3 3mo ago A 18 tokens original MIT

setup

15

typedef-ai/ade-bench-plugin

Command

Install or verify the ADE-Bench harness at /.ade-bench (clones repo, installs CLI, downloads bundled DuckDB databases).

3 3mo ago A 27 tokens original MIT