evaluation commands

44 tagged evaluation, measured the same way as everything else here.

Browse within: browser-automation 21code-agent 10harness 10llm-as-judge 6plugin 6quality 6

integrate-pipeline

02

run-llama/ParseBench

Command Claude Code

Integrate a new document parsing pipeline into ParseBench: $ARGUMENTS.

554 3d ago A 0 tokens original Apache-2.0

hegelion

03

Hmbown/Hegelion

Command Claude Code

Command "hegelion" from Hmbown/Hegelion, covering /hegelion, routing and autocoding loop.

171 5mo ago A 0 tokens original MIT

harness-gate

04

redhat-community-ai-tools/harness-eval

Command Cursor

Gate the agent setup on corpus-validated rules (gating tier). Fast, no LLM, exits nonzero on any finding. Suitable for CI and pre-commit.

27 5d ago A 0 tokens original Apache-2.0

harness-review

05

redhat-community-ai-tools/harness-eval

Command Cursor

Full qualitative review of the agent setup. Read every file, evaluate quality, redundancy, and optimization opportunities. Produce KEEP/REVIEW/REMOVE verdicts per component.

27 5d ago A 0 tokens original Apache-2.0

skill-review

06

redhat-community-ai-tools/harness-eval

Command Cursor

Deep-evaluate a single skill with static analysis and qualitative review, both individually and in context of the full setup.

27 5d ago A 0 tokens original Apache-2.0

arch-diff

07

JuanMarchetto/agent-skills

Command

Quick architecture drift summary β€” counts drift items by severity without producing a full report.

5 5mo ago A 19 tokens original MIT

arch-init

08

JuanMarchetto/agent-skills

Command

Generate an ARCHITECTURE.md from scanning the current codebase β€” shows draft for user approval before saving.

5 5mo ago A 24 tokens original MIT

review-arch

09

JuanMarchetto/agent-skills

Command

Run a full architecture drift audit β€” finds architecture docs, scans code, compares against implementation, and produces a detailed drift report.

5 5mo ago A 28 tokens original MIT

against

10

sattyamjjain/proofloop

Command

Side-by-side delta between two scorecards of the same skill.

5 2mo ago A 14 tokens original MIT

judge

12

sattyamjjain/proofloop

Command

Evaluate the execution quality of a skill or agent.

5 2mo ago A 11 tokens original MIT

proofrag

13

unshDee/proofrag

Command

Evaluate a RAG/LLM app β€” generate a golden set, judge it, and produce a scorecard.

2 22d ago A 22 tokens original MIT

skill-vetting

14

sina-heidariaan/cold-run

Command Claude Code

Vet a rule, principle, or heuristic with the Delta Test v2 and decide if it can become a real Claude skill or is a platitude that should die.

0 9d ago A 33 tokens

vet-my-skills

15

sina-heidariaan/cold-run

Command

Cold-run every skill in a directory against a baseline agent and report which ones change nothing, which make the output worse, and which earn their place.

0 9d ago A 0 tokens

vet-routing

16

sina-heidariaan/cold-run

Command

Test whether each skill's description actually gets that skill loaded, by routing written requests against descriptions alone.

0 9d ago A 20 tokens