Command Cursor
π§ Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 π.
44 tagged evaluation, measured the same way as everything else here.
Browse within: browser-automation 21code-agent 10harness 10llm-as-judge 6plugin 6quality 6
Command Cursor
π§ Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 π.
Command Claude Code
Integrate a new document parsing pipeline into ParseBench: $ARGUMENTS.
Command Claude Code
Command "hegelion" from Hmbown/Hegelion, covering /hegelion, routing and autocoding loop.
redhat-community-ai-tools/harness-eval
Command Cursor
Gate the agent setup on corpus-validated rules (gating tier). Fast, no LLM, exits nonzero on any finding. Suitable for CI and pre-commit.
redhat-community-ai-tools/harness-eval
Command Cursor
Full qualitative review of the agent setup. Read every file, evaluate quality, redundancy, and optimization opportunities. Produce KEEP/REVIEW/REMOVE verdicts per component.
redhat-community-ai-tools/harness-eval
Command Cursor
Deep-evaluate a single skill with static analysis and qualitative review, both individually and in context of the full setup.
Command
Quick architecture drift summary β counts drift items by severity without producing a full report.
Command
Generate an ARCHITECTURE.md from scanning the current codebase β shows draft for user approval before saving.
Command
Run a full architecture drift audit β finds architecture docs, scans code, compares against implementation, and produces a detailed drift report.
Command
Side-by-side delta between two scorecards of the same skill.
Command
Compare skill scores against ideal benchmarks.
Command
Evaluate the execution quality of a skill or agent.
Command
Evaluate a RAG/LLM app β generate a golden set, judge it, and produce a scorecard.
Command Claude Code
Vet a rule, principle, or heuristic with the Delta Test v2 and decide if it can become a real Claude skill or is a platitude that should die.
Command
Cold-run every skill in a directory against a baseline agent and report which ones change nothing, which make the output worse, and which earn their place.
Command
Test whether each skill's description actually gets that skill loaded, by routing written requests against descriptions alone.