awslabs/llm-evaluation-system

Agentic AI-guided evaluation system for comparing LLMs with multi-judge jury scoring

23Stars on the repository
5Mods indexed here, across every type
13d agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

aws-architecture

01

awslabs/llm-evaluation-system

Skill Claude CodeCodex ✓ vendor

Use when the user asks to create, update, or render an AWS architecture diagram (or a cloud/infrastructure diagram using AWS service icons) — e.g. "diagram our AWS setup", "architecture diagram with CloudFront/S3/EKS/RDS", "update docs/image.png", "make an AWS infra diagram". Produces a clean, poster-style diagram…

23 13d ago A 109 tokens original Apache-2.0

port-a-benchmark

02

awslabs/llm-evaluation-system

Skill Claude CodeCodex ✓ vendor

Port an existing third-party benchmark or eval into this repo as an Inspect AI task, faithfully. Use this whenever the task is "add benchmark X", "port this eval", "can we run here", or adapting any external eval harness (aiewf-eval, lm-evaluation-harness, a paper's repo, a colleague's script). Enforces the one rule…

23 13d ago A 118 tokens original Apache-2.0

ship-it

03

awslabs/llm-evaluation-system

Skill Claude CodeCodex ✓ vendor

Ship work in the llm-evaluation-system repo end-to-end — commit with conventional-commit messages, push to a feature branch (never directly to main), open a PR with the proper title format, and after merge either run make release to publish to PyPI or just clean up branches. Use this skill whenever the user says…

23 13d ago C 131 tokens original Apache-2.0