awslabs/llm-evaluation-system

Agentic AI-guided evaluation system for comparing LLMs with multi-judge jury scoring

About the project

LLM Evaluation System is an agent-guided platform for evaluating language models and agents, generating datasets and configuring multiple judges from natural-language requests before producing an analysis report. It is for comparing model responses, testing agents, and creating document-grounded evaluation data.

These files are awslabs/llm-evaluation-system's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

23Stars on the repository
5Files it configures its agents with
7,551Tokens loaded in every session
3Agents configured

Instructions

Skills