Agentic AI-guided evaluation system for comparing LLMs with multi-judge jury scoring
LLM Evaluation System is an agent-guided platform for evaluating language models and agents, generating datasets and configuring multiple judges from natural-language requests before producing an analysis report. It is for comparing model responses, testing agents, and creating document-grounded evaluation data.
These files are awslabs/llm-evaluation-system's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 102 tok CLAUDE.md A 7,449 tok .claude/skills/aws-architecture/SKILL.md A 109 tok .claude/skills/port-a-benchmark/SKILL.md A 118 tok .claude/skills/ship-it/SKILL.md C 131 tok