Scenarios and tasks for evaluating AI agents on real-world AWS automation tasks. Use the aws-bench framework (https://github.com/aws-bench/aws-bench) to provision isolated AWS environments, run agents, and evaluate their results.
These files are aws-bench/aws-bench-datasets's own configuration. They tell Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 1,228 tok CLAUDE.md A 5 tok