microsoft/a11y-llm-eval

An eval tool to benchmark how well LLMs generate accessible HTML

About the project

A11y LLM Evaluation Harness and Dataset is a test suite for measuring how well language models generate accessible HTML. It runs agent-generated pages in a browser and evaluates them with automated accessibility checks.

These files are microsoft/a11y-llm-eval's own configuration. They tell Claude Code and GitHub Copilot how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

59Stars on the repository
2Files it configures its agents with
533Tokens loaded in every session
2Agents configured

Instructions

Agents