One framework to evaluate any VLA model on any robot simulation benchmark.
Latest release v0.5.0 · 24 Aug 2026
These files are allenai/vla-evaluation-harness's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
CLAUDE.md A 1,298 tok .claude/skills/add-benchmark/SKILL.md A 78 tok .claude/skills/add-model-server/SKILL.md A 80 tok .claude/skills/run-evaluation/SKILL.md A 87 tok