allenai

4 mods across 2 repositories, 921 stars between them.

add-benchmark

01

allenai/vla-evaluation-harness

Skill Claude CodeCodex

Add a new simulation benchmark to the VLA evaluation harness. Use this skill whenever the user wants to integrate, create, or add a new benchmark or simulation environment — e.g. 'add ManiSkill3', 'integrate OmniGibson', 'hook up a new sim'. Also use when they ask how benchmarks are structured or want to understand…

573 8d ago A 78 tokens original Apache-2.0

add-model-server

02

allenai/vla-evaluation-harness

Skill Claude CodeCodex

Add a new VLA model server to the evaluation harness. Use this skill whenever the user wants to integrate, create, or add a new model — e.g. 'add OpenVLA server', 'integrate RT-2', 'hook up my model', 'write a model server'. Also use when they ask how model servers work or want to understand the server interface.

573 8d ago A 80 tokens original Apache-2.0

run-evaluation

03

allenai/vla-evaluation-harness

Skill Claude CodeCodex

Run a VLA model evaluation against a simulation benchmark. Use this skill whenever the user wants to evaluate, benchmark, test, or run a model on a sim environment — even if they say it casually like 'try OpenVLA on LIBERO' or 'get me CALVIN scores'. Covers the full workflow: serving the model, launching the…

573 8d ago A 87 tokens original Apache-2.0