Skill Claude CodeCodex
Run a standardized XORCISE agent benchmark — pick a ready-made playbook (or build your own from existing missions), point it at any OpenHands-supported model(s), and get a polished eval-card HTML report of how each model performed across the missions. Warns you about cost before spending anything.
2 28d ago B 66 tokens