Plugin Claude Code
Plugin marketplace listing 2 plugins: arc-agi-benchmarker, longmemeval-benchmarker.
Plugin Claude Code
Plugin marketplace listing 2 plugins: arc-agi-benchmarker, longmemeval-benchmarker.
Skill Claude CodeCodex
Given a benchmark repository URL (agentic envs like arc-agi, memory/eval benchmarks like longmemeval, QA/code/tool-use benchmarks, etc.), orchestrate the creation of a full Claude Code plugin that benchmarks the current harness setup against it. Wraps the babysitter:babysit skill with the benchmark-plugin-creator…
Skill Claude CodeCodex
Explore available ARC-AGI environments - lists games, shows details with ASCII grid visualization, and displays historical scores.
Skill Claude CodeCodex
Compare two or more ARC-AGI benchmark runs - shows score deltas, config changes, and trends to track improvement or regression.
Skill Claude CodeCodex
Cross-harness benchmarking - generate instructions for Codex/Gemini/OpenCode, import results, and compare across harnesses.
Skill Claude CodeCodex
Generate and display comprehensive reports from completed ARC-AGI benchmark runs - shows scores, per-game breakdowns, and performance analysis.
Skill Claude CodeCodex
Execute benchmark runs against ARC-AGI games - plays games with Claude Code as the agent and records scores.
Skill Claude CodeCodex
Set up the ARC-AGI benchmarking environment - installs dependencies, configures API access, and verifies the setup works.
Skill Claude CodeCodex
LLM-as-judge shim for LongMemEval - wraps upstream getanscheckprompt and calls Anthropic (default) or OpenAI (fallback) with exponential backoff.
Skill Claude CodeCodex
Detect an incomplete LongMemEval run and continue it from the last checkpoint.