tmuskal

10 mods across 1 repository, 3 stars between them.

benchmark-adder

02

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

Given a benchmark repository URL (agentic envs like arc-agi, memory/eval benchmarks like longmemeval, QA/code/tool-use benchmarks, etc.), orchestrate the creation of a full Claude Code plugin that benchmarks the current harness setup against it. Wraps the babysitter:babysit skill with the benchmark-plugin-creator…

3 4mo ago A 74 tokens

browse-tests

03

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

Explore available ARC-AGI environments - lists games, shows details with ASCII grid visualization, and displays historical scores.

3 4mo ago A 22 tokens

compare-runs

04

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

Compare two or more ARC-AGI benchmark runs - shows score deltas, config changes, and trends to track improvement or regression.

3 4mo ago A 26 tokens

cross-harness

05

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

Cross-harness benchmarking - generate instructions for Codex/Gemini/OpenCode, import results, and compare across harnesses.

3 4mo ago A 24 tokens

report

06

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

Generate and display comprehensive reports from completed ARC-AGI benchmark runs - shows scores, per-game breakdowns, and performance analysis.

3 4mo ago A 25 tokens

run-benchmark

07

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

Execute benchmark runs against ARC-AGI games - plays games with Claude Code as the agent and records scores.

3 4mo ago B 21 tokens

setup

08

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

Set up the ARC-AGI benchmarking environment - installs dependencies, configures API access, and verifies the setup works.

3 4mo ago C 23 tokens

judge

09

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

LLM-as-judge shim for LongMemEval - wraps upstream getanscheckprompt and calls Anthropic (default) or OpenAI (fallback) with exponential backoff.

3 4mo ago A 34 tokens

resume

10

tmuskal/arc-agi-benchmarker

Skill Claude CodeCodex

Detect an incomplete LongMemEval run and continue it from the last checkpoint.

3 4mo ago A 14 tokens