Skill Claude CodeCodex
I deploy the application to production via GitHub Actions.
724 tagged benchmark, measured the same way as everything else here.
Browse within: automatic 200continual-learning 200skill-generation 200skillsbench 197agent-evaluation 89ci 86llm-agents 84python-cli 83release-gate 83Evaluation 61bun 42game-development 39harness 39leaderboard 38
Skill Claude CodeCodex
I deploy the application to production via GitHub Actions.
Skill Claude CodeCodex
Walks a contributor through adding a new benchmark to OpenChainBench. Covers spec format, harness contract, Prometheus scrape wiring, local validation, and PR conventions.
Skill Claude CodeCodex
Find the best local LLM for your machine. Tests speed, quality and RAM fit, then tells you if a model is worth running on your hardware.
Yinzhanqing/embodied-eval-automation
Skill Claude CodeCodex
Plan, explain, build, run, monitor, validate, transfer, and audit reproducible embodied-model studies and batch episode collection. Use when a user wants to connect a local, SSH, or cloud GPU host; understand and compare a policy, VLA, world model, world-action model, or hybrid with a benchmark; discover and reuse…
Skill Claude CodeCodex
The end-to-end procedure for auditing one already-published post to a scientific-organization standard. It chains the adversarial skills and — critically — re-runs the auditor on the CORRECTED post to confirm it now passes clean before committing. We are a scientific organization; we do not ship missteps. Run EVERY…
Skill Claude CodeCodex
Persistent multi-agent economy where autonomous AI agents compete for resources, trade on a marketplace, and benchmark decision-making against a standing population of always-on agents. Invite other agents for energy rewards. Auto-registers — no API key needed.
Skill Claude CodeCodex
Run a standardized XORCISE agent benchmark — pick a ready-made playbook (or build your own from existing missions), point it at any OpenHands-supported model(s), and get a polished eval-card HTML report of how each model performed across the missions. Warns you about cost before spending anything.
Skill Claude CodeCodex
Use when the user asks to lint, audit, test, benchmark, validate, or fix an agent skill — from a single SKILL.md frontmatter check to trigger-accuracy measurement or a corpus-wide audit of installed skills for duplication — driving the shakespii CLI (init, lint --json, test --run, bench) to resolve findings until…
Skill Claude CodeCodex
Use this skill to perform actions inside an Unreal Engine project via a live-editor MCP connection. Trigger when the user wants to change, query, or run something in their Unreal Engine project, not for conceptual or docs questions. Concrete triggers: spawn/move/duplicate/transform actors in a level, open a .uproject…
Skill Claude CodeCodex
Independent adversarial security review of a custody-spine change against SECURITY.md. Spawns a FRESH reviewer subagent that did NOT write the code, prompted to break the diff (cap overspend, window-slide, early release, fail-open void-return, TOCTOU, dedup/replay bypass, honest-scope overclaim, DrainBench fairness…
Skill Claude CodeCodex
Use when the user wants to proactively push a project forward by implementing its next feature. Reads project artifacts (specs' "future work" sections, README, TODOs, recent commit trajectory, codebase gaps), infers candidate features, asks the user to pick one, then implements it on an isolated branch with…
Skill Claude CodeCodex
Maximum-savings variant of eco - the same frugality rules with a tighter reply budget, for routine chores (rename, small fix, quick lookup, boilerplate). Prefer plain eco for hard or high-stakes work. Works in any language.
Skill Claude CodeCodex
Build, launch, smoke-test and drive the InferBench FastAPI backend (:7777) and its inference-engine orchestration. Use when asked to run, start, launch, boot, smoke-test, benchmark, or drive the inferbench backend / API / engines (llama.cpp, ollama, vLLM, SGLang, TGI), or to verify an engine actually starts and runs a…
Skill Claude CodeCodex
Run, inspect, and compare multiple coding-agent CLIs on the same repository task with PatchLeague. Use when Codex should benchmark agents such as Codex, Claude Code, Gemini CLI, or OpenCode in isolated Git worktrees; verify their patches with tests; inspect saved runs or HTML reports; or preview and apply a selected…
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: