leaderboard

leaderboard is a skill for Claude Code from cukas/claudes-ai-buddies. It costs 14 tokens per session (502 once invoked), scanned A, original, MIT.

A tool that displays persistent ELO ratings for coding tasks. ELO is a scoring system that changes when one solution or agent wins or loses against another.

In plain words
What is it for?
Use it to show ratings overall or by task class, including algorithms, refactoring, bug fixes, features, tests, and documentation. Ratings begin building after the first forge run.
Why use it?
It provides a running view of comparative performance instead of relying only on individual results. Ratings can be viewed for all task types or for a specific class such as bug fixes or algorithms.

Skill for Claude Code

Written for Claude Code: ${CLAUDE_PLUGIN_ROOT} variable.

Runs only inside its plugin — its command needs a path that Claude Code sets for a plugin’s own hooks and for nothing else. Install the plugin, not this.

Part of the claudes-ai-buddies plugin — 12 skills, 1 command, 1 hook shipped together

Good fit Use it to show ratings overall or by task class, including algorithms, refactoring, bug fixes, features, tests, and documentation. Ratings begin building after the first forge run.

Compare 6 skills from other repositories ↓
Install

Getting it into your agent

This one installs as part of its plugin. Adding the marketplace and installing the plugin brings it with everything else the plugin ships.

Claude Code
/plugin marketplace add cukas/claudes-ai-buddies
Claude Code
/plugin install claudes-ai-buddies

Made for: Claude Code.

Or install claudes-ai-buddies, the plugin that ships this one along with the rest of its 12 skills, 1 command, 1 hook.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for leaderboard

README.md
[![agentmods](https://agentmods.dev/badge/skills/cukas/claudes-ai-buddies/leaderboard/github.svg)](https://agentmods.dev/skills/cukas/claudes-ai-buddies/leaderboard)
Your own site
<a href="https://agentmods.dev/skills/cukas/claudes-ai-buddies/leaderboard"><img src="https://agentmods.dev/badge/skills/cukas/claudes-ai-buddies/leaderboard/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for leaderboard

Your own site · 80×15
<a href="https://agentmods.dev/skills/cukas/claudes-ai-buddies/leaderboard"><img src="https://agentmods.dev/badge/skills/cukas/claudes-ai-buddies/leaderboard.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 14 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 502 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00014 $0.00502
Opus 5 $0.00007 $0.00251
Sonnet 5 $0.00003 $0.00100
Haiku 4.5 $0.00001 $0.00050

Measured 11d ago against content hash 41d6e8fec8d6, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

leaderboard scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/leaderboard/SKILL.md · 67 lines

What it actually says

/leaderboard — ELO Ratings

Show the persistent ELO ratings leaderboard. Ratings are updated after each /forge run based on who wins and loses.

How to invoke

/leaderboard              # show all task classes
/leaderboard algorithm    # show only algorithm class
/leaderboard bugfix       # show only bugfix class

Step-by-step workflow

  1. Parse optional task class from the user's message.
  2. Run the leaderboard formatter:
# All classes
bash "${CLAUDE_PLUGIN_ROOT}/scripts/elo-show.sh"

# Specific class
bash "${CLAUDE_PLUGIN_ROOT}/scripts/elo-show.sh" --task-class "algorithm"
  1. Present the output to the user. If no data exists yet, explain that ratings start building after the first /forge run.

Task classes

Ratings are tracked per task class (auto-detected from the forge task description):

Class Keywords
algorithm algorithm, sort, search, scoring, math, compute
refactor refactor, rename, extract, simplify, reorganize
bugfix fix, bug, error, crash, broken, regression
feature add, implement, create, build, feature, new
test test, spec, coverage, assert
docs doc, readme, comment, changelog
other (default)

How ELO works

  • All buddies start at 1200
  • K-factor: 32 (configurable)
  • Winner gains points, loser loses points (zero-sum)
  • Ratings below 100 are floored
  • Provisional status for < 10 games

Configuration

Key Default Description
elo_enabled true Enable/disable ELO tracking
elo_k_factor 32 ELO K-factor (higher = more volatile)

Rules

  • Read-only. This skill only displays ratings, never modifies them.
  • Suggest /forge if no data exists yet.
  • Keep it brief. Just show the table, no analysis unless asked.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 11d ago First seen · 67 lines · 14 tokens per session scan A 41d6e8fec8d6

Subscribe to this mod's changes

leaderboard is a skill published in the GitHub repository cukas/claudes-ai-buddies (9 stars, last pushed 4mo ago), licensed MIT. It adds 14 tokens to every session and 502 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

optimize-claude-opus-5-prompts

Clarify, audit, and rewrite rough or existing prompts for Claude Opus 5 using Anthropic's official prompting guidance. Use when a user asks to optimize, improve, migrate, debug, or design a prompt, system prompt, or agent harness for Claude Opus 5 (claude-opus-5), including requests phrased as Opus 5 prompt…

IchenDEV/prompt-optimizer-plugins · 182 tokens

optimize-claude-fable-5-prompts

A prompt-editing method for Claude Fable 5, an AI model. It turns rough or existing instructions into clearer prompts while keeping the original goal and important limits.

IchenDEV/prompt-optimizer-plugins · 137 tokens

optimize-gpt-5-6-prompts

A guide for improving prompts written for OpenAI’s GPT-5.6 models, including GPT-5.6 Sol.

IchenDEV/prompt-optimizer-plugins · 135 tokens

optimize-gpt-6-astra-prompts

Clarify, audit, migrate, and rewrite prompts for GPT-6 Astra using OpenAI's official latest-model prompting guidance. Use when a user asks to optimize, improve, rewrite, debug, migrate, or design a GPT-6 Astra prompt, including requests phrased as GPT-6, Astra, or gpt-6-astra prompt optimization, or when an…

IchenDEV/prompt-optimizer-plugins · 134 tokens

optimize-deepseek-v4-prompts

Clarify, audit, and rewrite rough or existing prompts specifically for DeepSeek-V4 Pro and Flash using the DeepSeek-AI V4 technical report as the model-specific evidence base. Use when a user asks to optimize, improve, rewrite, debug, migrate, or design a DeepSeek-V4 prompt, including requests phrased as DeepSeek…

IchenDEV/prompt-optimizer-plugins · 151 tokens

optimize-gemini-prompts

A prompt-editing skill for clarifying and rewriting prompts intended for Google’s Gemini models.

IchenDEV/prompt-optimizer-plugins · 126 tokens