cxcscmu/SkillLearnBench

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

83Stars on the repository
200Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

cxcscmu/SkillLearnBench

Skill Claude CodeCodex

Covers descriptive statistics (mean, median, percentiles) and weighted averages using Excel formulas.

not rated 83 +1 2mo ago A 24 tokens original MIT

cxcscmu/SkillLearnBench

Skill Claude CodeCodex

Guidelines for applying strict brand color palettes and minimalist technical design standards to generated imagery.

not rated 83 +1 2mo ago A 22 tokens original MIT

poetic-composition

196

cxcscmu/SkillLearnBench

Skill Claude CodeCodex

Guidelines and structures for composing classical Chinese poetry, specifically seven-character regulated verse (Qi-Yan Lu-Shi).

not rated 83 +1 2mo ago A 28 tokens original MIT

pdf-filler

197

cxcscmu/SkillLearnBench

Skill Claude CodeCodex

Provides guidelines for filling out PDF forms using programming tools.

not rated 83 +1 2mo ago A SkillSpector: pass 15 tokens original MIT

csv-reporting

198

cxcscmu/SkillLearnBench

Skill Claude CodeCodex

Creating structured CSV reports for security findings.

not rated 83 +1 2mo ago A SkillSpector: pass 12 tokens original MIT

trivy-audit

199

cxcscmu/SkillLearnBench

Skill Claude CodeCodex

Performing offline security audits on package-lock.json files using Trivy.

not rated 83 +1 2mo ago A 19 tokens original MIT

geospatial-processing

200

cxcscmu/SkillLearnBench

Skill Claude CodeCodex

Provides foundational techniques for loading, projecting, and manipulating geospatial datasets using GeoPandas.

not rated 83 +1 2mo ago A 22 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: