akitaonrails/llm-coding-benchmark

Simple benchmark to test the most popular open source and commercial LLMs with automated OpenCode

304Stars on the repository
2Mods indexed here, across every type
todayLast push, which is what freshness is scored on
noneNo LICENSE: all rights reserved, so bodies are not copied

benchmark-audit

01

akitaonrails/llm-coding-benchmark

Skill Claude CodeCodex

Automatically evaluates an LLM coding benchmark result using a standardized 0-100 rubric across 8 dimensions. Use when a benchmark finishes, when the user asks to evaluate a model, analyze a run's quality, or generate a score for a result. Also activates for 'score', 'evaluate benchmark', 'audit model', or 'analyze…

not rated 304 +6 today A 74 tokens