scoring skills

29 tagged scoring, measured the same way as everything else here.

Browse within: quality 10api 8openapi 8score 8scorecard 8Evaluation 5

by-scoring

01

001TMF/blatant-why

Skill Claude CodeCodex

Interpret and apply BY custom scoring metrics for protein and antibody design. This skill covers ipSAE (interface Predicted Structural Accuracy Error) — the primary custom metric that differentiates BY from generic structure prediction tools — along with ipTM, pLDDT, RMSD, liability scoring, and the BY composite…

114 15d ago A 3 tokens original MIT

sdd-implement-spec

02

jentic/jentic-api-scorecard

Skill Claude CodeCodex

Implement an existing feature spec end-to-end — pick an unprocessed feature spec (one whose ## Phase N — ... heading in specs/roadmap.md does not yet carry the ✅ lifecycle marker), cut a feature branch, walk plan.md task groups in order with one primary atomic Conventional-Commits commit per group (plus optional small…

21 7d ago A 215 tokens original Apache-2.0

sdd-new-spec

03

jentic/jentic-api-scorecard

Skill Claude CodeCodex

Scaffold a feature spec for a roadmap phase and open it as a PR for human review. Reads specs/roadmap.md, lets the user pick a phase (or accepts one as argument), runs preflight checks, cuts a feature branch, writes specs/YYYY-MM-DD- / (requirements.md, plan.md, validation.md — grounded in specs/mission.md and…

21 7d ago A 141 tokens original Apache-2.0

jentic-api-improve

04

jentic/jentic-api-scorecard

Skill Claude CodeCodex

Improve OpenAPI documents for AI-readiness by fixing issues and enriching content based on Jentic API AI-Readiness Framework (JAIRF) scoring. Use when you need to raise an API's quality score, fix diagnostics, or add missing descriptions/summaries/examples — whether directly requested ('improve my API', 'fix OpenAPI…

21 7d ago A 124 tokens original Apache-2.0

llm-as-judge

05

cobusgreyling/agent-skills

Skill Claude CodeCodex

Design and validate LLM-as-judge scoring — pairwise vs pointwise, bias correction, anchor calibration, and the cases where a judge is the wrong tool. Use when the user is building an eval, scoring open-ended outputs, or comparing model versions and mentions LLM-as-judge, model grader, pairwise comparison, position…

14 1mo ago A 106 tokens

bookkeeping

07

broomva/skills

Skill Claude CodeCodex

Universal knowledge engine — scores, promotes, and compounds knowledge across all sources into a permanent, query-able entity graph.

3 yesterday A 25 tokens original MIT

bde-score

08

hbhqq9/bde-score

Skill Claude CodeCodex

AI-powered multi-factor stock scoring MCP server. 7-factor composite model (VIX, Volume Profile, RSI, MACD, Bollinger, OBV, ATR) with real-time Yahoo Finance data. Zero-config, no API keys required.

3 1mo ago A 53 tokens