llm-eval skills

41 tagged llm-eval, measured the same way as everything else here.

Browse within: ai-monitoring 39ai-observability 39aiengineering 39datasets 39llmops 39

phoenix-cli

01

Arize-ai/phoenix

Skill Claude CodeCodex

Debug LLM applications using the Phoenix CLI. Fetch traces, analyze errors, structure trace review with open coding and axial coding, inspect datasets, review experiments, query annotation configs, and use the GraphQL API. Use whenever the user is analyzing traces or spans, investigating LLM/agent failures, deciding…

not rated 11k +28 today A 113 tokens

phoenix-github

02

Arize-ai/phoenix

Skill Claude CodeCodex

Manage GitHub issues, labels, project boards, sprint operations, and roadmap health for the Arize-ai/phoenix repository. Use when filing roadmap issues, triaging bugs, applying labels, running sprint close-out and rollover, auditing board hygiene, checking ticket-load balance across the team, keeping roadmap epics up…

not rated 11k +28 today A 90 tokens

pxi-eval-dataset

03

Arize-ai/phoenix

Skill Claude CodeCodex

Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/). Use whenever the user asks to create, author, draft, expand, or audit an eval dataset for a PXI tool, skill, or behavior — including phrases like "write evals for ", "test PXI behavior", "synthetic dataset for PXI", "cover this tool with…

not rated 11k +28 today A 139 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: