Galileo-Agent-Labs/eval-engineer

Bring Galileo powered eval workflows into Claude Code and Codex.

41Stars on the repository
9Mods indexed here, across every type
22d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

eval-audit

01

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when the user asks for an AI app audit, launch readiness review, safety/security review, OWASP agentic risk check, metric coverage review, or production RCA gap review.

not rated 41 22d ago A 40 tokens original MIT

eval-cost

02

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when the user asks to make an AI app cheaper or faster, reduce tokens, latency, model/tool/retrieval/rerank/self-check/retry/evaluator cost, or compare cost before/after.

not rated 41 22d ago A 45 tokens original MIT

eval-dataset

03

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when the user asks to turn a failure into an eval, create/review/accept/reject dataset cases, or convert Galileo traces, metric gaps, or production examples into cases.

not rated 41 22d ago A 41 tokens original MIT

eval-diagnose

04

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when Galileo evidence is available and the user asks why a trace, session, log stream, experiment, metric, or AI app behavior failed, regressed, or became unsafe.

not rated 41 22d ago A 41 tokens original MIT

eval-engineer

05

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when a user is unsure which Eval Engineer command to run for AI agents/RAG apps, needs onboarding/status for a .galileo workspace, or asks where to start.

not rated 41 22d ago A 39 tokens original MIT

eval-fetch

06

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when a user asks to fetch Galileo evidence, provides Galileo URLs or IDs, says "fetch this Galileo link", or needs traces, sessions, experiments, or log streams saved locally.

not rated 41 22d ago A 40 tokens original MIT

eval-measure

07

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when the user asks if an AI app is measured correctly, needs Galileo metrics, expected-output contracts, metric profiles, eval gates, or measurement before optimizing.

not rated 41 22d ago A 36 tokens original MIT

eval-setup

08

Galileo-Agent-Labs/eval-engineer

Skill Claude CodeCodex

Use when the user asks to set up Eval Engineer, check .galileo readiness, create workspace scaffolding, or configure editable files, verification commands, app type, or evidence paths.

not rated 41 22d ago A 41 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: