google cloud ml skills

3 tagged google cloud ml, measured the same way as everything else here.

agent-eval-workflow

01

GoogleCloudPlatform/professional-services

Skill Claude CodeCodex

This skill should be used when the user wants to evaluate an AI agent end-to-end: scaffold an evaluation, design metrics that test a real hypothesis, make an agent measurable, audit generated eval config, read evaluation results, or run an improvement ("hill climbing") loop. Covers evaluation methodology, metric…

3.1k 10d ago A 123 tokens original Apache-2.0

agent-eval

02

GoogleCloudPlatform/professional-services

Skill Claude CodeCodex

Executes high-performance agent evaluations, multi-turn UserSim simulations, and declarative metric grading aligned with google/agents-cli and the Quality Flywheel. Publishes benchmark artifacts to the GCS Evaluation Registry, executes automated head-to-head delta comparisons (--compare-to), and optimizes system…

3.1k 10d ago A 136 tokens original Apache-2.0

eval-breakdown

03

GoogleCloudPlatform/professional-services

Skill Claude CodeCodex

Performs an exhaustive, question-by-question narrative diagnostic breakdown of an agent-eval benchmark run by analyzing questionanswerlog.md, evalsummary.json, and raw trajectory traces. Use when diagnosing low score causes, investigating the Memory Reuse vs. Traceability rubric clash, performing pre-release failure…

3.1k 10d ago A 100 tokens original Apache-2.0