evaluation
73Agent
Learn how to evaluate agents and workflows in Agent Framework using local checks, custom evaluators, and Microsoft Foundry.
2,228 tagged Analytics and metrics, measured the same way as everything else here.
Browse within: Multi-Agent 43claude-plugin 43ai-workspace 37business-automation 37ai-skills 35seo 33claude-ai 32observability 32agentic-workflow 31ai-assistant 27agentic 26claude-plugins 26agent-orchestration 25claude-code-skill 25
Agent
Learn how to evaluate agents and workflows in Agent Framework using local checks, custom evaluators, and Microsoft Foundry.
Agent
Learn how to use observability with Agent Framework.
Agent
Analytics engineering specialist for event tracking implementation, analytics schemas, conversion funnels, A/B test design, and measurement planning. Use when the task requires instrumenting features with analytics, designing event taxonomies, building conversion funnels, or planning experiments. For example: adding…
Agent
Analyze experiment results from any tracking system. Use when asked to compare runs, generate reports, summarize training results, or monitor experiments. Triggers on phrases like "compare runs", "analyze results", "training report", "experiment summary", "monitor training", or "which run is best".
Agent
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
Agent
An agent that measures software performance and investigates CPU, memory, input/output, and network bottlenecks. A performance baseline is a repeatable set of measurements used to detect later changes.
Agent Claude Code
Data analysis specialist for BigQuery, Snowflake, GA4, Marimo. Accumulates domain knowledge and data quality patterns.
Agent Claude Code
Weekly standing patrol that mines telemetry — prunes never-consulted assets, recalibrates time-boxes, reclassifies "other" failures, and mines primitive candidates (duplicated code shapes, repeated raw hackage imports) into triager-fed issues.
Agent
Facilitates human-in-loop qualitative evaluation of skill executions.
Agent
Copyright (c) 2026 Advanced Micro Devices, Inc. All rights reserved.
Agent
Copyright (c) 2026 Advanced Micro Devices, Inc. All rights reserved.
Agent
You are a data analyst specializing in Korean public data and dataset analysis. You turn a user's goal (research real-estate prices X, screen court auctions Y, analyze stock Z, pull KOSIS statistic W, profile this CSV) into concrete, evidence-based deliverables: public-data research briefs, data tables, interactive…
Agent
Agent "storytelling" from ai-analyst-lab/ai-analyst, covering agent: storytelling, purpose, inputs, workflow and step 1: ingest the analysis outputs.
Agent
Use when monitoring Bosun task execution with the simulator first, then daemon health and throughput only after a real task completes the full PR/review/merge path end to end. Also use when resuming the local monitor from .bosun-monitor session notes.
pamirtuna/gamestudio-subagents
Agent
You are the Data Scientist Agent for game development projects. You collect, analyze, and interpret data to provide actionable insights for current projects and improve future iterations through machine learning and predictive analytics.
Agent Claude Code
A chat-analysis specialist for examining message history from a WeChat conversation or group. It looks at interaction patterns, communication habits, topics, emotional tone, and changes in the relationship.
thoughtbot/rails-audit-thoughtbot
Agent
You are a subagent responsible for collecting code quality metrics from a Rails application using RubyCritic (which wraps Reek, Flay, and Flog). The user has already confirmed they want code quality data. Follow the steps below in order. Return the results as described in the Output section.
Agent Claude Code
Veri bazlı çalışma koçu. Günlük notlar, görev panosu ve denetim kayıtlarını analiz ederek verimlilik kalıplarını, tükenmişlik sinyallerini ve kaçırılan fırsatları tespit eder. Kullanıcı çalışma performansını değerlendirmek istediğinde veya /coach komutuyla çağrılır. Motivasyon koçluğu yapmaz — sadece veriye dayalı…
zubair-trabzada/ai-trading-claude
Agent
Weight: 15% of composite Trade Score Output: riskscore (0-100), maxdrawdownestimate, positionsizerecommendation, keyrisks.
Agent
You are a media buying and budget allocation specialist. You design the optimal budget distribution across platforms, funnel stages, and campaign types, project returns, build scaling roadmaps, and assess financial risk for advertising campaigns.
TheMattBerman/google-ads-copilot
Agent
Specialist agent for landing page → conversion path diagnosis. Separates tracking problems from UX/path problems.
TheMattBerman/google-ads-copilot
Agent
Specialist agent for campaign/ad group architecture, intent mixing, routing, and structure cleanup.
Agent ✓ vendor
Analyzes crowdsourced subjective test results — runs resultparser.py for data cleaning, quality checks, and per-clip/per-worker MOS aggregation, and writes a re-runnable rerunresultparser.bat that re-runs resultparser.py.
zacharyfmarion/openscad-studio
Agent Codex
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.