Agent
Total unimplemented-tagged scenarios: 76 Classified: 76.
37 tagged evaluation, measured the same way as everything else here.
Browse within: browser-automation 18Multi-Agent 10octant 10public-goods 10
Agent
Total unimplemented-tagged scenarios: 76 Classified: 76.
redhat-community-ai-tools/harness-eval
Agent
General helper.
redhat-community-ai-tools/harness-eval
Agent
Run lint checks and fix issues automatically.
Agent
You are an analyzer agent for the michael-polanyi skill. Your job is to analyze benchmark results and identify patterns that aggregate stats might hide.
Agent
Performs blind A/B comparison between responses to determine which feels more like a practitioner's judgment.
Agent
You are a grader agent for the michael-polanyi skill. Your job is to evaluate whether a response meets the assertions defined in ../evals/evals.json.
Agent
You are the Synthesizer for the Architect evaluation system. You receive structured outputs from multiple evaluator agents and synthesize them into a unified evaluation with clear scores, actionable recommendations, and strategic direction.
Agent
You are the Analyzer agent for Founder Mode — Phase 7: Post-Run Analysis.
Agent
You are the Architecture Generator agent for Founder Mode — Phase 2.
Agent
Read-only isolated agent that evaluates skill/agent execution quality.
golemfoundation/octant-council-builder
Agent
Research on-chain activity, deployments, and usage metrics.
golemfoundation/octant-council-builder
Agent
Analyze project website, documentation, and public communications.
golemfoundation/octant-council-builder
Agent
Synthesize all evaluations into a final council report with recommendation.