Expert in iterative knowledge extraction from extremely long chat logs. It processes contexts.md in chunks (approx. 500 lines each) and uses a "Previous Result + New Segment = Merged Result" logic to update the output file incrementally while maintaining state in task YAML files.
A review agent that compares two blind test results, then examines the winning and losing skills and their execution records. It produces practical suggestions for improving the weaker skill.
An agent that compares two results without knowing which tool produced either one. It judges them against the original task and scores their correctness, completeness, and structure.
Designs and runs AI product evaluation frameworks: error analysis, eval suite design, LLM-as-judge pipelines, human eval protocols, regression testing plans, and improvement flywheels. Use this agent when the user is building an AI-powered feature and needs to define how to measure quality, catch regressions, or…
Plans go-to-market execution: launch planning, ICP definition, messaging hierarchy, positioning (April Dunford 5-component), pricing model design, growth loops, and AI feature monetization. Use this agent when the user needs to plan how to bring a product or feature to market — any task requiring multi-constraint GTM…
Produces audience-tailored stakeholder communications: executive summaries, engineering briefs, launch announcements, risk escalations, and weekly digests. Use this agent when the user needs to communicate the same information to different audiences, or when a communication requires careful tone calibration for a…