Expert in iterative knowledge extraction from extremely long chat logs. It processes contexts.md in chunks (approx. 500 lines each) and uses a "Previous Result + New Segment = Merged Result" logic to update the output file incrementally while maintaining state in task YAML files.
A review agent that compares two blind test results, then examines the winning and losing skills and their execution records. It produces practical suggestions for improving the weaker skill.
An agent that compares two results without knowing which tool produced either one. It judges them against the original task and scores their correctness, completeness, and structure.
An agent for checking whether expected results were achieved by another agent or workflow. It reads an execution record and the files produced, then marks each expectation as passed or failed with supporting evidence.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: