cursor
01Agent
Fetch and follow https://raw.githubusercontent.com/agentrhq/webcmd/main/start.md to set up Webcmd end to end.
11 tagged browser use, measured the same way as everything else here.
Browse within: browser-automation 8playwright 8self-learning-ai 8
Agent
Fetch and follow https://raw.githubusercontent.com/agentrhq/webcmd/main/start.md to set up Webcmd end to end.
Agent
Fetch and follow https://raw.githubusercontent.com/agentrhq/webcmd/main/start.md to set up Webcmd end to end.
Agent
Fetch and follow https://raw.githubusercontent.com/agentrhq/webcmd/main/start.md to set up Webcmd end to end.
Agent
Your job is to analyze benchmark results and surface insights about skill performance.
Agent
Your job is to perform blind A/B comparison of two skill outputs without knowing which is which.
Agent
Your job is to evaluate skill outputs against a set of assertions. For each assertion, determine if it passed or failed based on the actual output.