benchmark
01Skill Claude CodeCodex
Run SWE-bench Lite benchmarks against agtx coding agent workflows. Guides setup, configuration, execution, evaluation, and reporting.
🏄🏼♂️ The blackboard for coding agents - agentic development environment for claude code, codex, cursor, opencode, grok and more.
Skill Claude CodeCodex
Run SWE-bench Lite benchmarks against agtx coding agent workflows. Guides setup, configuration, execution, evaluation, and reporting.
Skill Claude CodeCodex
Execute an approved implementation plan. Implement the changes, then write a summary to .agtx/execute.md and stop.
Skill Claude CodeCodex
Plan a task implementation. Analyze the codebase, create a detailed plan, write it to .agtx/plan.md, then stop and wait for user approval before making any changes.
Skill Claude CodeCodex
Explore the codebase to understand a task before planning. Write findings to .agtx/research.md and stop. This is a read-only exploration — do not modify any files.
Skill Claude CodeCodex
Self-review completed work. Check for correctness, edge cases, and code quality. Write review to .agtx/review.md and stop.
Skill Claude CodeCodex
Enter brainstorm mode to explore a feature or enhancement idea. Stays in discussion mode only — no planning, no implementation. Use /agtx:sweep when ready to push outcomes to the board.
Skill Claude CodeCodex
Sweep this conversation into agtx tasks and push them to the kanban board. Use when the user wants to capture, decompose, or hand off conversation results to the agtx board.