Run multi-round AI brainstorming debates between multiple LLM providers (GPT, DeepSeek, Groq, Ollama). Claude actively participates as a debater alongside external models. Use when the user wants diverse perspectives, multi-model critiques, or synthesized answers from several AI models working together.
Use when the user wants a debate, second opinion, adversarial review, or iterative discussion between Claude and Codex (the OpenAI CLI) about a spec, design, plan, document, position, refactoring, or codebase — the subject can be anything, technical or non-technical, e.g. "ask Codex", "debate this with Codex", "have…
Run the local arena evaluation suite. Trigger when the user asks to "eval arena", "evaluate arena", "test arena quality", "score arena", "run arena evals", or anything semantically equivalent. Fans out parallel subagents to execute cases through the arena CLI and judge the resulting transcripts against per-case…