agent-evaluation
2185Skill Claude Code
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement quality.
23,505 tagged Building agents, measured the same way as everything else here.
Browse within: LLM 209agent 202agentic-ai 195skills 149agents 145openai 86antigravity 83agentic-workflow 79cli 75agentic-framework 73ai-skills 73Autonomous Agents 67ai-assistant 66agent-framework 59
Skill Claude Code
Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement quality.
Skill Claude CodeCodex
This Skill provides a complete management guide for OpenSquad Agents, covering the full workflow of creation, configuration, deployment, and troubleshooting.
Skill Claude CodeCodex
The multi-provider cockpit for parallel AI coding. Supervise Claude Code, Pi, Droid, and OpenCode across isolated git worktrees from one Textual control plane (three lanes: NEEDS YOU / READY TO SHIP / IN FLIGHT) with real-time cross-worktree Conflict Guard (file-overlap detection), pluggable multiplexer backends (tmux…
Skill Claude CodeCodex
Use when the user wants to install, list, sync, remove, or manage AI coding-agent skills on this machine. Scribe manages a canonical skill store and links skills into Claude Code, Cursor, Codex, and other supported tools.
Skill Claude CodeCodex
Create new agent skills with proper structure, progressive disclosure, and bundled resources. Use when user wants to create, write, or build a new skill.
ASMN-96/ai-agents-skills-toolkit
Skill Claude CodeCodex
Use for governed software work that needs source-of-truth checks, routing, token/context governance, branch/PR discipline, validation honesty, and release readiness. Do not use as approval for writes, installs, migrations, broad plugin use, external tool activation, or global config changes.
Skill Codex
Recruit ChatGPT Pro as an external expert ally for complex engineering, research, review, or artifact work through the user's Codex in-app Browser or Chrome. Package safe context, open a dedicated Pro conversation, monitor long-running work, recover and download deliverables, and independently validate them. Use when…
Skill Claude CodeCodex
Run a coding agent unattended overnight, long running agent safety, stop a runaway loop, cap agent token spend. Sets up four guardrails around an agent loop you already have: a resumable handoff file so an interrupted run continues, a verification gate that fails closed so unchecked work is never accepted, a runaway…
Skill Codex
Use when designing, iterating, or auditing a Multica agent squad. Pick topology, write a leader charter and role instructions, and keep only necessary hard gates. Strong models; few rules; every rule must earn its place with evidence.
smallTechOrg/zero-shot-claude-boilerplate
Skill Claude Code
Turn a zero-shot idea into a perfectly-working, thoroughly-tested, spec-driven agent. One intake round (which also collects the API keys into .env), then the agent-builder builds one phase at a time — autonomous within a phase, with a human testing gate between phases. Also used to add a new capability to an existing…
yossiovadia/claude-code-wingman
Skill Claude CodeCodex
Your Claude Code wingman - dispatch coding tasks via tmux for free/work-paid coding while keeping Clawdbot API costs minimal.
Skill Claude CodeCodex
This skill should be used when the user asks to "install FastContext", "set up FastContext", "configure FastContext MCP", "add FastContext", "update FastContext", "upgrade FastContext", "uninstall FastContext", "remove FastContext", or mentions needing to install, update, or remove the FastContext exploration agent.…
Skill Claude CodeCodex
Generate production-quality Dify workflow DSL/YAML files with precise node schemas from Pydantic source models. Covers all 22+ node types for Dify v1.12+.
blindlove200/sub-agents-skills
Skill Claude Code
Execute external CLI AIs as isolated sub-agents for task delegation, parallel processing, and context separation. Use when delegating complex multi-step tasks, running parallel investigations, needing fresh context without current conversation history, or leveraging specialized agent definitions. Returns structured…
Skill Claude CodeCodex
Use when you've drafted or improved a skill locally and want to publish it to a shared pack — walks through validating the draft, finding the target pack, running the promote script, and reviewing the staged diff before it ships.
Skill Claude Code
Integrate lifecycle hooks across AI coding tools (Claude Code, Gemini CLI, Cursor, OpenCode) and any CLI. Use when: (1) installing/removing hooks, (2) creating OpenCode plugins, (3) setting up auto-formatting, testing, notifications, security policies, (4) wrapping CLI tools without hooks API. Triggers on: "add hook"…
Skill Claude Code
Meta-skill for creating production-ready Claude Code skills using evaluation-driven development, automated eval loops, blind A/B comparison, benchmark aggregation, description optimization, and progressive disclosure patterns. Includes grader, comparator, and analyzer agents with schema-enforced data interchange.
Skill Claude CodeCodex
Use when migrating an OpenClaw installation to Claude Code. Triggers on "migrate openclaw", "openclaw to claude code", "convert openclaw", "import openclaw", or when a user provides an OpenClaw directory path and wants to move to Claude Code. Also use when the user says "switch from openclaw" or "leave openclaw".
Skill Claude CodeCodex
Monitor and guide a Codex (or other AI coding agent) session running in a tmux pane. Use this skill when the user wants to babysit, monitor, supervise, or watch over a Codex/AI agent session, or when they mention tmux + codex together, or ask you to keep an eye on another AI agent working on implementation tasks. Also…
tomc98/claude-code-codex-skill
Skill Claude CodeCodex
Delegate tasks to OpenAI Codex (GPT-6 Astra) as background tasks for precision coding, code review, deliberation, and complex implementation. Always launch in background (runinbackground=true), continue working, then collect results with TaskOutput when needed.
Skill Claude Code needs its repo
Anti-slop agentic engineering co-pilot. Teaches the Research-Plan-Implement (RPI) workflow, context management, quality gates, per-agent isolation, and anti-slop patterns for building software with AI coding agents. Produces agent-workflow.md or project configuration files. Part of the mlops-tabular skill family but…
Skill Codex
Executes a written implementation plan by fanning out one Opus 5 low subagent per vertical slice, each spawning Luna Max as a native Codex subagent through the Codex plugin, with regular check-ins, one review phase before the end-to-end phase, and a human-gated merge. Use when the user asks to execute or ship a plan…
Skill Claude Code
Security-hardened agentic coding assistant with 3 defense layers — input scanning, tool boundaries, and output scanning via Model Armor.
DennisWei9898/claude-dev-skills
Skill Claude CodeCodex
A review skill for Claude Code skills, which are instruction packages that guide an AI coding assistant. It checks them against two quality frameworks.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: