Plugin Claude Code
A workflow of skills for creating, updating, judging, reviewing, evaluating, and auditing AI skills and agent definitions.
Claude Code skills covering the skill-authoring lifecycle: create, judge, eval with persisted regression suites, audit for ecosystem-wide coherence, review, and update
Plugin Claude Code
A workflow of skills for creating, updating, judging, reviewing, evaluating, and auditing AI skills and agent definitions.
Instructions file CodexOpenCode
AGENTS.md instructions for robcsaszar/ai-forge, covering agents.md, mission, layout convention, judgment boundaries and pre-publish checks.
Skill Claude CodeCodex
HITL (Human-In-The-Loop) application of a numbered list one item at a time — status board upfront, per-item approve/skip, approve-all mode, and one opt-in commit bundling all approved items after the loop. Use when stepping through ai-forge-judge findings or any numbered changes. Triggers are apply these, go through…
Skill Claude CodeCodex
Batch-evaluate all skills and agents in the repo with ai-forge-judge and render a single consolidated grade report sorted by grade (worst first) so effort is directed correctly. Use when reviewing overall skill/agent quality, finding where to invest improvement effort, or after bulk changes. Triggers are audit all…
Skill Claude CodeCodex
Create a new skill (SKILL.md), agent definition, or instruction file. Use when converting ad-hoc knowledge into a reusable skill, scaffolding an agent for Claude Code, GitHub Copilot, OpenAI Codex, or Google Gemini, or creating instruction files for glob-pattern matching. Don't use for updating existing artifacts …
Skill Claude CodeCodex
Behavioral eval for skills and agents — whether the artifact actually changes model behavior, not just whether it scores well on a rubric. Use when verifying a skill or agent works in practice, or when tracking its performance across repeated runs to catch regressions. Triggers are test this skill, test this agent…
Agent
You make blind judgments between two outputs. You do not know which came from a skill/agent and which was a baseline — do not ask, do not infer.
Agent
You grade outputs against a list of expectations. You receive an output (text produced by an agent) and an expectations list (assertions about what the output should contain or demonstrate).
Agent
You receive a completed blind comparison (arbiter output) and a label mapping revealing which output was "withartifact" vs "baseline". Your job: explain why the winner won and surface targeted improvements to the artifact.
Skill Claude CodeCodex
Evaluate any LLM prompt (SKILL.md, agent definition, system prompts, instruction files) for quality — grouped dimensional scoring with letter grade and step-through-ready numbered improvements list. Triggers are judge/review/audit/score/evaluate this skill or prompt, grade this agent. Don't use for behavioral testing…
Skill Claude CodeCodex
Report what a skill or agent actually does versus what its description claims — drift, undeclared behaviors, verdict. Use before updating a skill or agent, or when a description is suspected of being out of date. Triggers are recap [skill], what does [skill] do, audit [skill] description, summarize [skill]. Don't use…
Skill Claude CodeCodex
Critically reviews and stress-tests agent, skill, and AI workflow definitions before they ship. Use whenever someone creates, modifies, or proposes an agent config, skill file, system prompt, or AI-powered workflow — including 'review this agent', 'check this skill', 'is this agent safe', or any request for feedback…
Skill Claude CodeCodex
Revise an existing SKILL.md or agent definition. Use when an existing skill or agent needs revision, modification, or improvement — including when it misfires, triggers too broadly or too rarely, or has drifted from what its description claims. Don't use for creating new artifacts — use ai-forge-create for that.…
At most 3 mods per repository are shown here — the rest are on their repository pages: