ralph

ralph is a skill for Claude Code, Codex from jmagly/aiwg. It costs 16 tokens per session (2,440 once invoked), scanned A, original, MIT.

An iterative task loop that keeps executing, checking, learning from failures, and retrying until stated completion criteria are met.

In plain words
What is it for?
Useful for tasks with checkable outcomes, such as passing tests, compiling code, or satisfying a command-based condition.
Why use it?
It prevents a task from ending after the first failed attempt. Each failure becomes input for the next attempt.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one. Also seen: mentions CLAUDE.md; names the TodoWrite tool; mentions Claude Code.

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/jmagly/aiwg/ralph
Any agent
npx skills add jmagly/aiwg --skill ralph
Clone the repo
git clone --depth 1 https://github.com/jmagly/aiwg

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for ralph

README.md
[![agentmods](https://agentmods.dev/badge/skills/jmagly/aiwg/ralph.svg)](https://agentmods.dev/skills/jmagly/aiwg/ralph)
Your own site
<a href="https://agentmods.dev/skills/jmagly/aiwg/ralph"><img src="https://agentmods.dev/badge/skills/jmagly/aiwg/ralph.svg" alt="Measured on agentmods" height="20"></a>
Per session 16 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,440 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00016 $0.02440
Opus 5 $0.00008 $0.01220
Sonnet 5 $0.00003 $0.00488
Haiku 4.5 $0.00002 $0.00244

Measured 6d ago against content hash 02cc3716dc6d, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-05, from the pricing page.

Security

Grade A, and why

ralph scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agentic/code/addons/agent-loop/skills/ralph/SKILL.md · 378 lines

How it starts

The opening of the file, as written. The whole thing — 378 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Agent Loop

You are the Agent Loop Orchestrator - executing iterative AI task loops until completion criteria are met.

Core Philosophy

"Iteration beats perfection" - errors become learning data within the loop rather than session-ending failures.

Your Role

You manage the iterative execution cycle:

  1. Parse task definition and completion criteria
  2. Execute the task
  3. Verify completion criteria
  4. Learn from failures and extract actionable insights
  5. Iterate if not complete (re-execute with learnings)
  6. Report final status with completion report

Natural Language Triggers

Users may say:

  • "ralph this: [task]"
  • "ralph [task]"
  • "loop until: [criteria]"
  • "keep trying until [condition]"
  • "iterate on [task] until [done]"
  • "agent loop [task]"

Parameters

Task (required)

The task to execute. Should be:

  • Specific and actionable
  • Measurable completion state
  • Self-contained (all context provided)

--completion (optional — inferred when omitted)

Success criteria. Must be:

  • Verifiable (tests, lint, compilation)
  • Specific (not subjective)
  • Checkable via commands

Good examples:

  • --completion "npm test passes with 0 failures"
  • --completion "npx tsc --noEmit exits with code 0"
  • --completion "all files in src/ have JSDoc comments"
  • --completion "coverage report shows >80%"

Poor examples (avoid these):

  • --completion "code looks good"
  • --completion "feature is done"

When omitted: the loop delegates to the infer-completion-criteria skill, which derives a measurable criterion from project docs (CLAUDE.md / AGENTS.md / AIWG.md), package manifests, CI configuration, and .aiwg/ artifacts. The proposed criterion is shown to the user for confirmation before the loop starts. Pass --auto-criteria to skip confirmation and use the inferred criterion directly (useful in CI / automation). Pass --no-infer-completion to require explicit --completion and fail fast if missing.

See @$AIWG_ROOT/agentic/code/addons/agent-loop/skills/infer-completion-criteria/SKILL.md for the inference pipeline.

Read the full file on GitHub · 378 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 378 lines · 16 tokens per session scan A 02cc3716dc6d

Subscribe to this mod's changes

ralph is a skill published in the GitHub repository jmagly/aiwg (209 stars, last pushed yesterday), licensed MIT. It adds 16 tokens to every session and 2,440 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

skeptical-triage

Reusable 3-round self-challenge + arbiter pattern for filtering false positives from findings/verdicts. Use when the cost of a false-positive gate block exceeds the cost of 4 extra LLM turns.

avelikiy/great_cto · 49 tokens

vertical-hr-recruiting

Domain-knowledge primer for the HR & recruiting vertical (ATS, onboarding, workforce scheduling, engagement). Applied by architect/pm during spec authoring so they aren't naive about hiring pipelines, the admitted offer→onboard data-carry gap, EEO/I-9 compliance, and shift-coverage rules. Stops the four products from…

avelikiy/great_cto · 90 tokens

vertical-real-estate

Residential-proptech domain knowledge so architect / pm aren't naive when speccing real-estate products (listings, lead-crm, transaction-coordination, property-mgmt). Codifies MLS/IDX reality, listing status lifecycle + syndication canonical-source, long-cycle lead nurture, transaction-coordination as the high-pain…

avelikiy/great_cto · 99 tokens

well-architected

6-pillar architecture review framework. Adapted from AWS Well-Architected for use by greatcto's architect agent on every non-nano ARCH document. Forces explicit answers across operational excellence, security, reliability, performance, cost, and sustainability — not just feature design.

avelikiy/great_cto · 60 tokens

product-economics

Does this product make money at a price someone will pay? Forces contribution margin, a price with a stated basis, and a bottom-up market size — each number labelled measured / assumed / unknown, so a guess can never be read as a calculation.

avelikiy/great_cto · 55 tokens

lifecycle-messaging

Email/SMS lifecycle and deliverability framework for SMB Product-Builder products that send transactional or lifecycle messages (booking reminders, CRM sequences, receipts, win-back). Codifies provider selection (Resend/Postmark/Twilio/SendGrid), domain auth (SPF/DKIM/DMARC), consent and compliance (TCPA, CAN-SPAM…

avelikiy/great_cto · 132 tokens