ralph-analytics

ralph-analytics is a skill for Claude Code, Codex from jmagly/aiwg. It costs 14 tokens per session (400 once invoked), scanned A, original, MIT.

A guide for calculating and displaying metrics from an agent loop’s execution history, including how often loops succeed or get stuck.

In plain words
What is it for?
Use it to inspect all loops or a selected date or loop, view a brief dashboard, identify patterns, and export the results.
Why use it?
It helps reveal recurring failures, the amount of iteration needed, when human help is required, and whether results are improving.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/jmagly/aiwg/ralph-analytics
Any agent
npx skills add jmagly/aiwg --skill ralph-analytics
Clone the repo
git clone --depth 1 https://github.com/jmagly/aiwg

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for ralph-analytics

README.md
[![agentmods](https://agentmods.dev/badge/skills/jmagly/aiwg/ralph-analytics.svg)](https://agentmods.dev/skills/jmagly/aiwg/ralph-analytics)
Your own site
<a href="https://agentmods.dev/skills/jmagly/aiwg/ralph-analytics"><img src="https://agentmods.dev/badge/skills/jmagly/aiwg/ralph-analytics.svg" alt="Measured on agentmods" height="20"></a>
Per session 14 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 400 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00014 $0.00400
Opus 5 $0.00007 $0.00200
Sonnet 5 $0.00003 $0.00080
Haiku 4.5 $0.00001 $0.00040

Measured 6d ago against content hash 844cd4b8edbe, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-05, from the pricing page.

Security

Grade A, and why

ralph-analytics scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agentic/code/addons/agent-loop/skills/ralph-analytics/SKILL.md · 53 lines

What it actually says

Al Analytics Command

Display aggregate analytics and metrics from agent loop execution history.

Instructions

When invoked, analyze agent loop data and present metrics:

  1. Scan Loop History

    • Load all loop records from .aiwg/ralph/
    • Load reflections from .aiwg/ralph/reflections/
    • Load debug memory from .aiwg/ralph/debug-memory/
  2. Calculate Metrics

    • Success rate: % of loops that completed successfully
    • Average iterations: Mean iterations to completion
    • Reflection reuse rate: % of reflections applied in subsequent loops
    • Stuck loop rate: % of loops that hit stuck detection
    • Escalation rate: % requiring human intervention
  3. Pattern Analysis

    • Most common failure types
    • Most effective fix patterns
    • Average time per iteration
    • Quality trajectory per loop
  4. Display Dashboard

    • Summary metrics table
    • Trend indicators (improving/stable/degrading)
    • Recommendations for improvement

Arguments

  • --since [date] - Analyze loops from date (default: all)
  • --loop [id] - Analyze specific loop
  • --export [path] - Export analytics to file
  • --brief - Show summary only

References

  • @$AIWG_ROOT/agentic/code/addons/ralph/schemas/reflection-memory.json - Reflection schema
  • @$AIWG_ROOT/agentic/code/addons/ralph/schemas/debug-memory.yaml - Debug memory schema
  • @$AIWG_ROOT/agentic/code/addons/ralph/docs/reflection-memory-guide.md - Guide
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 53 lines · 14 tokens per session scan A 844cd4b8edbe

Subscribe to this mod's changes

ralph-analytics is a skill published in the GitHub repository jmagly/aiwg (209 stars, last pushed yesterday), licensed MIT. It adds 14 tokens to every session and 400 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

skeptical-triage

Reusable 3-round self-challenge + arbiter pattern for filtering false positives from findings/verdicts. Use when the cost of a false-positive gate block exceeds the cost of 4 extra LLM turns.

avelikiy/great_cto · 49 tokens

vertical-hr-recruiting

Domain-knowledge primer for the HR & recruiting vertical (ATS, onboarding, workforce scheduling, engagement). Applied by architect/pm during spec authoring so they aren't naive about hiring pipelines, the admitted offer→onboard data-carry gap, EEO/I-9 compliance, and shift-coverage rules. Stops the four products from…

avelikiy/great_cto · 90 tokens

vertical-real-estate

Residential-proptech domain knowledge so architect / pm aren't naive when speccing real-estate products (listings, lead-crm, transaction-coordination, property-mgmt). Codifies MLS/IDX reality, listing status lifecycle + syndication canonical-source, long-cycle lead nurture, transaction-coordination as the high-pain…

avelikiy/great_cto · 99 tokens

well-architected

6-pillar architecture review framework. Adapted from AWS Well-Architected for use by greatcto's architect agent on every non-nano ARCH document. Forces explicit answers across operational excellence, security, reliability, performance, cost, and sustainability — not just feature design.

avelikiy/great_cto · 60 tokens

product-economics

Does this product make money at a price someone will pay? Forces contribution margin, a price with a stated basis, and a bottom-up market size — each number labelled measured / assumed / unknown, so a guess can never be read as a calculation.

avelikiy/great_cto · 55 tokens

lifecycle-messaging

Email/SMS lifecycle and deliverability framework for SMB Product-Builder products that send transactional or lifecycle messages (booking reminders, CRM sequences, receipts, win-back). Codifies provider selection (Resend/Postmark/Twilio/SendGrid), domain auth (SPF/DKIM/DMARC), consent and compliance (TCPA, CAN-SPAM…

avelikiy/great_cto · 132 tokens