agentkaizen

A method for measuring whether an AI coding agent followed its instructions during a session. It checks behavior such as tool use, branching, and compliance with project rules.

In plain words
What is it for?
It is for evaluating recent Codex or Claude Code sessions and checking whether changes to agent instructions improved compliance.
Why use it?
It turns a subjective review of an agent session into a structured score and can reveal repeated problems.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/theillusionoflife/agentkaizen/optimize-coding-agent-skill
Any agent
npx skills add TheIllusionOfLife/AgentKaizen --skill optimize-coding-agent-skill
Clone the repo
git clone --depth 1 https://github.com/TheIllusionOfLife/AgentKaizen

Made for: Claude Code, Codex.

Per session 114 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,240 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00114 $0.02240
Opus 5 $0.00057 $0.01120
Sonnet 5 $0.00023 $0.00448
Haiku 4.5 $0.00011 $0.00224

Measured 2d ago against content hash b53c79dfad4f, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

agentkaizen scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (scripts/check_setup.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skill/optimize-coding-agent-skill/optimize-coding-agent-skill/SKILL.md · 208 lines

How it starts

The opening of the file, as written. The whole thing — 208 lines — stays where its author put it; the contents beside it link to each section on GitHub.

AgentKaizen

Measure and improve CLI-based AI coding agent behavior. Works entirely through agent-native tools (Read/Glob/Bash/Write) — no Python install or CLI required.


Section 1 — Detect Environment

First, determine which agent context is active:

echo $CLAUDECODE
  • Non-empty → Claude Code context: sessions at ~/.claude/projects/, one-shot via claude -p
  • Empty → check for Codex: command -v codex or presence of ~/.codex/sessions/Codex context
  • If running as Claude Code (responding to this prompt), you are always in Claude Code context

Section 2 — Score a Session (native, no CLI)

Claude Code path

Session discovery:

  1. Glob ~/.claude/projects/ for project directories (skip */subagents/)
  2. In the most recently modified project dir, select the latest *.jsonl NOT under */subagents/
  3. Hard cap: read at most 500 records; skip early if file is very large

Record parsing rules:

  • Skip records where type is in: progress, system, file-history-snapshot, queue-operation
  • user records → user turns (content may be string or list of blocks; extract text)
  • assistant records → assistant turns (content blocks: text, tool_use, thinking)

Completion detection:

  • Last assistant record with message.stop_reason == "end_turn""complete"
  • Any record with type == "last-prompt""complete"
  • Otherwise → "incomplete"

Codex path

Session discovery:

  1. Read ~/.codex/session_index.jsonl — parse lines, sort by updated_at desc
  2. Take most recent entry; resolve session file path from id under ~/.codex/sessions/
  3. Fallback: glob ~/.codex/sessions/**/*.jsonl sorted by mtime if index absent

Record parsing rules:

  • response_item records → messages (access payload.role and payload.content)
  • event_msg records → metadata (usage, completion)

Completion detection:

  • event_msg with payload.type == "task_complete""complete"
  • Otherwise → "incomplete"

Read the full file on GitHub · 208 lines

Files

What ships with it

5 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 208 lines · 114 tokens per session scan A b53c79dfad4f

Subscribe to this mod's changes

agentkaizen is a skill published in the GitHub repository TheIllusionOfLife/AgentKaizen (2 stars, last pushed 5mo ago), licensed Apache-2.0. It adds 114 tokens to every session and 2,240 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens

chat-perf

Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.

microsoft/vscode · 51 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens