Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/vladolaru/claude-code-plugins/analyzing-codex-sessionsnpx skills add vladolaru/claude-code-plugins --skill analyzing-codex-sessionsgit clone --depth 1 https://github.com/vladolaru/claude-code-pluginsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/vladolaru/claude-code-plugins/analyzing-codex-sessions)<a href="https://agentmods.dev/skills/vladolaru/claude-code-plugins/analyzing-codex-sessions"><img src="https://agentmods.dev/badge/skills/vladolaru/claude-code-plugins/analyzing-codex-sessions.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00082 | $0.05213 |
| Opus 5 | $0.00041 | $0.02606 |
| Sonnet 5 | $0.00016 | $0.01043 |
| Haiku 4.5 | $0.00008 | $0.00521 |
Grade A, and why
analyzing-codex-sessions scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 253 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Skill Directory
Resolve SKILL_DIR to the absolute directory containing this SKILL.md, as
shown by the current host, before using any bundled path below.
Before a shell command uses $SKILL_DIR, assign it in that command or replace
it with the resolved path. It is not a host-exported environment variable.
Analyzing Codex Raw Sessions
Overview
Codex CLI writes every thread — the root conversation and each subagent alike — as its own JSONL rollout file, one JSON object per line. Each line has a top-level type, and the two types that matter carry a nested payload.type that does the real discriminating.
Core principle: event_msg/item_completed is the digested layer; response_item is the raw layer. Analyze items first — they carry semantic types with timing and exit codes already attached. Drop to response_item only for prompt text and reasoning content.
A CommandExecution item hands you the command, its exit_code, and its duration in a single object. The equivalent question asked of response_item requires correlating a custom_tool_call with its custom_tool_call_output and reconstructing timing from entry timestamps. Start at the digested layer and stay there unless you need something it does not carry.
When to Use
- Analyzing what happened in a Codex run (commands, failures, file changes, cost)
- Reconstructing a thread tree — which subagents a root spawned and what each did
- Extracting metrics (tokens, wall duration, command counts, failure counts) across Codex threads
- Debugging why a Codex subagent failed or produced poor results
- Comparing Codex reviewer roles (
code-reviewer,php-tests-reviewer, …) across runs - Building or maintaining Codex rollout analysis scripts
Before You Start
| Your goal | Start here |
|---|---|
| Understand what a whole Codex run did | codex_session_analyzer.py — it resolves the root and walks its subagent tree for you |
| Analyze one thread you already have the id for | The rollout file whose name ends -{thread-id}.jsonl, then its item_completed entries |
| Find sessions for a specific project | Scan line 1 of each rollout in the date window and match session_meta.payload.cwd — there is no per-project directory |
| Compare metrics across threads or roles | codex_session_metrics.py |
| Debug a failing command | item_completed → CommandExecution items with non-zero exit_code; stderr and aggregated_output are inline |
| Recover the prompt text or the model's reasoning | response_item entries of type message and reasoning — items do not carry full prompts |
| Attribute cost to a thread | The last event_msg/token_count entry's info.total_token_usage |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 253 lines · 82 tokens per session scan A 764073648e17
analyzing-codex-sessions is a skill published in the GitHub repository vladolaru/claude-code-plugins (8 stars, last pushed yesterday), licensed MIT. It adds 82 tokens to every session and 5,213 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.
chat-perf
Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…