cache-report

A command-line report for examining prompt-cache and reused-input data from Claude and Codex. Prompt caching stores or reuses repeated input so usage can be analyzed over time.

In plain words
What is it for?
It reports usage by date, session, provider or project, including cache rates, estimated savings, wasted cost, net cost and possible anomalies.
Why use it?
It helps reveal caching regressions, wasted cache writes and differences between the two providers without mixing their accounting methods.

Command

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/omrikais/cctally/cache-report
Clone the repo
git clone --depth 1 https://github.com/omrikais/cctally
Per session 0 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 2,642 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.02642
Opus 5 $0.00000 $0.01321
Sonnet 5 $0.00000 $0.00528
Haiku 4.5 $0.00000 $0.00264

Measured 2d ago against content hash d45d2a306ea6, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

cache-report scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

docs/commands/cache-report.md · 198 lines

How it starts

The opening of the file, as written. The whole thing — 198 lines — stays where its author put it; the contents beside it link to each section on GitHub.

cache-report

Claude cache diagnostics or Codex cached-input/token-reuse analytics across days or sessions. Provider sections stay separate because their token and cache semantics differ.

Claude cost coverage: Claude dollar and token totals are transcript-derived lower bounds, not exact /usage billing totals. Codex accounting is unaffected.

Synopsis

cctally cache-report
    [--days DAYS] [--since SINCE] [--until UNTIL]
    [--by-session]
    [--offline] [--project PROJECT] [--json]
    [--anomaly-threshold-pp PP] [--anomaly-window-days N] [--no-anomaly]
    [--sort {date,net,cache,recent,cost,anomaly,reuse}]
    [--source {claude,codex,all}] [--speed {auto,standard,fast}]

Purpose

For Claude, surface cache behavior so a prompt-caching regression becomes visible in days rather than dollars. For Codex, surface cached-input/token reuse without relabeling it as Claude cache behavior.

What it shows

Claude cache diagnostics show:

  • Cache % = cache_read_tokens / (input + cache_create + cache_read)
  • $ Saved — counterfactual no-cache cost minus actual cost
  • $ Wasted — cache-write premium that did not yield enough reads
  • Net $Saved – Wasted; more than 1e-9 USD below zero means caching is costing you
  • Anomaly glyph (⚠) — Net $ is more than 1e-9 USD below zero, or Cache % drops ≥15pp vs. the trailing median
  • Eval — which of the four evaluation states the row is in (see Evaluation states)

Claude financial fields use each retained response's effective message.usage.speed. Current Opus 5/4.8 fast rows use the 2x fast rate; historical retained Opus 4.6/4.7 fast rows use 6x. Cache reads and both cache-write TTLs stack on that effective base rate, so Total, Saved, Wasted, and Net remain internally consistent. Standard, missing, malformed, unsupported, and recorded-cost rows keep their existing behavior; Claude has no inferred or user-selected speed flag.

Read the full file on GitHub · 198 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 198 lines · 0 tokens per session scan A d45d2a306ea6

Subscribe to this mod's changes

cache-report is a command published in the GitHub repository omrikais/cctally (5 stars, last pushed 3d ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 2,642 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.