Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/claesbackman/ai-research-feedback/explain-diffnpx skills add claesbackman/AI-research-feedback --skill explain-diffgit clone --depth 1 https://github.com/claesbackman/AI-research-feedbackWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/claesbackman/ai-research-feedback/explain-diff)<a href="https://agentmods.dev/skills/claesbackman/ai-research-feedback/explain-diff"><img src="https://agentmods.dev/badge/skills/claesbackman/ai-research-feedback/explain-diff.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00051 | $0.01827 |
| Opus 5 | $0.00026 | $0.00914 |
| Sonnet 5 | $0.00010 | $0.00365 |
| Haiku 4.5 | $0.00005 | $0.00183 |
Grade A, and why
explain-diff scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 99 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Explain a Change
Produce one self-contained HTML page that explains a code change well enough that someone who did not write it could defend it, ending in a quiz that tests whether they actually followed.
Scope
Set the target from $ARGUMENTS:
- No argument — uncommitted working tree against
HEAD. Usegit status --shortandgit diff(addgit diff --cachedif anything is staged). - One ref (
abc123,main) — that ref against the working tree, or againstHEADif the tree is clean. - A range (
main..HEAD,abc123..def456) — use it as given.
New files are part of the change. git diff shows nothing for an untracked file, so a change that only adds scripts looks empty. Always read git status --short for ?? entries and treat every new file as an addition to explain: read it in full, since there is no diff to read. A new robustness script is exactly the kind of change worth explaining, and it is the one the diff will hide.
Stop only when git status --short and the diff are both empty.
Read from scratch
Work only from the code and the diff. If the conversation already contains an account of this change — yours or the user's — ignore it and read the files. A summary written earlier is a hypothesis, not evidence.
Do not stop at the diff hunks. For every changed block, read the function, script, or stage that contains it, and read what consumes its output. A three-line addition to a pipeline stage is usually only comprehensible from the two files on either side of it.
Verify before you claim
The single most common failure here is reporting a change in results that did not happen, or missing one that did. Check, do not infer:
- Generated files. Tables,
.texbodies, and rendered figures often show as modified when only a timestamp comment or a nondeterministic tie-break changed. Diff them and look at what actually differs before calling it a result change. - Binary artifacts. If an image changed, compare both versions visually. The old version lives in git rather than on disk, so extract it first —
git show HEAD:path/to/figure.png > /tmp/old-figure.png— then read the extracted copy alongside the working-tree one. A re-render with identical content is not a finding. - Logs. Compare sample sizes, observation counts, and headline coefficients between the old and new run. Identical numbers are the proof that an added block was inert; do not assert it from code reading alone.
- Numbers you quote. Every figure in the page should come from an output file you opened, not from the prose of the diff. If the change writes a new CSV, read it.
- Samples. When a new block filters data, compare its filter line by line against the filters of the existing estimation samples. Differences of one clause are the interesting ones.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 99 lines · 51 tokens per session scan A d39eb5a3f205
explain-diff is a skill published in the GitHub repository claesbackman/AI-research-feedback (476 stars, last pushed 6d ago), licensed MIT. It adds 51 tokens to every session and 1,827 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.
chat-perf
Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…