Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add FarzamMohammadi/the-engineer --skill assess-rungit clone --depth 1 https://github.com/FarzamMohammadi/the-engineerWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/farzammohammadi/the-engineer/assess-run)<a href="https://agentmods.dev/skills/farzammohammadi/the-engineer/assess-run"><img src="https://agentmods.dev/badge/skills/farzammohammadi/the-engineer/assess-run/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/farzammohammadi/the-engineer/assess-run"><img src="https://agentmods.dev/badge/skills/farzammohammadi/the-engineer/assess-run.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00192 | $0.01863 |
| Opus 5 | $0.00096 | $0.00932 |
| Sonnet 5 | $0.00038 | $0.00373 |
| Haiku 4.5 | $0.00019 | $0.00186 |
Grade A, and why
assess-run scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 103 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Assess Run
Post-mortem one run of The Engineer — a task moving through (or stopped in, or crashed out of) the pipeline — and turn it into improvements to The Engineer itself.
Hold two questions apart:
- Was this run good? — the machinery, each phase, the commits, the PR.
- What should change in The Engineer so the next run is better?
Deliverable: a short report at .claude/temp/assess-run/<task-short>-<date>.md, then walk it with the owner.
Principles
- The trail is the evidence. Judge only from what the system recorded — DB, traces, commits, PR. If you can't reconstruct the story from the trail, that gap is itself a finding.
- Honest, not flattering. A post-mortem, not a status report. If a phase phoned it in, say so with the artifact. If it was clean, say why, so it's repeatable.
- Improvements are principles, never patches. A prompt fix must sharpen the agent's judgment across the next hundred tasks, not the one case. "Gather must persist acceptance criteria to the structured field" generalizes; "when the issue mentions a prefix, check the classifier" does not.
- Tiered, not exhaustive. Reconstruct cheaply from the DB first; open a raw trace only when a hypothesis needs the agent's actual words. Most of a clean run isn't worth reading.
1. Pick the run and the mode
The argument is a task id or prefix. With none, run triage with no argument to list recent runs and ask. The owner often hands you the situation ("stopped at self-review", "crashed") — treat it as the lead, confirm it against the trail.
Mode shapes emphasis: completed (did it finish good, not just finish?) · stopped (active/blocked/queued — why, and is the stop legitimate or a defect?) · crashed (root cause, and did recovery work?).
Home defaults to ~/.engineer; honor --home / $ENGINEER_HOME.
2. Triage — reconstruct the lifecycle
.claude/skills/assess-run/scripts/triage.sh <task-id-or-prefix> # --home <dir> for a custom data dir
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 103 lines · 192 tokens per session scan A acc682890266
assess-run is a skill published in the GitHub repository FarzamMohammadi/the-engineer (12 stars, last pushed 2mo ago), licensed MIT. It adds 192 tokens to every session and 1,863 once invoked, about $0.0010 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
doubt-driven-review
In-flight adversarial check on a non-trivial decision BEFORE it stands — distinct from post-hoc review of a finished diff. Use on "stress-test this decision", "are we sure about this", "verify before commit", "poke holes in this", when working in unfamiliar code, or before an irreversible step (migration, prod deploy…
release-cut
Cut a new pi-agent-dashboard release: promote ## [Unreleased] in CHANGELOG.md, bump every workspace package.json per SemVer, commit, tag v , and push — triggering the Release workflow that publishes every non-private workspace, builds the Electron artifacts, and creates a GitHub Release. Use on "cut a release"…
spec-coherence-check
Sweep all active OpenSpec proposals for staleness, conflicts, and obsolescence against the current codebase and archived changes. Use when proposals may be outdated, when checking cross-proposal conflicts, or before starting a batch of implementations. Produces a gap-analysis report, updates a priority queue file, and…
ship-it
Worktree-side implementation orchestrator for an OpenSpec change. Idempotent: gates automated scenarios on filesystem reality, owns the red-test fix loop, runs the docker harness with always-teardown, then drives ship-change inline. Escape hatch writes SHIPITBLOCKED.md. Runnable headless. Triggers: "ship it", "build…
faq-mine
Mine docs/faq.md from README.md, docs/.md, and the pi-hermes memory stores. Dispatches @fast subagents per source, dedupes against the existing FAQ, and merges entries in caveman style. Use when asked to "build / regenerate / extend the FAQ", "mine docs into FAQ", "mine hermes memory into FAQ", "surface runtime…
session-to-guideline
Turn a pi session into a Markdown "how-we-did-it" collaboration guideline: reads the session's JSONL transcript and synthesizes a reusable playbook of which prompts worked, what had to be steered, and how to reproduce the result faster. Use when: "document this session", "write up how we did X with the AI", "make a…