Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add shahidshabbir-se/my-pi-setup --skill agent-engineeringgit clone --depth 1 https://github.com/shahidshabbir-se/my-pi-setupWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/shahidshabbir-se/my-pi-setup/agent-engineering)<a href="https://agentmods.dev/skills/shahidshabbir-se/my-pi-setup/agent-engineering"><img src="https://agentmods.dev/badge/skills/shahidshabbir-se/my-pi-setup/agent-engineering/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/shahidshabbir-se/my-pi-setup/agent-engineering"><img src="https://agentmods.dev/badge/skills/shahidshabbir-se/my-pi-setup/agent-engineering.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00170 | $0.03375 |
| Opus 5.5 | $0.00068 | $0.01350 |
| Sonnet 5.5 | $0.00034 | $0.00675 |
| Haiku 4.5 | $0.00017 | $0.00337 |
Grade A, and why
agent-engineering scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 362 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Agent Engineering
A durable decision framework for your own agent system. It exists because you will forget the reasoning behind past choices — this skill is the memory of how to decide, not a record of what you've already built.
The one rule that matters more than any mechanism
Do not build agent infrastructure because it sounds useful. Let a real, repeated problem determine what gets promoted, and promote the smallest thing that reliably solves it.
Every section below exists to serve that rule. If you're ever unsure, come back to this sentence before adding anything.
First questions, always: is there evidence, and has this recurred?
Before consulting anything else, ask:
- What's the actual evidence this is a problem — not a hunch, a single transcript, a diff.
- Did this happen more than once, or am I reacting to a single event?
Default: don't promote a one-off mistake. Just fix it and move on. If you catch yourself wanting to build something after one occurrence, that's the overengineering instinct — stop and just fix the thing.
Exceptions require a concrete reason, not a hunch — either:
- the cost of recurrence would be unacceptable even once (security, data loss, irreversible action), or
- the mechanism is genuinely cheap and strongly justified even for one occurrence (e.g. a one-line hard constraint that removes a whole class of mistake outright).
Continue below if the answer is "yes, this keeps happening," "I genuinely don't know, and not knowing is itself the problem" (→ Observability), or one of the exceptions above applies.
Relationship to self-improve
self-improve
↓
"I found a specific recurring mistake.
How do I make this lesson durable?"
→ turns a confirmed recurring mistake into durable project
knowledge — normally a rule in AGENTS.md, or an update to an
appropriate skill.
agent-engineering (this skill)
↓
"What kind of agent-engineering mechanism
should exist for this problem, if any?"
→ resolves to: nothing, a rule (→ hand to self-improve), a skill,
a hard constraint, verification, an eval, an eval harness,
observability, a feature map, or a trust/automation step.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 362 lines · 170 tokens per session scan A f9eb4faef484
agent-engineering is a skill published in the GitHub repository shahidshabbir-se/my-pi-setup (2 stars, last pushed 4d ago), licensed MIT. It adds 170 tokens to every session and 3,375 once invoked, about $0.0007 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-27.
Other skills, from other repositories
release-notes
Write release notes for a tagged release — a business-impact summary of what changed for the user, plus a thanks section naming every contributor whose commits are in the release. Short by construction.
review-this-branch
Run the nohuman review gate (fresh-session adversarial reviewer + tamper guard) over the current branch or a GitHub pull request, with no server, no database, and no onboarding, and relay the pass/fail checklist.
audit-agents-skills
Audit Claude Code agents, skills, and commands for quality and production readiness. Use when evaluating skill quality, checking production readiness scores, or comparing agents against best-practice templates.
issue-triage
3-phase issue backlog management with audit, deep analysis, and validated triage actions. Use when triaging GitHub issues, sorting bug reports, cleaning up stale tickets, or detecting duplicate issues. Args: 'all' to analyze all, issue numbers to focus (e.g. '42 57'), 'en'/'fr' for language, no arg = audit only.
pr-triage
4-phase PR backlog management with audit, deep code review, validated comments, and optional worktree setup. Use when triaging pull requests, catching up on pending code reviews, or managing a backlog of open PRs. Args: 'all' to review all, PR numbers to focus (e.g. '42 57'), 'en'/'fr' for language, no arg = audit…
check-cache-bugs
Audit Claude Code setup for cache bugs (CC#40524): sentinel, --resume/--continue, attribution header + ArkNill B3/B4/B5.