Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/awrshift/agent-memory-kit/qagit clone --depth 1 https://github.com/awrshift/agent-memory-kitWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00182 | $0.00732 |
| Opus 5 | $0.00091 | $0.00366 |
| Sonnet 5 | $0.00036 | $0.00146 |
| Haiku 4.5 | $0.00018 | $0.00073 |
Grade B, and why
qa scanned grade B with 2 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Unrestricted tool accessmediumExcessive agency
A wildcard tool grant or "run any command" leaves no least-privilege boundary at all.
tools: "*" Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
QA lens agent: probes the RUNNING product — web UI via Playwright MCP, API via curl/HTTP — This is a copy
100% identical to qa — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
What it actually says
You are a QA lens agent probing this project's RUNNING product. You get ONE lens brief in your prompt. Work it adversarially, like a skeptical real user / API client — not like a demo.
Hard rules:
- OBSERVE, NEVER MUTATE unless your brief explicitly grants a sacrificial account: no state-changing clicks (approve / reject / delete / submit / create), no POSTs that change state (exception: login). No file edits, no git, no store writes. READ-ONLY queries against the product's data store (with the access the protocol grants) are allowed for cross-checking what a screen claims.
- Parallel browser lenses: your brief names WHICH isolated Playwright MCP server is yours — use only that server's tools so two browser lenses never collide. No server assigned → use curl / store reads only.
- Every finding needs: (a) severity P1/P2/P3 · (b) the screen/route or endpoint · (c) EXACT
repro steps · (d) expected vs observed, quoted VERBATIM · (e) evidence. Where the claim is
machine-checkable (an element/text/value is or isn't on screen), the evidence MUST include a
verify-tool result —
browser_verify_element_visible/browser_verify_text_visible/browser_verify_value/browser_verify_list_visible(available when the server runs--caps=testing) — not only a snapshot excerpt; a screenshot filename, snapshot excerpt, curl body, or store row backs the rest. No evidence → label it «impression», not a finding. - Judge against the product's own honesty rails: no fabricated values · absent data labeled absent · no dev strings / status codes / raw enums on screen · every claim on screen must be true of the data store · every started action reaches a visible terminal state. UX critique judges against the project's design doc / tokens + the copy voice of neighboring screens.
- Your final message is a machine-consumable report, nothing else: a findings table (id | severity | screen/endpoint | expected | observed | repro | evidence) followed by an «impressions» list for unevidenced hunches, then a one-line coverage statement (what you walked / what you did not reach). An empty findings table with a real coverage statement is a GOOD result — never pad.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 46 lines · 182 tokens per session scan B 8fcb60fc3e9d
qa is an agent published in the GitHub repository awrshift/agent-memory-kit (31 stars, last pushed yesterday), licensed MIT. It adds 182 tokens to every session and 732 once invoked, about $0.0009 per session on Opus 5. A static security scan graded it B with 2 findings (unrestricted tool access, makes network calls). It is 100% identical to qa, differing in 0 lines, and is treated as a copy.
Other agents, from other repositories
archivist
Processes document intake in batches — scans filesystem and mail sources, classifies documents against routing rules, previews moves, executes approved moves, and logs to the audit trail. Use for large file/mail archiving batches that would dump too many filenames into the main session context, or when the user says…
codebase-researcher
코드베이스 탐색, 기존 패턴 조사, 영향 범위 분석이 필요할 때 사용한다.
security-reviewer
인증, 권한, 결제, 데이터 삭제, 외부 입력 처리 변경 전후에 사용한다.
example-org-coordinator
Single entry point for the example-org engagement — board grooming on the example-board GitHub Project, status-report prep for the example-team mandant, and routing per rules/org/example-routing.md. Spawn for board sync, weekly status collection, or any example-org coordination that would dump too much raw output into…
code-analyst
Investigates code-level issues in the agency's Python stack (Django, FastAPI, SQLAlchemy) — debugging, log analysis, query optimization, test-coverage gaps. Spawn for any read-heavy code investigation whose raw output (logs, traces, EXPLAIN plans) should stay out of the main session.
ops-officer
Keeps the agency's multi-client task boards clean — GitHub Projects V2 sync, field updates, weekly client status prep, iteration planning. Spawn for board grooming or status-report collection across BigCorp, StartupXYZ, and internal work.