Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/anhnguyen0905/codex-mcp/exec-deliverablenpx skills add anhnguyen0905/codex-mcp --skill exec-deliverablegit clone --depth 1 https://github.com/anhnguyen0905/codex-mcpWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/anhnguyen0905/codex-mcp/exec-deliverable)<a href="https://agentmods.dev/skills/anhnguyen0905/codex-mcp/exec-deliverable"><img src="https://agentmods.dev/badge/skills/anhnguyen0905/codex-mcp/exec-deliverable.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00066 | $0.01125 |
| Opus 5 | $0.00033 | $0.00562 |
| Sonnet 5 | $0.00013 | $0.00225 |
| Haiku 4.5 | $0.00007 | $0.00112 |
Grade A, and why
exec-deliverable scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 80 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Deliverable Standards for Non-Code Output (embed into Codex prompts)
The core flow assumes code, so Phase 4 embeds exec-coding-standards + exec-self-testing by
default. When a task's output is content, not code (a data analysis, a marketing brief, a
launch email, documentation, a research summary, a plan), those blocks don't fit. Embed THIS block
instead — plus any domain skills selected via skill-selection.
Standards block
Deliverable standards (mandatory — this task produces content, not code):
- Accuracy first: every factual claim, number, or quote must be verifiable. Do not invent data,
statistics, sources, quotes, or citations. If a figure is estimated or assumed, label it as such.
- Ground it: derive conclusions from the provided inputs/files/data; when you use an external fact,
name the source. Distinguish "the data shows X" from "I recommend Y".
- Structure for the audience: lead with the answer/recommendation, then support it; use headings,
short paragraphs, and lists so the reader can scan. Match the requested format and length exactly.
- Match voice & conventions: follow the project's existing tone, terminology, and templates (read a
sample first) over a generic voice. Respect brand/style guides when provided.
- Scope discipline: deliver what the task asked for and nothing extra; flag gaps or missing inputs
rather than filling them with speculation.
- No fabrication of authority: never imply endorsement, real people's words, or official records
that don't exist.
Data tooling block (embed additionally when the task processes a dataset)
Embed this block whenever the task reads or transforms data files beyond ~50 MB, in ANY lane — a content task analyzing an export, or a code task that happens to crunch data. Tool choice follows the data, not the repo's language: a TypeScript project does not mean Node scripts are the right way to scan an 800 MB CSV.
Data tooling rules (mandatory — this task processes a large dataset):
- Measure before choosing: run `du -h` on the inputs (and `wc -l` when cheap) BEFORE picking
tooling; these rules bind when inputs exceed ~50 MB — never guess sizes.
- Ingest once, query many: convert raw CSV/JSON exports into columnar form first (DuckDB database
file or Parquet — e.g. `duckdb analysis.duckdb "CREATE TABLE events AS SELECT * FROM
read_csv('<file>', union_by_name=true)"`), then run every question as a query against that.
Never re-parse the raw file per question or per report.
- Never write row-by-row scan scripts (Node readline, Python line loops, etc.) over large raw
files when columnar tooling can express the aggregation — regardless of the project's language.
- Sample-first iteration: develop and debug every query/script against a small sample (e.g. the
first 10-50k rows) and run the full dataset exactly once, after the logic passes on the sample.
Report the full-pass wall-clock time and row count in the deliverable.
- One pass, many outputs: when several reports derive from the same raw data, build shared
intermediate tables (per-user, per-day aggregates) in the ingest step and point every report at
those — never give each report its own full scan of the raw file.
- Keep heavy I/O local: if the input lives in a cloud-synced folder (OneDrive, Dropbox, Google
Drive), copy it to a local temp dir before ingesting and write outputs locally; sync overhead
can multiply runtimes.
- Memory discipline: never accumulate per-row objects for the whole dataset in RAM; aggregate
incrementally or let the columnar engine do it.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 80 lines · 66 tokens per session scan A eae9337c9e4a
exec-deliverable is a skill published in the GitHub repository anhnguyen0905/codex-mcp (3 stars, last pushed 4d ago), licensed MIT. It adds 66 tokens to every session and 1,125 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
insight-error-page
Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…