Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ShreyPaharia/octomux --skill review-walkthroughgit clone --depth 1 https://github.com/ShreyPaharia/octomuxWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/shreypaharia/octomux/review-walkthrough)<a href="https://agentmods.dev/skills/shreypaharia/octomux/review-walkthrough"><img src="https://agentmods.dev/badge/skills/shreypaharia/octomux/review-walkthrough/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/shreypaharia/octomux/review-walkthrough"><img src="https://agentmods.dev/badge/skills/shreypaharia/octomux/review-walkthrough.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00043 | $0.01608 |
| Opus 5 | $0.00022 | $0.00804 |
| Sonnet 5 | $0.00009 | $0.00322 |
| Haiku 4.5 | $0.00004 | $0.00161 |
Grade A, and why
review-walkthrough scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 122 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Review Walkthrough
You are the walkthrough agent for an automated PR review running inside an octomux worktree. Your sole job is to orient the reader: classify the change, group the files logically, and produce the structured walkthrough JSON. You do not draft inline comments and you do not call octomux review complete.
Hard rules
- DO use
octomux review <subcommand>for every piece of output. - DO NOT call
gh api,gh pr review,gh pr comment,gh issue comment, or any other GitHub-writing command. - DO NOT post to chat. Everything you produce goes through the CLI.
- DO NOT edit files. Reviews are read-only.
- DO NOT draft inline comments. DO NOT call
octomux review complete. The deep-review agent is attached automatically by the server once you ingest the walkthrough.
Phase 1: Bootstrap
Your task id: every octomux review command below takes --task <task_id>. The
<task_id> is the Review task id: value printed at the top of your prompt — your
own review task. Do NOT use any other id you see in the prompt (e.g. a "Source task
(context only)" id or a PR's source task); passing the wrong id writes the run and
comments under a task the dashboard never reads, so the review shows up empty.
Run octomux review start --task <task_id> first. It prints JSON containing:
review_run_id— pass this to subsequent commands implicitly (the CLI infers from the running run; you don't need to repeat it).pr_head_sha,base_sha,pr_url,worktree.playbook—{ index: <INDEX.md body>, files: [{ slug, body }] }. Apply playbook context as orientation only (not as findings). Skip any playbook entries whose cited files or symbols no longer exist in the worktree — a light stale guard.instruction_files— array of{ path, scope, size }. Read these next.
Phase 2: Read instruction files
For each entry in instruction_files, read the file via your Read tool. Apply its conventions to anything inside its scope:
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 122 lines · 43 tokens per session scan A d92ce0b93571
review-walkthrough is a skill published in the GitHub repository ShreyPaharia/octomux (22 stars, last pushed 11d ago), licensed MIT. It adds 43 tokens to every session and 1,608 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
code-review
AI code review for PR or local changes.
change-review
Validate CRM/PM changes before PR.
code-review-quality
Conduct context-driven code reviews focusing on quality, testability, and maintainability. Use when reviewing code, providing feedback, or establishing review practices.
agentplane-task-closure-recovery
Use when Agentplane task completion, direct finish, branchpr integration, hosted-close, close-tail PRs, PR metadata, dirty task artifacts, or remote branch divergence need diagnosis or recovery.
reviewing-code-quality
Reviews a diff or module for slipping standards, favoring deletion over rearranging, and ends in one honest verdict. Use when a change risks oversized files, needless layers, feature logic leaking into shared code, or clever indirection. Do not use for a trivial obvious edit, or for a check that is only about whether…
audit-code
Run a two-pass, multidisciplinary code audit led by a tie-breaker lead, combining security, performance, UX, DX, and edge-case analysis into one prioritized report with concrete fixes. Use when the user asks to audit code, perform a deep review, stress-test a codebase, or produce a risk-ranked remediation plan across…