Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add adonai-labs/agent-runway --skill po-evalgit clone --depth 1 https://github.com/adonai-labs/agent-runwayWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/adonai-labs/agent-runway/po-eval)<a href="https://agentmods.dev/skills/adonai-labs/agent-runway/po-eval"><img src="https://agentmods.dev/badge/skills/adonai-labs/agent-runway/po-eval/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/adonai-labs/agent-runway/po-eval"><img src="https://agentmods.dev/badge/skills/adonai-labs/agent-runway/po-eval.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00070 | $0.00858 |
| Opus 5 | $0.00035 | $0.00429 |
| Sonnet 5 | $0.00014 | $0.00172 |
| Haiku 4.5 | $0.00007 | $0.00086 |
Grade A, and why
po-eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 123 lines — stays where its author put it; the contents beside it link to each section on GitHub.
PO Evaluation
Invoke Command
/po-eval [Jira key or path to .md]
Examples:
/po-eval PROJ-501/po-eval .agent-runway/specs/order-flow/spec.md/po-eval .agent-runway/specs/order-flow/tickets/task-03-email-notifications.md
Goal
Evaluate whether a spec or ticket is product-ready before delivery work starts.
This skill does not do technical code review. It evaluates product quality and delivery readiness.
Workflow - 4 Phases
Phase 1 - Load artefact and context
- Read the provided spec/ticket content completely.
- Extract explicit business objective, target user/persona, scope, constraints, and acceptance expectations.
- If business context exists in
.agent-runway/docs/business/*, use it to validate alignment. - If key sections are missing, continue evaluation and mark gaps clearly.
Phase 2 - Score product readiness
Evaluate the artefact against these criteria:
- Problem clarity:
- Is the user/business problem explicit and non-ambiguous?
- Outcome and value:
- Is the expected business/user outcome explicit?
- Is value described (impact, risk reduction, revenue, time saved, etc.)?
- Scope quality:
- Are in-scope and out-of-scope boundaries explicit?
- Are assumptions and constraints listed?
- Success metrics:
- Are measurable success indicators defined?
- Are acceptance outcomes testable at product level?
- Dependency and risk coverage:
- Are product dependencies identified (teams, upstream/downstream flows, legal/compliance, external services)?
- Are major product risks and mitigations listed?
- Release and adoption readiness:
- Are rollout expectations, communication needs, and operational readiness considered when relevant?
Phase 3 - Verdict
Set one verdict:
PRODUCT READY - YESPRODUCT READY - CONDITIONALPRODUCT READY - NO
Rules:
YES: all critical product criteria are covered with concrete detail.CONDITIONAL: core criteria pass, but non-blocking gaps remain.NO: one or more critical criteria are missing or too vague.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 123 lines · 70 tokens per session scan A bb0423b99430
po-eval is a skill published in the GitHub repository adonai-labs/agent-runway (2 stars, last pushed 22d ago), licensed MIT. It adds 70 tokens to every session and 858 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
sdd-tasks
Break an SDD change into implementation tasks. Trigger: orchestrator launches task planning for a change.
pr-blocker-summarizer
Summarizes open pull requests into a blockers-first standup digest. Activates when the user asks to summarize open PRs, find blocked pull requests, generate a PR standup, or triage review backlog from a PR export.
review-work
Post-implementation gate review: run manual QA on the real surface yourself, then launch ONE gate reviewer (never a panel) to audit goal, constraints, code quality, security, missed context, and QA evidence. Use before a PR handoff or when the user explicitly asks to review completed work.
ponytail-lazy-senior-dev
Applies the "lazy senior developer" mindset. Use this skill whenever generating, modifying, reviewing code, or fixing bugs to prioritize code reuse, minimalism, YAGNI principles, and root-cause fixes. Also use whenever the user says "ponytail", "be lazy", "lazy mode", "simplest solution", "minimal solution", "yagni"…
full-audit
Exhaustive, denominator-driven audit of an entire area (codebase, docs, memory, skills, DB, config). Runs a 6-phase pipeline: scope agreement + prior-map diff -> deterministic sweep (counts/versions/paths/parsing plus cross-index reconciliation) -> parallel read-only content review (citations forced, rule dry-run) ->…
tech-lead/code-review
A checklist-based method for examining code for correctness, security, performance, clarity, and ease of maintenance.