Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/tobiasblask/open-paper-machine/evaluate-ideagit clone --depth 1 https://github.com/TobiasBlask/open-paper-machineWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/tobiasblask/open-paper-machine/evaluate-idea)<a href="https://agentmods.dev/commands/tobiasblask/open-paper-machine/evaluate-idea"><img src="https://agentmods.dev/badge/commands/tobiasblask/open-paper-machine/evaluate-idea.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00084 | $0.00344 |
| Opus 5 | $0.00042 | $0.00172 |
| Sonnet 5 | $0.00017 | $0.00069 |
| Haiku 4.5 | $0.00008 | $0.00034 |
Grade A, and why
evaluate-idea scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
/evaluate-idea — Research Idea Stress-Test
Read skills/idea-engine/SKILL.md and execute the full 6-phase idea evaluation pipeline.
The user provides a research idea, topic, or question. Run the complete evaluation: Phase 1 (SEED) → Phase 2 (DIVERGE) → Phase 3 (EVALUATE) → Phase 4 (DEEPEN) → Phase 5 (FRAME) → Phase 6 (DECIDE).
Save the evaluation to research-evaluations/YYYY-MM-DD-<topic-slug>.md.
If the verdict is PURSUE and the user wants to proceed, hand off to the paper machine pipeline (Phase 1: Reconnaissance).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 37 lines · 84 tokens per session scan A d47de697f049
evaluate-idea is a command published in the GitHub repository TobiasBlask/open-paper-machine (18 stars, last pushed 4mo ago), licensed MIT. It adds 84 tokens to every session and 344 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
psql-query
Run ad-hoc PostgreSQL analytics queries against dev/test database.
github-actions
Design, review, secure, and debug GitHub Actions workflows — reusable workflows, OIDC federation, SHA pinning, token scoping, promotion orchestration, and CI failure diagnosis.
fluxcd
FluxCD entry point — routes to the right workflow based on what you need. Live cluster issue → structured 5-workflow debug trace. Repo health check → 6-phase audit (discovery, validation, API compliance, best practices, security). Helm chart review → helmchart. Starts by asking one question to confirm the right mode.
genshijin-compress
Markdown・テキストファイルを原始人形式へ安全圧縮.
import
Import a shared context bundle, or take a teammate's newer copy of one you already have.
scan
Launch a comprehensive security audit on the current project. Detects vulnerabilities, scans dependencies, checks OWASP Top 10, and generates a structured report.