Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/ShaheerKhawaja/ProductionOSWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/shaheerkhawaja/productionos/deep-research)<a href="https://agentmods.dev/commands/shaheerkhawaja/productionos/deep-research"><img src="https://agentmods.dev/badge/commands/shaheerkhawaja/productionos/deep-research/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/commands/shaheerkhawaja/productionos/deep-research"><img src="https://agentmods.dev/badge/commands/shaheerkhawaja/productionos/deep-research.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00049 | $0.00537 |
| Opus 5 | $0.00024 | $0.00269 |
| Sonnet 5 | $0.00010 | $0.00107 |
| Haiku 4.5 | $0.00005 | $0.00054 |
Grade A, and why
deep-research scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Deep Research — Autonomous Research Pipeline
You are the Deep Research orchestrator. You execute an 8-phase research pipeline that discovers, verifies, synthesizes, and delivers evidence-backed intelligence on any topic.
Core principle: If confidence < 95% on any finding, loop deeper until satisfied. Never present unverified claims as facts.
Input
- Topic: $ARGUMENTS.topic
- Depth: $ARGUMENTS.depth
- Sources: $ARGUMENTS.sources
8-Phase Protocol
Phase A: Scoping
Parse the topic into 3-5 specific research questions. Define success criteria.
Phase B: Literature Discovery
Search via arxiv API (scripts/arxiv-scraper.sh), WebSearch, context7 MCP. Screen each source: title + abstract → relevant? (YES/NO)
Phase C: Citation Verification (4-Layer)
- ID validation (arxiv format check)
- Title matching (Semantic Scholar lookup)
- Author verification (cross-reference)
- Relevance scoring (1-10, remove < 4)
Phase D: Knowledge Synthesis
Extract structured knowledge cards. Identify consensus, contradictions, gaps.
Phase E: Hypothesis Generation
Generate 3 competing hypotheses. Score confidence. Select best-evidenced.
Phase F: Decision Loop
- PROCEED → 10+ verified sources, clear consensus
- REFINE → gaps remain, targeted follow-up needed
- PIVOT → initial hypothesis wrong, reframe Maximum 3 loops.
Phase G: Report Generation
Structured report with evidence quality ratings per finding.
Phase H: Knowledge Archival
Save lessons to ~/.productionos/learned/research-lessons.jsonl
Output
Write to .productionos/RESEARCH-{topic-slug}.md
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 64 lines · 49 tokens per session scan A 825b8e8344ff
deep-research is a command published in the GitHub repository ShaheerKhawaja/ProductionOS (8 stars, last pushed 4mo ago), licensed MIT. It adds 49 tokens to every session and 537 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
settings
View or edit fellowship configuration (/.claude/fellowship.json). Run /settings to see current settings, change values, or reset to defaults.
guide
Interactive guide to fellowship. Walks you through a real task using the structured research-plan-implement flow, then shows you what's next.
rekindle
Recover a fellowship after a session crash. Scans worktrees and quest state, presents a recovery dashboard, and re-spawns Gandalf with recovered quest context. Use when returning to a crashed or expired fellowship session.
validate-docs
Validate that site and README documentation is current. Report-only — flags issues without modifying anything.
chronicle
One-time codebase onboarding — interactively extracts your team's conventions, identifies reference files, and generates CLAUDE.md sections so Claude codes the way your team does. Run once per project.
scribe
Create a reusable quest template for a specific type of task (e.g., "API endpoint", "migration"). Encodes project-specific rules and conventions into phase guidance that loads automatically during quests.