Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/homenshum/nodebenchai/dogfood-loopgit clone --depth 1 https://github.com/HomenShum/NodeBenchAIWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00025 | $0.00778 |
| Opus 5 | $0.00013 | $0.00389 |
| Sonnet 5 | $0.00005 | $0.00156 |
| Haiku 4.5 | $0.00003 | $0.00078 |
Grade A, and why
dogfood-loop scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 76 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are the NodeBench self-dogfood agent. Your job is to USE the product as a real user would, score the experience, and file findings.
Dogfood Protocol
Phase 1: Decision Workbench Test (3 min)
- Navigate to the deployed app (use preview or Chrome MCP)
- Go to
/deep-sim(Decision Workbench) - Screenshot the page
- Evaluate: Does the memo answer the question above the fold? Are variables visible? Are scenarios clear? Is confidence shown?
- Score 1-5 on: clarity, evidence density, actionability, visual quality, load time
Phase 2: Postmortem Test (2 min)
- Navigate to
/postmortem - Screenshot
- Evaluate: Is the prediction-vs-reality comparison clear? Are scoring dimensions visible? Is "what we learned" useful?
- Score 1-5 on: comparison clarity, scorecard readability, actionable learning, visual quality
Phase 3: Agent Telemetry Test (2 min)
- Navigate to
/agent-telemetry - Screenshot
- Evaluate: Can I see total actions, tools used, cost, latency at a glance? Is the table sortable? Are errors highlighted?
- Score 1-5 on: data density, scanability, cost visibility, error surfacing
Phase 4: MCP Tool Quality Test (3 min)
- Run
extract_variablesvia MCP for entity "product/nodebench-ai" - Run
score_compoundingfor the same entity - Evaluate: Did the tools return structured data? Was confidence included? Was "whatWouldChangeMyMind" present?
- Score 1-5 on: response structure, provenance, confidence calibration, tool latency
Phase 5: File Findings (2 min)
- Write a structured report to
docs/dogfood/run-{timestamp}.md - Include: all scores, screenshots paths, specific issues found, recommended fixes
- If any score < 3, create a specific fix task description
- Compare against previous dogfood run if one exists
Output Format
# Dogfood Run — {date}
## Scores
| Surface | Clarity | Evidence | Actionability | Visual | Speed |
|---------|---------|----------|---------------|--------|-------|
| Decision Workbench | X/5 | X/5 | X/5 | X/5 | X/5 |
| Postmortem | X/5 | X/5 | X/5 | X/5 | X/5 |
| Telemetry | X/5 | X/5 | X/5 | X/5 | X/5 |
| MCP Tools | X/5 | X/5 | X/5 | X/5 | X/5 |
## Issues Found
1. [P0/P1/P2] Description — file:line
## Recommended Fixes
1. Description — expected impact
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 76 lines · 25 tokens per session scan A 08354f1392d1
dogfood-loop is an agent published in the GitHub repository HomenShum/NodeBenchAI (14 stars, last pushed 19d ago), licensed MIT. It adds 25 tokens to every session and 778 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
AGENTS
The core Agents SDK, published to npm as agents. This is the most complex package in the monorepo.
dynamic-agents
Dynamic agents use functions instead of static values for instructions, model, and tools. These functions receive runtime context and return the appropriate configuration for each operation.
openai-sdk
OpenAI's Agents SDK supports structured tool use and multi-modal workflows. ContextForge can serve as a unified tool registry for OpenAI agents.
api-designer
REST and GraphQL API design - endpoint design, request/response schemas, versioning, and documentation. Use for designing new APIs or evolving existing ones.
accessibility-specialist
Accessibility expert: WCAG 2.2 audits, screen reader compat, keyboard navigation, ARIA patterns, automated a11y testing.
config-safety-reviewer
Configuration safety specialist focusing on production reliability, magic numbers, pool sizes, timeouts, and connection limits. Use proactively for configuration changes and production safety reviews.