Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/hahaxiang27/FlowHarnessWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/hahaxiang27/flowharness/harness.e2e)<a href="https://agentmods.dev/commands/hahaxiang27/flowharness/harness.e2e"><img src="https://agentmods.dev/badge/commands/hahaxiang27/flowharness/harness.e2e.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00082 | $0.01745 |
| Opus 5 | $0.00041 | $0.00873 |
| Sonnet 5 | $0.00016 | $0.00349 |
| Haiku 4.5 | $0.00008 | $0.00175 |
Grade A, and why
harness-e2e scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Harness [E2E_TOOL] E2E æµè¯
**ä¸ä¸æç®¡ç?*: ð æ¸ 空ä¸ä¸æ?â?使ç¨åä»£çæ§è¡ï¼ç¡®ä¿å¹²åç¯å¢
æä»¤
为å½å宿çç¨æ·æ
äºç¼åå¹¶è¿è¡?[E2E_TOOL] E2E æµè¯ã?
è¾å ¥åæ°
$ARGUMENTS â?ç¨æ·æ äºç¼å·ï¼å¦ "US1"ï¼æ "sprint" 表示å½å Sprint çææ?E2E
æ§è¡æ¥éª¤
ä½¿ç¨ Agent å·¥å ·å¯å¨å代çï¼ä¼ å ¥ä»¥ä¸ä»»å¡ï¼?
ä½ æ¯ Harness E2E æµè¯æ§è¡å¨ã?
1. 读å `specs/[FEATURE_ID]/spec.md`ï¼æ¾å?{USç¼å·} çå
¨é¨éªæ¶åºæ?2. 读å `.harness/prompts/evaluator.md` ç?Level 3 模æ¿ï¼äºè§?E2E ç¼åè§å
3. 读å `.harness/prompts/generator.md` ç?E2E 模æ¿ï¼äºè§?`[E2E_TOOL]` æµè¯çæè§è
æ§è¡ï¼?a. 为æ¯ä¸ªéªæ¶åºæ¯ç Given/When/Then ç¼åä¸ä¸?test case
b. ä½¿ç¨ Page Object 模å¼ç»ç»é¡µé¢äº¤äº
c. åå
¥ `[TEST_ROOT]/e2e/{story-name}[TEST_FILE_SUFFIX]`
d. è¿è¡ `[E2E_COMMAND]`
e. å¦æå¤±è´¥ï¼ä¿®æ£æµè¯ä»£ç åéè·
f. 确认éè¿åéè·?次éªè¯ç¨³å®æ?g. å
³é®æ¥éª¤æªå¾ä¿åå?`[E2E_SCREENSHOT_DIR]`
æ¥åï¼?- æµè¯ç¨ä¾æ»æ°ãéè¿æ°ã失败æ°
- 失败ç¨ä¾ç详æ
åæªå¾è·¯å¾
- ç¨³å®æ§ç»æï¼3次éè·æ¯å¦å
¨é¨éè¿ï¼?- æ´æ° sprint-*-progress.md ä¸ç E2E éªè¯ç¶æ?```
### 注æ
- E2E æµè¯éè¦ç¸å
³åºç¨ãç¨æ·çé¢åä¾èµæå¡é½å¨è¿è¡
- 妿åºç¨æç¨æ·ç颿ªå¯å¨ï¼å
æç¤ºç¨æ·æ§è¡ `[APP_START_COMMAND]` / `[UI_START_COMMAND]`
- 馿¬¡è¿è¡å¯è½éè¦å®è£
æåå§å?`[E2E_TOOL]` çè¿è¡æ¶èµæº
## SDD Step Gate
When specs/{REQUIREMENT_ID}/dashboard-state.json exists (SDD workflow active), after this command completes follow .harness/prompts/command-step-gate.md:
1. Update dashboard-state.json and dashboard.html when applicable.
2. Mark this command done, next step next, workflow_plan.phase = awaiting_user.
3. **Stop immediately** - do not chain the next internal command in the same turn.
4. Hand off with .harness/prompts/step-gate-handoff.md.
Skip only for standalone invocation without dashboard state, or when the user explicitly asks to batch remaining steps.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 57 lines · 0 tokens per session scan A 159f1092b002
harness-e2e is a command published in the GitHub repository hahaxiang27/FlowHarness (4 stars, last pushed 2mo ago), licensed MIT. It adds 82 tokens to every session and 1,745 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
paul:verify
Guide manual user acceptance testing of recently built features.
webapp-testing
Toolkit for interacting with and testing local web applications using Playwright. Supports verifying frontend functionality, debugging UI behavior, capturing browser screenshots, and viewing browser logs.
ui-snapshot.template
This prompt was authored for Claude-style slash workflows. In Codex runtime, adapt tool calls as follows.
laravel-playwright
E2E Playwright patterns; use the laravel:e2e-playwright skill exactly as written.
qa
Smoke or browser-walk a running app. Report only. Do not implement. Do not merge.
verify-pr
Verify a PR's frontend changes through browser automation.