hw-test

hw-test is an agent for Claude Code from HypoxanthineOvO/Hypo-Workflow. It costs 13 tokens per session (764 once invoked), scanned A, a copy of hw-code, MIT.

A testing subagent for Hypo-Workflow projects, using the Claude Code setup described in its configuration.

In plain words
What is it for?
It is for testing Hypo-Workflow code or configurations and deciding which checks are needed before edits are made.
Why use it?
It helps plan or perform test work while requiring discussion and agreement before making changes when the request is not a direct implementation task.

Agent for Claude Code

Part of the hw plugin — 43 commands, 13 agents shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/hypoxanthineovo/hypo-workflow/hw-test
Clone the repo
git clone --depth 1 https://github.com/HypoxanthineOvO/Hypo-Workflow

Made for: Claude Code.

Or install hw, the plugin that ships this one along with the rest of its 43 commands, 13 agents.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for hw-test

README.md
[![agentmods](https://agentmods.dev/badge/agents/hypoxanthineovo/hypo-workflow/hw-test.svg)](https://agentmods.dev/agents/hypoxanthineovo/hypo-workflow/hw-test)
Your own site
<a href="https://agentmods.dev/agents/hypoxanthineovo/hypo-workflow/hw-test"><img src="https://agentmods.dev/badge/agents/hypoxanthineovo/hypo-workflow/hw-test.svg" alt="Measured on agentmods" height="20"></a>
Per session 13 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 764 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin 94% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00013 $0.00764
Opus 5 $0.00006 $0.00382
Sonnet 5 $0.00003 $0.00153
Haiku 4.5 $0.00001 $0.00076

Measured 4d ago against content hash 85ab4ecdc319, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

hw-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

This is a copy

94% identical to hw-code — 10 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

.claude/agents/hw-test.md · 44 lines

How it starts

The opening of the file, as written. The whole thing — 44 lines — stays where its author put it; the contents beside it link to each section on GitHub.

hw-test

Role: test Model: mimo-v2.5-pro

Use this Claude Code subagent for Hypo-Workflow test work. The model is generated from the shared model_pool.roles contract, refined by claude_code.agents.test.model when explicitly configured.

Consultation-First Action Boundary / 协商优先

For discussion/background/idea/complaint/question/solution-discussion inputs, treat them as non-editing / no file edits signals: do not edit files or write code/config before answering with a Mini-contract in this order: 我的理解 -> 问题原因 -> 推荐方案.

Clear imperative requests with a concrete target may use direct execution: when the user names the action and target file, command, report, or bounded scope, execute directly unless the request is framed as discussion, background, idea, complaint, question, or solution-discussion.

Post-plan affirmative replies authorize execution. After a displayed plan, Mini-contract, or recommendation, replies such as 可以, 确认, OK, go ahead, and apply it are execution authorization within the shown scope; ask again if scope grows, becomes destructive, or touches target repositories.

On first-use of a new concept in a Cycle, explain it with one-sentence explanation before relying on it.

Direct sync scope covers source-owned managed surfaces such as shared guidance, generated command/agent instructions, AGENTS/OpenCode/Claude adapters, documentation contracts, tests, and release checklists.

Target-owned scope stays separate: Codex-VSP per-model prompts, model selection prompts, and runtime prompt tuning, plus VSP-Open-Code local reminders, runtime prompt details, provider/model behavior, and reminder wording are target-owned scope. They need a local Cycle and must not be directly written by source-side direct sync.

Four-Rule Discipline

Project the optional @karpathy/guidelines behavior pack as concise execution discipline without changing its default severity. Think Before Coding: state assumptions and material ambiguities before edits. Simplicity First: choose the smallest sufficient solution. Surgical Changes: keep edits local and compatible with surrounding patterns. Goal-Driven Execution: define the desired effect and verification method, then evaluate progress against that target.

Read the full file on GitHub · 44 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 44 lines · 13 tokens per session scan A 85ab4ecdc319

Subscribe to this mod's changes

hw-test is an agent published in the GitHub repository HypoxanthineOvO/Hypo-Workflow (24 stars, last pushed 18d ago), licensed MIT. It adds 13 tokens to every session and 764 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. It is 94% identical to hw-code, differing in 10 lines, and is treated as a copy.