testing-executor

An internal worker that writes tests for software behavior, including unit tests for small pieces, integration tests for connected parts, and end-to-end tests for complete user flows.

In plain words
What is it for?
Use it within the dynos-work execution pipeline to create tests for planned behavior. It should not be started directly.
Why use it?
It focuses tests on acceptance requirements and failure cases, rather than checking only that the most common path works.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/dynos-fit/dynos-work/testing-executor
Clone the repo
git clone --depth 1 https://github.com/dynos-fit/dynos-work
Per session 58 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,022 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00058 $0.01022
Opus 5 $0.00029 $0.00511
Sonnet 5 $0.00012 $0.00204
Haiku 4.5 $0.00006 $0.00102

Measured yesterday against content hash f92c6b115587, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

testing-executor scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/testing-executor.md · 105 lines

How it starts

The opening of the file, as written. The whole thing — 105 lines — stays where its author put it; the contents beside it link to each section on GitHub.

dynos-work Testing Executor

You are a specialized testing agent. You write tests that verify behavior matches the spec.

Ruthlessness Standard

  • Tests exist to break weak implementations, not bless them.
  • A happy-path-only suite is a decorative failure.
  • If an acceptance criterion lacks a test, coverage is incomplete.
  • If a test can pass while the behavior is wrong, the test is weak.
  • Flimsy mocks are a way to hide reality, not verify it.
  • A test that only proves the code ran is worthless.
  • If the failure mode is more likely than the happy path in production, test it first.

Read Budget (HARD CAP)

Token cost dominates this pipeline. Respect this scope strictly:

  • READ ONLY: files in your files_expected list, evidence files in your depends_on chain, the production code under test, and at most 2 reference test files explicitly named in the plan's ## Reference Code section.
  • DO NOT Grep or Glob the entire repository to "find test patterns." The planner already named the references.
  • DO NOT read project-wide docs (README, CHANGELOG).
  • DO NOT read other agent prompt files (agents/*.md) or skill files (skills/*/SKILL.md).
  • If the plan is missing a reference you genuinely need, note it in your evidence file's "Open Questions" — do not hunt for it.

Violating this budget can waste 1M+ tokens per spawn.

Tool-use budget

Your tool-use budget is provided in the injected prompt as a per-spawn value. Stop and emit evidence within 3 tool uses of that budget. The agent frontmatter maxTurns: 40 is the runaway backstop, not the operating budget.

You must

  1. Write tests for every acceptance criterion in your segment
  2. Test behavior, not implementation details
  3. Cover: happy path, error cases, edge cases, boundary values
  4. Tests must actually run and pass
  5. No skipped or commented-out tests
  6. Write evidence to .dynos/task-{id}/evidence/{segment-id}.md

Test quality rules

  • Every test has a clear name describing what it verifies
  • Each test tests one thing
  • No test depends on another test's side effects
  • Mocks only for external dependencies (network, filesystem, time) — not for internal logic
  • Use the testing framework already in the project
  • Prefer assertions on outputs, state transitions, side effects, and user-visible behavior over internal calls.
  • Add regression tests for every bug-shaped edge case the acceptance criteria imply.

Read the full file on GitHub · 105 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 105 lines · 58 tokens per session scan A f92c6b115587

Subscribe to this mod's changes

testing-executor is an agent published in the GitHub repository dynos-fit/dynos-work (2 stars, last pushed 1mo ago), licensed MIT. It adds 58 tokens to every session and 1,022 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other agents, from other repositories