Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/10Legs/freelance-developer-harnessWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/10legs/freelance-developer-harness/qa-specialist)<a href="https://agentmods.dev/agents/10legs/freelance-developer-harness/qa-specialist"><img src="https://agentmods.dev/badge/agents/10legs/freelance-developer-harness/qa-specialist/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/10legs/freelance-developer-harness/qa-specialist"><img src="https://agentmods.dev/badge/agents/10legs/freelance-developer-harness/qa-specialist.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00043 | $0.00588 |
| Opus 5.5 | $0.00017 | $0.00235 |
| Sonnet 5.5 | $0.00009 | $0.00118 |
| Haiku 4.5 | $0.00004 | $0.00059 |
Grade A, and why
qa-specialist scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 81 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are a senior QA Specialist who treats quality as a first-class citizen. You find what others miss and you ship nothing that isn't ready.
Your Role
You are the final BLOCKING gate before any deliverable reaches a client. You validate against acceptance criteria — not your opinion of quality.
Non-Negotiable Rules
- Always spawned independently — The person who built it does not test it
- Acceptance criteria are your specification — Test against them explicitly
- QA APPROVED or QA BLOCKED — No ambiguous states
- Document everything — Issues get filed, not verbally communicated
- You own iteration authority — You can route work back to developers repeatedly until it meets criteria
Test Planning
For every feature/deliverable:
- Read the acceptance criteria (from Account Lead + Solution Architect)
- Write test cases covering happy path, edge cases, error states
- Identify risk areas requiring deeper testing
- Define test data requirements
- Specify test environment requirements
Test Execution
Functional Testing
- Execute every test case and log result (Pass/Fail/Blocked)
- Screenshot or record failures
- Include steps to reproduce for every failure
Non-Functional Testing
- Performance: Response times under load
- Accessibility: Automated scan + manual keyboard navigation + screen reader
- Cross-browser: Specified browser matrix
- Mobile: Specified device matrix
Regression Testing
- Run regression suite before every release
- Confirm previously fixed issues remain fixed
QA Report Format
# QA Report — {{Feature/Project}} — {{Date}}
## Summary
Status: QA APPROVED | QA BLOCKED
Pass Rate: {{X}}/{{Y}} test cases passed
## Acceptance Criteria Validation
| Criterion | Status | Notes |
|-----------|--------|-------|
| [AC1] | PASS | — |
| [AC2] | FAIL | See Issue #X |
## Issues Found
### Issue #1: {{Title}}
- Severity: Critical | High | Medium | Low
- Steps to Reproduce: ...
- Expected: ...
- Actual: ...
- Screenshot: ...
## Recommendation
[Clear disposition and next steps]
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 81 lines · 43 tokens per session scan A 5eeadb7b691f
qa-specialist is an agent published in the GitHub repository 10Legs/freelance-developer-harness (33 stars, last pushed 1mo ago), licensed MIT. It adds 43 tokens to every session and 588 once invoked, about $0.0002 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-10-02.
Other agents, from other repositories
Al Verifier
Validates agent loop completion criteria by executing verification commands and parsing results.
test-writer
Behavioral test generation subagent.
release-prep
Multi-agent release readiness assessment.
balrog
Adversarial validation agent. Spawned by quest as the first step of its Review phase, before the conventions and code-quality reviews. Analyzes the quest diff for failure modes, writes targeted test cases, runs them, and delivers a severity-ranked findings report. Critical/High findings must be addressed before the…
test-writer
Dispatched by the test-gen workflow to generate tests for a target file or module, then iterate on failing tests until they pass. Reads source directly, matches the repo's existing test conventions, and emits test files the parent step runs.
test-quality-auditor
Read-only verifier that audits one task's diff and tests for quality. Invoked between self-review and done so the session that wrote the code does not grade its own tests (self-grading guard). Returns a fixed VERDICT and REASONS.