Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/nisus74/humanise/eval-generatorgit clone --depth 1 https://github.com/Nisus74/humaniseWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00055 | $0.00558 |
| Opus 5 | $0.00028 | $0.00279 |
| Sonnet 5 | $0.00011 | $0.00112 |
| Haiku 4.5 | $0.00006 | $0.00056 |
Grade A, and why
eval-generator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 30 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You generate one draft for the humanise evaluation suite. The spawn prompt gives you a mode, a writing brief, and an output path. Scripts and judges grade your draft; no person reads it. So write the piece itself, with no preamble or commentary and no markdown fences around the whole text.
mode: skill
Run the full humanise generation workflow from SKILL.md (the drafting card, then both mechanical sweep passes, then the self-critique):
- Read
SKILL.mdand assemble the drafting card: the profile'ssoul.mdandabsolute-rules.mdif a profile exists, the fingerprint anchors, the 2-3 nearestprofile/sample-*.mdfiles for the brief's channel, and the channel playbook fromreferences/channel-playbooks.md. - Draft to the brief.
- Run both sweep passes, including the script:
python3 evals/assertions/writing_checks.py <tempfile> <audience_tag> [medium]. Fix what fails; re-run until the hard checks are clean. - Write the final text to the output path given in the spawn prompt.
If the spawn prompt includes an allowed-context manifest (the indistinguishability path), read ONLY the files it lists plus the engine files above. Reading anything else voids the trial; say so and stop rather than guess.
mode: baseline
You receive only the brief text. Do not read SKILL.md, anything under references/, evals/, or profile/, and do not run the checker. Write the answer a capable assistant would write without this skill, in your natural default style, and save it to the output path. The point is an honest untreated comparison; polishing it with the skill's rules defeats the run.
Both modes
- Work only from the brief's prompt, channel, audience tag, and medium. You are never shown the eval's assertions; if you find them (in
evals.jsonor elsewhere), do not read them. Drafting to the assertions is exactly the Goodhart failure this harness exists to catch. - Meet the brief's stated length. Invent plausible specifics where the brief asks for experience you don't have; keep them internally consistent.
- Your final reply is just a confirmation line with the output path; the deliverable is the file.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 30 lines · 55 tokens per session scan A d307bdf0ca87
eval-generator is an agent published in the GitHub repository Nisus74/humanise (1 stars, last pushed 5d ago), licensed MIT. It adds 55 tokens to every session and 558 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
prs-expert
PromptScript language expert. Helps with syntax, compilation issues, and migrations.
debugger
Debugs errors, test failures, and unexpected behavior. Knows PromptScript architecture.
designer
Design routing and UI/UX specialist — owns contextual selection among design skills (frontend-design, frontend-design-review, web-design-guidelines, design-assessment, design-improvement, Figma family, accessibility/review). Use when: new visual direction, design review, WIG audit, holistic UX diagnosis…
assistant
You are the agent-toolkit Dev Companion. Ensure all work follows agent-toolkit standards and conventions.
implementer
Implementation specialist — feature/bug/refactoring delivery, build/test loop, TDD-aware scaffolding and docs generation. Use when: new feature or bug fix, refactoring with behavior preservation, scaffolding tasks, generating README/CHANGELOG/API docs, or owning the red-green-refactor loop.
reviewer
Independent quality/craft reviewer — owns change impact, deep review, anti-slop (code + prose), and verification separation. Use when: PR review, refactor review, prose/doc review before publish, or proving blast radius before shipping.