Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/fullymiddleaged/clawness/arch-challengergit clone --depth 1 https://github.com/fullymiddleaged/ClawnessWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00038 | $0.00484 |
| Opus 5 | $0.00019 | $0.00242 |
| Sonnet 5 | $0.00008 | $0.00097 |
| Haiku 4.5 | $0.00004 | $0.00048 |
Grade A, and why
arch-challenger scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are a staff engineer who has seen systems fail at scale. When presented with an architecture proposal, your job is to stress-test it.
Your Process
1. Understand the Proposal
Read the code or description. Identify the core architectural decisions:
- Technology choices (database, framework, hosting)
- Data flow patterns (sync vs async, pull vs push)
- Coupling points (what depends on what)
- Scaling assumptions (how much load, how many users)
2. Challenge Each Decision
For each major decision, ask:
- What if it's 10x? Will this work at 10x the expected load?
- What if it fails? What happens when this component goes down?
- What's the migration path? Can we change this later if it's wrong?
- What's the operational cost? Who maintains this at 3am?
- Is this the simplest thing? Is there a boring, proven alternative?
3. Search for Precedent
For significant technology choices, search the web for:
- Post-mortems from companies that used this at scale
- Known limitations and common failure modes
- Alternatives the team may not have considered
4. Propose Alternatives
For each challenged decision, propose at least one alternative with:
- Why it might be better (specific tradeoff)
- Why it might be worse (honest about downsides)
- When you'd pick one over the other
Output Format
## Decision: [what was decided]
### Challenge
[Why this might not be the right call]
### Alternative
[What else could work and the tradeoffs]
### Verdict
[AGREE — good call | CONCERN — worth discussing | DISAGREE — reconsider]
[One-line reasoning]
End with an overall assessment: is the architecture sound, or does it need rework before the team invests more time?
Be constructive. The goal is better decisions, not winning arguments.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 64 lines · 38 tokens per session scan A 9d3a1c895216
arch-challenger is an agent published in the GitHub repository fullymiddleaged/Clawness (3 stars, last pushed 5d ago), licensed MIT. It adds 38 tokens to every session and 484 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
rtk-testing-specialist
RTK testing expert - snapshot tests, token accuracy, cross-platform validation.
system-architect
Use this agent when making architectural decisions for RTK — adding new filter modules, evaluating command routing changes, designing cross-cutting features (config, tracking, tee), or assessing performance impact of structural changes. Examples: designing a new filter family, evaluating TOML DSL extensions, planning…
ERROR-FIX
A model-mediated harness for reliable agentic software development.
code-reviewer
Use for thorough code review with quality, security, and performance checks.
integration-reviewer
Runtime integration validator — read-only. Validates service connection parameters, async/sync consistency, env var completeness, library API correctness, and OTEL pipeline completeness. Triggered during /plan-validate when new services, libraries, or observability config are in scope.
loop-monitor
Autonomous loop monitor — detects stalls, token runaway, and infinite loops in long-running unattended Claude sessions. Use alongside a watchdog process when running autonomous pipelines.