Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/tyox-all/weave_protocol/agentsecbenchnpx skills add Tyox-all/Weave_Protocol --skill agentsecbenchgit clone --depth 1 https://github.com/Tyox-all/Weave_ProtocolWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00622 |
| Opus 5 | $0.00000 | $0.00311 |
| Sonnet 5 | $0.00000 | $0.00124 |
| Haiku 4.5 | $0.00000 | $0.00062 |
Grade A, and why
agentsecbench scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 70 lines — stays where its author put it; the contents beside it link to each section on GitHub.
🎯 AgentSecBench skill — standardized agent security benchmarking
You have access to @weave_protocol/agentsecbench — a benchmark package that runs locked, versioned attack suites against AI agents and produces tier-graded reports.
When to invoke
Use AgentSecBench when the user:
- Wants a citable score for an agent ("Tier B, 87/100 on ASB-Browser-v1")
- Asks to compare two agents or two versions of the same agent
- Needs a paste-ready security review artifact for a blog, RFP, or audit
- Asks whether their WARD policy actually does anything
- Wants to detect a regression between runs (e.g. in CI)
For raw attack execution without the interpretation layer, use @weave_protocol/adversary directly instead.
Key commands
# Run the canonical browser suite
agentsecbench run --suite=ASB-Browser-v1
# Save report
agentsecbench run --json=./report.json --md=./report.md
# Measure WARD contribution (runs twice)
agentsecbench run --measure-ward-delta
# Compare two reports
agentsecbench compare baseline.json new.json
# Show suite info / list suites
agentsecbench suite ASB-Browser-v1
agentsecbench suite # lists all
What's in a Report
- Tier (A/B/C/D/F) — headline grade
- Score (0-100) — raw Adversary score
- Category gaps — which attack classes failed
- Trophy performance — pass/fail against 4 named real-world attacks (Atlan, EchoLeak, Brave/Comet, Forcepoint)
- WARD delta (optional) — does the policy contribute to defense?
- Interpretation prose — paste-ready summary
- Full Adversary scorecard — embedded for auditability
Suites are locked
ASB-Browser-v1's 40 attacks will never change. Methodology improvements ship as ASB-Browser-v2. This is the central property that makes scores comparable.
Programmatic API
import { runSuite, ASB_BROWSER_V1, renderMarkdownReport } from '@weave_protocol/agentsecbench';
import { BrowserTarget } from '@weave_protocol/adversary';
const report = await runSuite({
target: new BrowserTarget({ runAgent }),
suite: ASB_BROWSER_V1,
targetMeta: { name: 'My Agent', type: 'browser-agent' },
measureWardDelta: true,
});
console.log(renderMarkdownReport(report));
What ships with it
18 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- .gitignore 54 B
- .npmrc 0 B
- METHODOLOGY.md 8.6 KB
- package.json 1.3 KB
- README.md 7.8 KB
- src/cli.ts 7.5 KB runs code
- src/compare/index.ts 8.0 KB runs code
- src/index.ts 1.2 KB runs code
- src/report/index.ts 157 B runs code
- src/report/interpret.ts 3.2 KB runs code
- src/report/json.ts 147 B runs code
- src/report/markdown.ts 5.0 KB runs code
- src/runner.ts 4.0 KB runs code
- src/suites/browser-v1/manifest.ts 4.7 KB runs code
- src/suites/index.ts 405 B runs code
- src/tier.ts 2.7 KB runs code
- src/types.ts 6.1 KB runs code
- tsconfig.json 459 B
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 70 lines · 0 tokens per session scan A ecab3ad567de
agentsecbench is a skill published in the GitHub repository Tyox-all/Weave_Protocol (0 stars, last pushed 9d ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 622 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
rtc-balance
Check RustChain wallet balance, epoch info, and network status via the public RPC.
panguard
AI agent security platform — audit skills, scan for threats, and run 24/7 protection with 9,700+ detection rules.
lictor-security-check
Pre-release security audit for ANY project — AI-built or hand-written, web or mobile. Scans the codebase for the full range of real-world risks that get apps breached: leaked API & AI-provider keys, exposed configs/secrets, broken auth & access control (IDOR), injection (SQL/XSS/command), SSRF, open databases & cloud…
lictor-rotate
Walks the user through rotating a leaked API key — step by step, provider-specific. Knows the exact URL to visit, the exact button to click, and how to verify the rotation worked. Supports Stripe, OpenAI, Anthropic, Google Cloud / AI Studio, GitHub, AWS, Slack, Supabase, Firebase, Postmark, and generic OAuth providers.
alchemy-agentic-gateway
Use when accessing Alchemy APIs for RPC calls, token balances, NFT metadata, asset transfers, transaction simulation, or Alchemy-specific features. Also use when the user mentions "SIWE", "SIWS", "x402", "MPP", "mppx", or "agentic gateway" — this skill covers wallet-based auth flows for Alchemy's x402 and MPP…
alchemy-api
Integrates Alchemy blockchain APIs using an API key. Requires $ALCHEMYAPIKEY to be set; if unavailable, use the alchemy-agentic-gateway skill instead. Use when user asks about EVM JSON-RPC calls, token balances, NFT ownership or metadata, transfer history, token prices, portfolio data, transaction simulation…