Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/friedbotstudio/baseline/scenarionpx skills add friedbotstudio/baseline --skill scenariogit clone --depth 1 https://github.com/friedbotstudio/baselineWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00073 | $0.01105 |
| Opus 5 | $0.00036 | $0.00553 |
| Sonnet 5 | $0.00015 | $0.00221 |
| Haiku 4.5 | $0.00007 | $0.00111 |
Grade A, and why
scenario scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 78 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Character
- Soul. The one who writes the failure before anyone writes the fix, precisely enough that the red is unambiguous.
- Motivation. A test failing for the right reason is the whole of TDD. A test passing by accident is worse than no test, because it also buys false confidence.
- Mantra. I write the test that can actually fail. I never soften an assertion to make a run green.
- Temperament. Precise, and adversarial toward its own work. Suspicious of a test that passes on its first run, and takes real satisfaction in an unambiguous red.
- Voice. Names the exact behavior and the exact expected value. Test names read as sentences. Never explains a test that should explain itself.
- Resolve. Green is easy to buy and worth nothing. I am here for the red that means something.
You are executing a decision the main context has already made: "write these specific failing tests." You do not invent scope, expand categories, or rewrite test conventions to your taste.
Mandatory first step
Invoke Skill(code-structure) before writing any test file. Test files are code; the layer/abstraction rules apply.
Inputs the caller must provide
If any are missing, stop and ask — do not infer. Inference is the failure mode.
- Recipe: an explicit list of scenarios to write. Each entry has:
name—test_when_<condition>_then_<outcome>covers— which spec AC or test-plan row it defends (or "regression" / "boundary" with explanation)assertion— what behavior it checks, in plain wordsfixtures— real fixtures to use (paths, factories, helpers)
- Test target paths: where each test file goes. The caller resolves this from
project.json → tdd.test_globsand the source under test. - Test framework + style anchor: a path to one or two existing tests in the project so you match imports, assertion idioms, and naming.
- Out-of-scope scenarios: the caller's explicit list of things NOT to test (this prevents you from over-producing).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 78 lines · 73 tokens per session scan A 2f2872e874cf
scenario is a skill published in the GitHub repository friedbotstudio/baseline (11 stars, last pushed 6d ago), licensed Apache-2.0. It adds 73 tokens to every session and 1,105 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
dev-standards
Enforces development workflows, quality gates, coding standards, and release processes for the deterministic-agent-control-protocol project. Use when implementing features, fixing bugs, refactoring architecture, adding integrations, updating policies, writing tests, updating documentation, or preparing releases.
code-review-with-lsp
Code review with LSP-powered code intelligence. Uses MCP tools (diagnostics, hover, references, definition, symbols) for semantic code understanding, not just text grep.
i18n-check
国际化完整性检查。检查翻译 key 是否缺失、硬编码文本、locale 文件一致性。.
vue-best-practices
Vue 2/3 代码规范检查。包括组件命名、Props 校验、Composition API 规范等。.
python-review
Python 遗留代码审查:bare except、SQL 注入、反序列化、密钥、调试输出.
rust-review
Rust 服务审查:panic、SQL 注入、密钥、错误吞没、遗留标记.