Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/managedcode/geminisharpsdk/mcaf-testingnpx skills add managedcode/GeminiSharpSDK --skill mcaf-testinggit clone --depth 1 https://github.com/managedcode/GeminiSharpSDKWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00075 | $0.00997 |
| Opus 5 | $0.00037 | $0.00498 |
| Sonnet 5 | $0.00015 | $0.00199 |
| Haiku 4.5 | $0.00007 | $0.00100 |
Grade A, and why
mcaf-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 68 lines — stays where its author put it; the contents beside it link to each section on GitHub.
MCAF: Testing
Outputs
- New/updated automated tests that encode documented behaviour (happy path + negative + edge), with integration/API/UI preferred
- For new behaviour and bugfixes: tests drive the change (TDD: reproduce/specify → test fails → implement → test passes)
- Updated verification sections in relevant docs (
docs/Features/*,docs/ADR/*) when needed (tests + commands must match reality) - Evidence of verification: commands run (
build/test/coverage/analyze) + result + the report/artifact path written by the tool (when applicable)
Workflow
- Read
AGENTS.md:- commands:
build,test,format,analyze, and the repo’s coverage path (either a dedicatedcoveragecommand or atestcommand that generates coverage) - testing rules (levels, mocks policy, suites to run, containers, etc.)
- commands:
- Start from the docs that define behaviour (no guessing):
docs/Features/*for user/system flows and business rulesdocs/ADR/*for architectural decisions and invariants that must remain true- if the docs are missing/contradict, fix the docs first (or write a minimal spec + test plan in the task/PR)
- follow
AGENTS.mdscoping rules (Architecture map → relevant docs → relevant module code; avoid repo-wide scanning)
- Follow
AGENTS.mdverification timing (optimize time + tokens):- run tests/coverage only when you have a reason (changed code/tests, bug reproduction, baseline confirmation)
- start with the smallest scope (new/changed tests), then expand to required suites
- Define the scenarios you must prove (map them back to docs):
- positive (happy path)
- negative (validation/forbidden/unauthorized/error paths)
- edge (limits, concurrency, retries/idempotency, time-sensitive behaviour)
- for ADRs: test the invariants and the “must not happen” behaviours the decision relies on
- Choose the highest meaningful test level:
- prefer integration/API/UI when the behaviour crosses boundaries
- use unit tests only when logic is isolated and higher-level coverage is impractical
- Implement via a TDD loop (per scenario):
- write the test first and make sure it fails for the right reason
- implement the minimum change to make it pass
- refactor safely (keep tests green)
- Write tests that assert outcomes (not “it runs”):
- assert returned values/responses
- assert DB state / emitted events / observable side effects
- include negative and edge cases when relevant
- Keep tests stable (treat flakiness as a bug):
- deterministic data/fixtures, no hidden dependencies
- avoid
sleep-based timing; prefer “wait until condition”/polling with a timeout - keep test setup/teardown reliable (reset state between tests)
- Coverage (follow
AGENTS.md, optimize time/tokens):- run coverage only if it’s part of the repo’s required verification path or if you need it to find gaps
- run coverage once per change (it is heavier than tests)
- capture where the report/artifacts were written (path, summary) if generated
- If the repo has UI:
- run UI/E2E tests
- inspect screenshots/videos/traces produced by the runner for failures and obvious UI regressions
- Run verification in layers (as required by
AGENTS.md):
- new/changed tests first
- then the related suite
- then broader regressions if required
- run
analyzeif required
- Keep docs and skills consistent:
- ensure
docs/Features/*anddocs/ADR/*verification sections point to the real tests and real commands - if you change test/coverage commands or rules, update
AGENTS.mdand this skill in the same PR
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 68 lines · 75 tokens per session scan A aa5400032d50
mcaf-testing is a skill published in the GitHub repository managedcode/GeminiSharpSDK (2 stars, last pushed 4mo ago), licensed MIT. It adds 75 tokens to every session and 997 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
security-patcher
Invoke this as your absolute first action before using any other tools whenever a user requests to fix, patch, or remediate a vulnerability. Do not perform manual research first.
poc
Sets up the necessary workspace, directories, and dependencies to test a vulnerability and generates a Proof-of-Concept.
dependency-manager
Safely resolve and install isolated dependencies for isolated sandboxes (PoC execution).
architecture
Project architecture and file structure conventions for all process types. Use when: (1) Creating new files or modules, (2) Deciding where code should go, (3) Converting single-file components to directories, (4) Reviewing code for structure compliance, (5) Adding new bridges, services, agents, or workers.
testing
Testing workflow and quality standards for writing and running tests. Use when: (1) Writing new tests, (2) Adding a new feature that needs tests, (3) Modifying logic that has existing tests, (4) Before claiming a task is complete.
widget-creator
Step-by-step guide for creating new widgets in the Hex1b TUI library. Use when implementing new widgets from scratch, including widget records, nodes, extension methods, theming, reconciliation, and tests.