Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/babywyrm/mcpnukeWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/rules/babywyrm/mcpnuke/mcpnuke-tests)<a href="https://agentmods.dev/rules/babywyrm/mcpnuke/mcpnuke-tests"><img src="https://agentmods.dev/badge/rules/babywyrm/mcpnuke/mcpnuke-tests.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.00499 |
| Opus 5 | $0.00000 | $0.00249 |
| Sonnet 5 | $0.00000 | $0.00100 |
| Haiku 4.5 | $0.00000 | $0.00050 |
Grade A, and why
mcpnuke-tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
mcpnuke Test Conventions
TDD Workflow (superpowers: test-driven-development)
- RED — Write the failing test first. Run it. Watch it fail.
- GREEN — Write the minimal code to make it pass. Run it. Watch it pass.
- REFACTOR — Clean up. Run tests again. Still green.
Never write implementation before the test exists.
Test Structure
"""Tests for <check_name>."""
from mcpnuke.core.models import TargetResult
from mcpnuke.checks.<module> import check_<name>
def test_<name>_positive(result_with_tools):
r = result_with_tools([{"name": "vuln_tool", "description": "...", "inputSchema": {}}])
check_<name>(r)
assert any(f.check == "<name>" for f in r.findings)
def test_<name>_clean(result_with_tools):
r = result_with_tools([{"name": "safe_tool", "description": "Nothing here", "inputSchema": {}}])
check_<name>(r)
assert not any(f.check == "<name>" for f in r.findings)
def test_<name>_timing(result_with_tools):
r = result_with_tools([{"name": "x", "description": "y", "inputSchema": {}}])
check_<name>(r)
assert "<name>" in r.timings
Running Tests
uv run pytest tests/ -v # Full suite
uv run pytest tests/test_foo.py -v # Single file
uv run pytest tests/ -v -x # Stop on first failure
Verification (superpowers: verification-before-completion)
Before claiming work is done, run and confirm:
uv run pytest tests/ -v— 1040+ passed, 37 skipped, 0 faileduv run ruff check .— zero findings, anduv run mypy mcpnuke/at or below the ceiling in.github/workflows/tests.yml- No import errors or deprecation warnings
- New tests cover positive, negative, and timing
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 59 lines · 0 tokens per session scan A 514445fa84a5
mcpnuke-tests is a cursor rule published in the GitHub repository babywyrm/mcpnuke (2 stars, last pushed 2d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 499 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other cursor rules, from other repositories
workflow
BDD delivery order, Gherkin spec before tests before code, definition of done, Rules Sync.
debug-issue
When the user reports a bug, an error, or unexpected behavior. Enforces four structured phases — reproduction, failing test, root cause isolation, fix and verify — to stop guess-and-check loops.
common-testing
Testing requirements: 80% coverage, TDD workflow.
tdd
Test-driven development — red-green-refactor cycle.
testing_rules
A testing guide covering TDD, where tests are written before code, and BDD, where expected behavior is described in Given-When-Then form.
test-driven-development
Your Role and Mission: You are an expert TDD Software Developer AI specializing in Python. Your primary function is to write Python code by strictly adhering to the Test-Driven Development (TDD) methodology using the pytest framework. Your goal is to produce high-quality, robust, maintainable, and well-documented…