field-test

field-test is a skill for Claude Code from cyanheads/open-meteo-mcp-server. It costs 93 tokens per session (9,103 once invoked), scanned B, a copy of field-test, Apache-2.0.

A live test workflow for tools, resources, and prompts exposed by an HTTP server through JSON-RPC, a structured format for requests and responses. It starts the server, calls its tools, and reports concrete findings.

In plain words
What is it for?
Use it after adding or changing server definitions to test normal and adversarial inputs, inspect the tool catalogue, verify output formats, and find transport-level bugs.
Why use it?
Unit tests use mocked surroundings, while live testing checks what a real client receives, including input handling, errors, and returned content.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin. Also seen: positional $N argument.

Part of the open-meteo-mcp-server plugin — 33 skills, 1 MCP server shipped together

Good fit Use it after adding or changing server definitions to test normal and adversarial inputs, inspect the tool catalogue, verify output formats, and find transport-level bugs.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/cyanheads/open-meteo-mcp-server/field-test
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add cyanheads/open-meteo-mcp-server --skill field-test
Clone the repo
git clone --depth 1 https://github.com/cyanheads/open-meteo-mcp-server

Made for: Claude Code.

Or install open-meteo-mcp-server, the plugin that ships this one along with the rest of its 33 skills, 1 MCP server.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for field-test

README.md
[![agentmods](https://agentmods.dev/badge/skills/cyanheads/open-meteo-mcp-server/field-test/github.svg)](https://agentmods.dev/skills/cyanheads/open-meteo-mcp-server/field-test)
Your own site
<a href="https://agentmods.dev/skills/cyanheads/open-meteo-mcp-server/field-test"><img src="https://agentmods.dev/badge/skills/cyanheads/open-meteo-mcp-server/field-test/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for field-test

Your own site · 80×15
<a href="https://agentmods.dev/skills/cyanheads/open-meteo-mcp-server/field-test"><img src="https://agentmods.dev/badge/skills/cyanheads/open-meteo-mcp-server/field-test.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 93 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 9,103 The whole file, excluding the scripts and references it only reads on demand.
Security scan B 2 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin 100% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00093 $0.09103
Opus 5 $0.00046 $0.04551
Sonnet 5 $0.00019 $0.01821
Haiku 4.5 $0.00009 $0.00910

Measured today against content hash 370664dc34d6, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade B, and why

field-test scanned grade B with 2 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Sends data to an external URLmediumData exfiltration

A POST to an outside endpoint may be telemetry or may be exfiltration; either way the mod talks to somewhere, and you should know where.

stats=$(curl -sS -o "$resp_file" -w '%{http_code} %{time_total}' -X POST "$url" "${headers[@]}" -d "$body")

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

Exercise tools, resources, and prompts against a live HTTP server via MCP JSON-RPC over curl. Starts the server, surfaces the catalog, runs real and adversarial inputs, measures every call (bytes, token estimate, wall-cl
Origin

This is a copy

100% identical to field-test — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

skills/field-test/SKILL.md · 512 lines

How it starts

The opening of the file, as written. The whole thing — 512 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Context

Unit tests (add-test skill) verify handler logic with mocked context. Field testing exercises the real HTTP transport with real JSON-RPC: starts the server, calls initialize, surfaces the catalog, runs inputs, and checks what a client actually sees. It catches what unit tests miss — awkward input shapes, unhelpful errors, missing format output, drift between structuredContent and content[], edge-case surprises.

Actively call the tools. Don't read code and guess.

Transport coverage

This skill drives an HTTP server because curl + JSON-RPC is the most reliable harness for shell-based agents. The same handlers run on both transports — only the framing differs — so HTTP exercises the full functional surface. Both HTTP session modes are covered: a durable Mcp-Session-Id session, and the sessionless initialization a MCP_SESSION_MODE=stateless server performs.

Stdio coverage is a boot check only — run this before Step 1. Run bun run rebuild && bun run start:stdio, confirm the startup logs look clean (banner, expected tool/resource counts, no errors/warnings, no missing-config gripes), then kill it. Pino logs go to stderr in stdio mode (stdout is reserved for JSON-RPC), so they print straight to the terminal when you run interactively. No need to call tools over stdio — the HTTP pass already covered handler behavior.


Steps

1. Start the server

Generate a 10-character alphanumeric ID (e.g. 9DJ73-K103L) and write the helper to /tmp/<project-name>-field-test-<ID>.sh. Use that exact path in every subsequent Bash call. Two agents in the same project tree must pick different IDs — that's what keeps their helper files, server logs, and call scratch from colliding.

The helper itself is stateless — every function takes the IDs it needs (server pid, url, port, MCP sid, server log path) as positional args. mcp_start prints them; the agent threads them through every later call. No env vars, no shared state files.

Read the full file on GitHub · 512 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. today Changed · +119 lines · +17 tokens per session scan A → B 370664dc34d6
  2. 10d ago First seen · 393 lines · 76 tokens per session scan A 5e0878fedc2a

Subscribe to this mod's changes

field-test is a skill published in the GitHub repository cyanheads/open-meteo-mcp-server (5 stars, last pushed today), licensed Apache-2.0. It adds 93 tokens to every session and 9,103 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it B with 2 findings (sends data to an external url, makes network calls). It is 100% identical to field-test, differing in 0 lines, and is treated as a copy.

Related

Other skills, from other repositories

api-testing

Testing patterns for MCP tool/resource handlers using createMockContext and Vitest. Covers mock context options, handler testing, McpError assertions, format testing, Vitest config setup, and test isolation conventions.

cyanheads/eur-lex-mcp-server · 46 tokens

add-test

Scaffold a test file for an existing tool, resource, or service. Use when the user asks to add tests, improve coverage, or when a definition exists without a matching test file.

cyanheads/eur-lex-mcp-server · 41 tokens

api-testing

Testing patterns for MCP tool/resource handlers using createMockContext and Vitest. Covers mock context options, handler testing, McpError assertions, format testing, Vitest config setup, and test isolation conventions.

cyanheads/gdelt-mcp-server · 46 tokens

field-test

Exercise tools, resources, and prompts against a live HTTP server via MCP JSON-RPC over curl. Starts the server, surfaces the catalog, runs real and adversarial inputs, and produces a tight report with concrete findings and numbered follow-up options. Use after adding or modifying definitions, or when the user asks to…

cyanheads/gdelt-mcp-server · 76 tokens

add-test

Scaffold a test file for an existing tool, resource, or service. Use when the user asks to add tests, improve coverage, or when a definition exists without a matching test file.

cyanheads/gdelt-mcp-server · 41 tokens

api-testing

Testing patterns for MCP tool/resource handlers using createMockContext and Vitest. Covers mock context options, handler testing, McpError assertions, format testing, Vitest config setup, and test isolation conventions.

cyanheads/aviation-weather-mcp-server · 46 tokens