buildkite-plugin-testing

A testing guide for Buildkite plugins, which are small extensions that add behavior to Buildkite builds. It covers Bats tests, linting, shell checking, and CI setup.

In plain words
What is it for?
Use it to write or fix Bats tests, stub commands in tests, validate plugin.yml, run shellcheck, and create the plugin’s Buildkite pipeline.
Why use it?
It catches broken hook behavior, invalid plugin configuration, shell mistakes, and other problems before release.

Skill for Claude CodeCodex

Part of the buildkite-developer-toolkit plugin — 7 skills shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/hasithaishere/buildkite-developer-toolkit/buildkite-plugin-testing
Any agent
npx skills add hasithaishere/buildkite-developer-toolkit --skill buildkite-plugin-testing
Clone the repo
git clone --depth 1 https://github.com/hasithaishere/buildkite-developer-toolkit

Made for: Claude Code, Codex.

Or install buildkite-developer-toolkit, the plugin that ships this one along with the rest of its 7 skills.

Per session 105 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,798 The whole file, excluding the scripts and references it only reads on demand.
Security scan C 2 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00105 $0.01798
Opus 5 $0.00053 $0.00899
Sonnet 5 $0.00021 $0.00360
Haiku 4.5 $0.00011 $0.00180

Measured 2d ago against content hash 721ee744b29d, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade C, and why

buildkite-plugin-testing scanned grade C with 2 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Recursive force deletehighDestructive command

rm -rf with a variable or a broad path is one typo away from removing the wrong tree.

teardown() { rm -rf "${STUB_DIR:-}"; }

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

### Stubbing external commands (docker, git, aws, curl)
skills/buildkite-plugin-testing/SKILL.md · 190 lines

How it starts

The opening of the file, as written. The whole thing — 190 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Buildkite plugin testing

Plugins are shell, so they're tested with Bats (Bash Automated Testing System) run inside the buildkite/plugin-tester image, linted with buildkite/plugin-linter, and statically checked with shellcheck.

Verify current image tags / behavior via WebSearch on buildkite.com:

The three checks every plugin needs

  1. shellcheck every hook script (catches unquoted vars, bad tests, etc.).
  2. plugin-linter validates plugin.yml + README against your plugin id.
  3. Bats exercises hook behavior, including failure paths.

Running locally

The tester image mounts your plugin at /plugin and runs bats tests/ by default:

docker run --rm -v "$PWD:/plugin:ro" buildkite/plugin-tester:latest

Or via docker-compose.yml (preferred, keeps lint + tests together):

services:
  tests:
    image: buildkite/plugin-tester:latest
    volumes: [".:/plugin:ro"]
  lint:
    image: buildkite/plugin-linter:latest
    command: ["--id", "your-org/your-plugin"]
    volumes: [".:/plugin:ro"]
docker-compose run --rm tests
docker-compose run --rm lint

Writing good Bats tests

Bats basics:

@test "description" {
  run "$PWD/hooks/command"      # captures $status and $output
  [ "$status" -eq 0 ]
  [ "$output" = "expected" ]
  [[ "$output" == *"substring"* ]]
}

Rules for plugin hook tests:

  • Set the exact env vars the hook reads (BUILDKITE_PLUGIN_<NAME>_<KEY>). Getting the names right is most of the battle — see the buildkite-plugin-dev reference/env-var-mapping.md.
  • Test the failure paths: missing required settings, invalid values. Assert both the non-zero $status and the error message.
  • Cover array settings (_0, _1) and defaults (var unset).
  • Don't call the network or real tools. Stub them (see below).
  • Use setup() / teardown() for per-test fixtures and cleanup.

Read the full file on GitHub · 190 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 190 lines · 105 tokens per session scan C 721ee744b29d

Subscribe to this mod's changes

buildkite-plugin-testing is a skill published in the GitHub repository hasithaishere/buildkite-developer-toolkit (2 stars, last pushed 1mo ago), licensed MIT. It adds 105 tokens to every session and 1,798 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it C with 2 findings (recursive force delete, makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

multi-agent-release-manager

Cleans up the workspace, formats code, runs presubmit checks, and uploads CLs to Gerrit.

chromium/chromium · 27 tokens

dsh-web-pre-push-checks

Use before pushing, opening or updating a pull request, or claiming dsh-web checks pass. Selects the required repository gates and diff-specific generation, build, and GUI evidence.

zhu1090093659/dsh-web · 45 tokens

babysit

Same-session monitoring loop for PRs, CI runs, tickets, and deployments using the monitorstart / monitorupdate / autonudgestop MCP tools. The loop re-injects your check instructions into THIS session on an idle interval — same context, same tools — and works from dashboard chat, Slack threads, and Discord DMs. Use…

kirodotdev/KiroCrew · 137 tokens

azsdk-common-pipeline-analysis

Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format. Load this skill before calling azsdkanalyzepipeline, which returns raw failure data that this skill interprets and formats. USE FOR: "pipeline failed", "build failure", "CI check failing", "tests failing in…

Azure/azure-sdk-for-net · 192 tokens

harness-setup

HAR: Project init, tool setup, agent config, memory setup, skill mirror sync. Trigger: setup, init, new project, CI/Codex setup, harness-mem, mirror. Do NOT load for: implementation, review, release, planning.

Chachamaru127/claude-code-harness · 57 tokens

managing-github-actions-secrets

Creates and updates GitHub Actions secrets for PostHog workflows. Use when adding a new CI secret, rotating an existing secret, wiring a workflow to an API token, package registry credential, deploy key, or any value referenced via ${{ secrets. }} in .github/workflows/.

PostHog/posthog-foss · 67 tokens