stackhawk-data-seed

A workflow for preparing checked-in test data so authenticated HawkScan security scans can reach meaningful application paths. HawkScan is a security scanner, and seed data is sample data placed in an application for testing.

In plain words
What is it for?
Running the seed preflight, creating a manifest of required sample data, validating it, and finalizing reproducible files under data-seed/. It does not run the application or configure HawkScan itself.
Why use it?
A scan may miss authenticated routes when the application has no suitable records to work with. This workflow designs and validates the smallest seed-data set based on the repository's detected storage and services.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/stackhawk/agent-skills/stackhawk-data-seed
Any agent
npx skills add stackhawk/agent-skills --skill stackhawk-data-seed
Clone the repo
git clone --depth 1 https://github.com/stackhawk/agent-skills

Made for: Claude Code, Codex.

Per session 159 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,427 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00159 $0.02427
Opus 5 $0.00079 $0.01213
Sonnet 5 $0.00032 $0.00485
Haiku 4.5 $0.00016 $0.00243

Measured 2d ago against content hash e782413abcd8, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

stackhawk-data-seed scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/stackhawk-data-seed/skills/stackhawk-data-seed/SKILL.md · 189 lines

How it starts

The opening of the file, as written. The whole thing — 189 lines — stays where its author put it; the contents beside it link to each section on GitHub.

StackHawk Data Seed Skill

This skill produces checked-in, reproducible seed-data artifacts for a target repo so authenticated HawkScan finds non-empty results. The hawk perch seed command provides the deterministic steps — a static repo pre-flight (storage + upstream detection), a manifest validator, and an artifact finalizer. This skill supplies the reasoning between them: it reads the pre-flight's digest and designs the minimum seed manifest. It works the same across every agent that can run a subprocess and read its output.

It does NOT run the artifacts, start the environment, or write stackhawk.yml — those belong to the human, the user's tooling, and the hawkscan skill respectively.


When to Run

Invoke explicitly when:

  • User says "set up data for HawkScan" / "seed this repo" / "my scan has no data to hit."
  • Configuring HawkScan against a repo for the first time and the app needs authenticated routes to scan.
  • A previous data-seed/ exists but the data shape changed (new entity types, new upstream service).

Do NOT run autonomously after code changes — this is a setup tool, not a per-commit safety net.


Phase 0: Preflight

0.1 — Confirm working directory

The user must invoke from inside the target repo (the repo HawkScan will scan):

test -d .git || echo "NOT-A-REPO"
pwd

If not a git repo, ask the user to cd to the target repo and re-invoke.

0.2 — Confirm hawk supports the seed flow (capability gate)

This skill drives the caller-driven hawk perch seed subcommands (validate and finalize). Probe for them directly:

# Identify the driving skill for CLI usage telemetry (read by hawk/hawkop).
export _STACKHAWK_SKILL=stackhawk-data-seed
if hawk perch seed validate --help >/dev/null 2>&1 && hawk perch seed finalize --help >/dev/null 2>&1; then
  echo "SEED-FLOW-OK"
else
  echo "SEED-FLOW-UNSUPPORTED"
fi
hawk version 2>/dev/null || hawk --version 2>/dev/null   # for the message only; never gates

Read the full file on GitHub · 189 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 189 lines · 159 tokens per session scan A e782413abcd8

Subscribe to this mod's changes

stackhawk-data-seed is a skill published in the GitHub repository stackhawk/agent-skills (15 stars, last pushed 12d ago), licensed MIT. It adds 159 tokens to every session and 2,427 once invoked, about $0.0008 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

next-cache-components-adoption

Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…

vercel/next.js · 95 tokens

babysit-pr

Babysit a GitHub pull request after creation by continuously polling review comments, CI checks/workflow runs, and mergeability state until the PR is merged/closed or user help is required. Diagnose failures, retry likely flaky failures up to 3 times, auto-fix/push branch-related issues when appropriate, and keep…

openai/codex · 114 tokens

imagegen

Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output…

openai/codex · 113 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

next-cache-components-optimizer

Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…

vercel/next.js · 170 tokens