Babysitter is a workflow engine for AI coding agents that enforces predefined steps, quality checks, human approvals, and decision records. It is used to coordinate complex, repeatable agent workflows across supported coding tools. The catalogue contains skills, agents, instructions, settings, a plugin, and an MCP integration for its workflow.
Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add a5c-ai/babysitter --skill gsd-toolsgit clone --depth 1 https://github.com/a5c-ai/babysitterWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/a5c-ai/babysitter/gsd-tools)<a href="https://agentmods.dev/skills/a5c-ai/babysitter/gsd-tools"><img src="https://agentmods.dev/badge/skills/a5c-ai/babysitter/gsd-tools.svg" alt="Measured on agentmods" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to high
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- high Memory Poisoning · line 234 Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.Fix: Protect agent memory and state from modification by untrusted content. Use read-only memory for critical instructions and validate all state changes.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00062 | $0.01961 |
| Opus 5 | $0.00031 | $0.00981 |
| Sonnet 5 | $0.00012 | $0.00392 |
| Haiku 4.5 | $0.00006 | $0.00196 |
Grade A, and why
gsd-tools scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 244 lines — stays where its author put it; the contents beside it link to each section on GitHub.
- Load and manage
.planning/config.jsonconfiguration - Generate deterministic slugs from descriptions (kebab-case, deduplication)
- Format timestamps for GSD artifact headers and frontmatter
- Resolve paths within the
.planning/directory structure - Parse and normalize phase numbers (integer and decimal)
- Manage sequential quick task numbering in
.planning/quick/ - Initialize and validate
.planning/directory structure - Detect project state (new project, existing project, mid-phase, etc.)
Capabilities
1. Configuration Loading
Load and parse .planning/config.json with defaults:
{
"profile": "balanced",
"autoCommit": true,
"autoVerify": true,
"planCheckerEnabled": true,
"maxPlanRevisions": 2,
"waveParallelization": true,
"contextMonitor": true,
"contextWarningThreshold": 70,
"contextCriticalThreshold": 85,
"summaryVariant": "standard",
"tddEnabled": false,
"questioningDepth": "adaptive"
}
Read config:
cat .planning/config.json
If config does not exist, create with defaults. Merge user overrides with defaults (user values take precedence).
2. Slug Generation
Generate kebab-case slugs from descriptions:
Input: "Add user authentication with OAuth2"
Output: "add-user-authentication-with-oauth2"
Input: "Fix bug #123 in payment processing"
Output: "fix-bug-123-in-payment-processing"
Rules:
- Lowercase all characters
- Replace spaces and special characters with hyphens
- Remove consecutive hyphens
- Trim leading/trailing hyphens
- Maximum 60 characters (truncate at word boundary)
- Check for duplicates in target directory, append
-2,-3if needed
3. Timestamp Formatting
Format timestamps for GSD artifacts:
Header: "2026-03-02T14:30:00Z"
Frontmatter: "2026-03-02"
Filename: "20260302-143000"
Display: "Mar 2, 2026 2:30 PM"
4. Path Operations
Resolve paths within .planning/ structure:
Phase directory: .planning/phase-{N}/
Plan file: .planning/phase-{N}/PLAN-{M}.md
Summary file: .planning/phase-{N}/SUMMARY.md
Quick task: .planning/quick/{NNN}-{slug}/
Debug session: .planning/debug/{slug}.md
Codebase docs: .planning/codebase/
Research docs: .planning/research/
Milestone archive: milestones/v{X}.{Y}/
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 244 lines · 62 tokens per session scan A 5dd998f56e77
gsd-tools is a skill published in the GitHub repository a5c-ai/babysitter (1,778 stars, last pushed 2d ago), licensed MIT. It adds 62 tokens to every session and 1,961 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
kano-backlog
Prioritize and refine a GitHub Issues backlog with the Kano model — categorize every open issue as Must-be, Performance, Attractive, Indifferent, or Reverse, apply Kano + priority labels back to GitHub automatically, and recommend the single best next issue to pick up. Use this whenever the user wants to triage…
team-structure
Breaks a reviewed design into verified slices. Trigger on "slice this up", "break the design into steps", or "/team-structure".
tracking-tickets
Defines tracker status transitions and closing rules. Load when a pipeline run is linked to a ticket.
team-pr
Opens a pull request after verification. Trigger on "open the PR", "open a draft PR", or "/team-pr" only; never infer the phase from passed verification.
pr-watch-mechanics
Bounded watch-loop mechanics for the pr-watch skills: cycle timing, soft cap, handoff. Load when running or authoring a PR watch loop.
team-question
Decomposes a feature into task and question artifacts. Trigger on "shape this idea", "decompose this task", or "/team-question".