Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add cckyros/goal-acceptance --skill validate-criteriongit clone --depth 1 https://github.com/cckyros/goal-acceptanceWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/cckyros/goal-acceptance/validate-criterion)<a href="https://agentmods.dev/skills/cckyros/goal-acceptance/validate-criterion"><img src="https://agentmods.dev/badge/skills/cckyros/goal-acceptance/validate-criterion.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00026 | $0.00506 |
| Opus 5 | $0.00013 | $0.00253 |
| Sonnet 5 | $0.00005 | $0.00101 |
| Haiku 4.5 | $0.00003 | $0.00051 |
Grade A, and why
validate-criterion scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
"evidence": "curl -s localhost:3000/health → 200 {\"status\":\"ok\"}", What it actually says
After executing work, call this to validate each acceptance criterion.
Statuses passed and failed require concrete evidence from actual
execution.
💡 Fast-path: Use the run_and_validate tool to automatically execute a
shell command and record stdout/stderr/exitCode as high-confidence evidence in
a single call.
Evidence Requirements
You MUST run the actual command or check before validating. Do NOT validate based on memory, assumption, or "it should work".
| method | What to do | Evidence to provide |
|---|---|---|
command |
Run the exact command in a shell | Paste the real stdout/stderr + exit code |
file |
Read the file and check the content | Paste the relevant lines from the file |
url |
Make the HTTP request | Paste the response status + body |
What NOT to do
- ❌ Validate
passedwithout running anything - ❌ Write "should work" or "looks correct" as evidence
- ❌ Copy evidence from a previous run without re-running
- ❌ Use
evidence_type=textfor a criterion withmethod=command - ❌ Validate before implementation is actually complete
Self-Claimed Status
When role=agent (default), passed criteria are marked selfClaimed=true.
This means can_complete_goal will block until an independent reviewer
calls confirm_criterion with fresh evidence.
This is by design — your self-assessment is not trusted for completion.
Use confirm_criterion (as a separate reviewer agent) to convert
self-claimed passes to formal passes.
Example
{
"criterion_id": "api-200",
"status": "passed",
"evidence": "curl -s localhost:3000/health → 200 {\"status\":\"ok\"}",
"evidence_type": "command"
}
{
"criterion_id": "api-500",
"status": "failed",
"evidence": "curl -s localhost:3000/users → 500 Internal Server Error: TypeError: Cannot read property 'name' of undefined",
"evidence_type": "command"
}
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 62 lines · 26 tokens per session scan A f2873259fcb5
validate-criterion is a skill published in the GitHub repository cckyros/goal-acceptance (3 stars, last pushed 10d ago), licensed MIT. It adds 26 tokens to every session and 506 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
daytona-windows-cert
Run the verified Windows repro for iPolloWork enterprise TLS behavior: install a fake corporate CA into the Windows machine store, serve healthy and broken HTTPS control planes, install a Windows build, and prove the desktop app and spawned runtimes use the operating system trust path.
daytona-recording-artifacts
Use this skill to collect proof that a Daytona UI flow works. Use fraimz before declaring the flow passed. If the user asks to do e2e tests for a feature, frame proof is required unless they explicitly ask for non-UI or mock-only validation.
fraimz
create a fraimz, make fraimz, prove it works, frame proof, PR proof, validate experience, e2e evidence, fraimz.html. The full fraimz loop — frame the claim, drive the real app via CDP, validate/repair, output fraimz.html. Use whenever a task ends with "please create a fraimz" or any change needs end-to-end proof.
daytona-electron-den
Electron and Den, desktop plus cloud, two-sandbox e2e, cloud auth, marketplace, org policy, worker proxy, provider sync, desktop handoff. Validate Electron against a Daytona Den server with unified proof.
run-evals
Launch a real iPolloWork app and run coded eval flows against it. This skill owns launch + run; the prove/repair/verdict loop and evidence standard live in the fraimz skill — load that too for anything that ends in a verdict.
code2skill-review-flow
A read-only reviewer for the main user flows in a package generated by Code2Skill. It checks whether representative paths are basically usable, without proving that every source-code detail is included.