Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add danielgap/openclaw-planitor --skill pipeline-judgegit clone --depth 1 https://github.com/danielgap/openclaw-planitorWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/danielgap/openclaw-planitor/pipeline-judge)<a href="https://agentmods.dev/skills/danielgap/openclaw-planitor/pipeline-judge"><img src="https://agentmods.dev/badge/skills/danielgap/openclaw-planitor/pipeline-judge/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/danielgap/openclaw-planitor/pipeline-judge"><img src="https://agentmods.dev/badge/skills/danielgap/openclaw-planitor/pipeline-judge.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00029 | $0.00922 |
| Opus 5 | $0.00015 | $0.00461 |
| Sonnet 5 | $0.00006 | $0.00184 |
| Haiku 4.5 | $0.00003 | $0.00092 |
Grade A, and why
pipeline-judge scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 85 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Pipeline Judge — JudgeAgent
Input
| Archivo | Descripción |
|---|---|
plan-final.md |
Documento final |
Todos los artefactos projects/{proyecto}/v{n}/ |
Para cross-check |
research-web.json |
Datos de investigación web (OBLIGATORIO — si falta, penalizar datos_verificados -3) |
research-benchmarks.json |
Datos de Investigación web (OBLIGATORIO — si falta, penalizar validacion_cruzada -1) |
ground-truth.json |
Datos base del plan |
Output
projects/{proyecto}/v{n}/JUDGE-REPORT.json
{
"verdict": "approved|approved_with_suggestions|rejected",
"score_calidad": 0,
"score_honestidad": 0,
"score_viabilidad": 0,
"issues_criticos": [],
"warnings": [],
"sugerencias": [],
"fact_check": {
"datos_verificados": 0,
"datos_estimados": 0,
"datos_sin_fuente": 0,
"inconsistencias": []
},
"redundancia": {
"datos_repetidos": [],
"secciones_duplicadas": []
}
}
Si rejected, indicar a qué fase volver y por qué.
Frameworks
- Business Plan Validation — checklist de fact-check
- Stress testing (6 preguntas clave)
Prompt Guía
Lee el plan de negocio y todos los artefactos en projects/{proyecto}/v{n}/.
Ejecuta una review adversarial:
1. FACT-CHECK: ¿Números coinciden con Ground Truth? ¿Break-even coherente? ¿Márgenes realistas?
2. COHERENCIA INTERNA: ¿Buyer persona consistente? ¿Pricing coincide en todas las secciones?
3. ANÁLISIS FINANCIERO: ¿Sensibilidad 3 escenarios? ¿Cash flow desglosado? ¿ROI con coste oportunidad?
4. RIESGOS: ¿Cada riesgo alto tiene mitigación concreta? ¿Plan de contingencia?
5. REDUNDANCIA: ¿Se repiten datos? ¿Secciones duplicadas?
6. HONESTIDAD: ¿Optimismo infundado? ¿Datos estimados marcados? ¿Supuestos sin validar?
7. CALIDAD: ¿Legible para no técnicos? ¿Sin jerga innecesaria?
8. FUENTES DE BÚSQUEDA: ¿research-web.json y research-benchmarks.json existen? ¿Contienen datos reales (no solo errores)? ¿Los datos del plan coinciden con los del research?
9. VALIDACIÓN CRUZADA: ¿Se compararon datos de investigación web vs Investigación web? ¿Hay conflictos sin resolver? ¿Competidores de motor de búsqueda coinciden con los listados en ground-truth?
10. ANTIGÜEDAD DE DATOS: ¿Hay datos con fuente > 6 meses? Si sí, ¿están marcados ⚠️?
11. CONSULTORÍA DE EXPERTOS: ¿Se consultaron ≥3 skills de expertos? ¿Se aplicaron sus frameworks? (weight: 0.5 puntos)
12. CURVA DE INGRESOS: ¿La proyección de ingresos sigue la curva realista del sector? ¿Se investigó la curva típica del tipo de negocio? ¿Año 1 no se asume como menor que años posteriores sin justificación? ¿Se incluyen al menos 5 años de proyección? ¿Hay escenario pesimista con efecto novedad agotado? (weight: 1 punto)
Clasificar: crítico (bloquea) / warning (mejorable) / sugerencia (nice-to-have).
Scores 1-10.
Guarda como JUDGE-REPORT.json en projects/{proyecto}/v{n}/
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 85 lines · 29 tokens per session scan A 09da5353ff3b
pipeline-judge is a skill published in the GitHub repository danielgap/openclaw-planitor (5 stars, last pushed 4mo ago), licensed MIT. It adds 29 tokens to every session and 922 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
insight-error-page
Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…
next-cache-components-optimizer
Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…
next-partial-prefetching-adoption
Turn on Partial Prefetching in a Next.js app and work through the insights it surfaces. Use when the user wants to enable or adopt Partial Prefetching, flip the partialPrefetching flag, opt routes in with export const prefetch = 'partial', audit Link prefetch={true} behavior, preserve existing prefetched UI with…