Guide the user through an independent iFixAi audit of their own agent, checking whether it does the job it is supposed to do given their business rules and org structure. Prefer pointing it at the user's REAL deployed agent over its HTTP endpoint (its actual tools, retrieval, and governance) with --provider http…
Give your AI agent a professional health checkup (AgentVitals). Use when the user asks the agent to run a checkup / test itself / benchmark itself ("run a checkup", "check your vitals", "test yourself", "how stable are you", "/checkup"), or an advanced personality checkup (backbone, proactivity, creativity). 给 AI…
Use when a user is unsure which Eval Engineer command to run for AI agents/RAG apps, needs onboarding/status for a .galileo workspace, or asks where to start.
Use when a user asks to fetch Galileo evidence, provides Galileo URLs or IDs, says "fetch this Galileo link", or needs traces, sessions, experiments, or log streams saved locally.
Use when the user asks to set up Eval Engineer, check .galileo readiness, create workspace scaffolding, or configure editable files, verification commands, app type, or evidence paths.
Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR…
Run coding agent benchmarks and verify results with nasde. Use this skill when the user wants to: Run a benchmark (all tasks, single task, specific variant) Re-run assessment evaluation on existing trial results Check or verify results in Opik (traces, feedback scores, experiments) Troubleshoot a failed benchmark run…
Scope what quality should be measured, convert it into one or more actionable binary evaluations, deploy those evaluations through Truesight MCP, and generate a companion skill that applies them correctly. Use when a user wants to create new evals, quality checks, guardrails, or pass/fail criteria for AI outputs.
Systematically identify and categorize failure modes in evaluated traces using Truesight datasets and error-analysis tools. Use when quality issues are unclear, after major pipeline changes, or when incidents indicate drift.
Generate synthetic test data for LLM evaluations using dimension-based tuple expansion. Use when the user needs synthetic traces, test cases, eval datasets, or when create-evaluation needs synthetic fallback data.
Operates Harneloop harness units through adaptive intake, environment mapping, artifact-aware attempts, explicit evaluation decisions, candidates, and evidence-gated promotion. Use when creating, continuing, evaluating, or recovering context for a Harneloop harness unit.
Use this skill to evaluate the quality of a RAG pipeline on faithfulness, answer relevancy, context precision, context recall, and hallucination rate. Activates after a RAG system is implemented or when retrieval quality is in question. Produces a structured evaluation report with measurable results.
Use when user has multiple analysis outputs (risk, feasibility, architecture, verification) and needs a combined decision. Triggers: "combine these analyses", "decision brief", "GO/NO-GO", "aggregate results", "what's the overall verdict".
Use when user wants to find or create a business opportunity — whether discovering market gaps, inventing novel concepts, or both. Combines signal harvesting, gap detection, novel synthesis, and revenue architecture into one pipeline. Triggers: "find a niche", "what should I build", "business opportunity", "invent…
Use when user asks to verify, fact-check, validate, or check correctness of code, documents, specs, claims, or LLM-generated artifacts. Triggers: "verify this", "is this correct", "check against spec", "find contradictions", "review this document". Do NOT use for code review (style/quality) — this is for…
Use when the user wants a WCAG 2.2 accessibility review of a UI — semantics, keyboard operability, focus management, color contrast, ARIA, forms/labels, reduced-motion, and screen-reader support, including streaming AI output via live regions. Triggers on "accessibility audit", "is this WCAG compliant", "a11y review"…
Use when the user wants to reduce LLM token usage, context-window pressure, or API cost in an AI/agent codebase without hurting quality — reviewing prompt construction, context assembly, chat-history retention, tool definitions, retrieval, caching, batching, and output verbosity. Triggers on "reduce token usage", "cut…
Use when the user wants a UX / UI / interaction-design review or redesign of an app, dashboard, editor, canvas, AI/agentic product, web app, or mobile app — including microinteractions, motion, loading, error recovery, empty states, accessibility, perceived performance, and AI trust/observability. Triggers on "audit…
Mide qué fracción de las reglas propias del proyecto se cumple de verdad, contándolo sobre el historial de sesiones que Claude Code ya guarda en disco. Se activa al preguntar si una norma, protocolo o convención se está aplicando, al revisar si un CLAUDE.md sirve de algo, al decidir si una regla debe pasar a ser hook…
Connect the current Codex task to a consented, redacted public experiment journal on joinmultiplayer.ai. Use when the user asks to start, continue, inspect, finish, or publicly document a Pocket i / joinmultiplayer experiment without leaving Codex. Do not use for ordinary private coding tasks or publish anything…
Use this skill whenever a user or autonomous agent needs scientifically grounded, machine-readable knowledge about natural systems, ecological relationships, geometry in nature, established mathematical references including Hopf Fibration, topology, fiber bundles, state-space geometry, weather, water systems, ocean…
A minimal, deliberately clean SKILL.md fixture used to CI-test scripts/lint-skills.mjs against a known-good file. It has valid frontmatter, no dollar-digit sequences, no secret-shaped strings, and no personal paths.