eval-engineer
529Galileo-Agent-Labs/eval-engineer
Skill Codex
Use when a user is unsure which Eval Engineer command to run for AI agents/RAG apps, needs onboarding/status for a .galileo workspace, or asks where to start.
16,632 tagged codex, measured the same way as everything else here.
Browse within: claudecode 297agentic-ai 232coding-agents 200claudecode-skills 192skills 182ai-coding 176agent 166copilot 143ai-tools 141openai 138awesome 136awesome-list 136codex-plugins 136cli 97
Galileo-Agent-Labs/eval-engineer
Skill Codex
Use when a user is unsure which Eval Engineer command to run for AI agents/RAG apps, needs onboarding/status for a .galileo workspace, or asks where to start.
IAPro-Community/Orquestrador-Maestro
Skill Claude Code
Socratic deep interview with mathematical ambiguity gating before execution.
Skill Claude CodeCodex needs its repo
Use for the H100 Triton fused RMSNorm + SiLU gate optimization task. Provides the workflow and MCP tooling for objective runs, rollback-safe config sweeps, source diagnostics, benchmark history, and final held-out verification.
Skill Claude CodeCodex
The First-Time Principle at four layers: thought, planning, execution, output (including conversation). Write everything as if for the first time. Triggers on implementation, refactor, bug fix, doc/memory edit, planning, and any "just patch it" urge. Cross-platform: Claude Code, OpenAI Codex, OpenCode, OpenClaw…
BusyBee3333/sol-governed-codex
Skill Codex
Use when Codex work should route bounded implementation to lower-cost workers and reserve Sol for risky plan review or final evidence gates, or when installing or validating this workflow.
eddyzzl/imagegen-pptx-pipeline
Skill Codex
End-to-end stateful workflow for designing and generating editable PowerPoint decks from briefs, templates, reference decks, brand assets, data, ImageGen comps, or user-supplied slide images. Uses multi-round content/style confirmation, reviewed per-slide comps, mandatory Real-ESRGAN 4K comp/icon upscaling, built-in…
Skill Claude Code
Use when solving CUMCM, MCM, ICM, or other mathematical modeling tasks that need contest problem intake, model selection, Python solving, figures, paper writing, quality audit, and packaged deliverables.
Skill Claude CodeCodex
A writing-style guide for producing natural-sounding Chinese conversations and creative fiction. It defines rules for avoiding artificial wording, unsupported claims, and unwanted changes to supplied text.
Skill Claude CodeCodex
Product analytics with your AI agent: set up consent-based tracking, read funnels, paths, retention, experiments, and context, then recommend the smallest growth action using the official Agent Analytics CLI.
Skill Claude CodeCodex
Create top-level Wolfpack sessions, spawn child agents, and inspect or control local/Tailscale sessions through the canonical CLI.
Skill Claude Code needs its repo
A reusable client foundation for connecting to Lingxing ERP’s OpenAPI, including authentication, request signing, documentation caching, and basic connection checks.
himynameisben/macos-disk-cleanup
Skill Claude CodeCodex
A macOS disk-space diagnostic and cleanup workflow, including hidden “System Data” storage. It examines developer files, package-manager caches, containers, virtual machines, browser profiles, and leftover app data.
Skill Claude CodeCodex
Skill "openclaw-plugin" from christopherkarani/ryk, covering ryk, when to use ryk, how it works, prerequisites and example policy behavior.
Skill Claude Code
Event schemas for ml-ralph log.jsonl. Use when logging any event.
Skill Codex
Run a strict repository-native Spec programming loop that keeps requirements, owning decisions, implementation, tests, and current documentation consistent in one bounded change. Use when explicitly invoked; when repository instructions require a Spec, RFC, ADR, proposal, design doc, or decision record for…
Skill Claude Code
Search the user's own past AI coding-agent conversations (across Claude Code, Codex, Cursor, Gemini, Qwen, Goose, OpenCode, Continue, Cline, and in-app chats) before redoing work. Use when the user references something they "did before", "talked about", "figured out earlier", asks "how did we…/what did we decide…
Skill Claude Code
Use when someone asks for an AIOS audit, asks to score their setup against the Four Cs, or says "is my AIOS working" / "audit my setup" / "find gaps in my AIOS". Produces a Four-Cs scoreboard with top-3 fixes ranked by leverage.
Skill Codex
Coordinate Planr-backed or project-configured Herdr workflows in which Sol owns intake, architecture, design-sensitive implementation, integration, and the final decision; Luna Max serves only as an explicitly selected read-only silent sentinel, fresh bounded verifier, or interactive operator; and Fable supplies at…
Skill Claude CodeCodex
5-phase repeatable structure for autonomous agent sessions: context-load, tiered work-selection, coordination claim, execute, and persist-learning. Prevents duplicate work across concurrent sessions and ensures every run produces durable artifacts. Runtime-agnostic — Claude Code, gptme, Codex, or any…
Skill Codex
Coordinates custom specialists through workflow phases, artifact handoffs, approval gates, and autonomous implementation. Use when executing a workflow recipe or routing its next specialist.
Skill Codex
Local open-source Agent Skill for authorized web application, API, browser, business-logic, and supporting server-integrity penetration testing. Use for scoped OWASP WSTG, ASVS, and API Security assessments; authenticated and multi-role validation; attack-surface mapping; bounded DAST; evidence-backed findings…
Skill Claude CodeCodex
Drive the cultivar CLI to test whether an agent skill improves behavior — scaffold tasks, run with/without the skill across Claude/Copilot/Gemini (locally or on Modal), grade against a rubric, and read the results.
paradoxie/cpa-codex-auth-sweep
Skill Claude CodeCodex
A tool for scanning local Codex authentication files and identifying credentials that no longer work. It can also remove credentials confirmed as invalid.
Skill Claude Code
Triage memory, swap, disk and CPU where you are running, and rank Docker reclaim levers by safety×yield. Docker triage is host-only: inside a sandbox with no Docker socket the skill states that once and emits a procedure for the orchestrator to run at the host project root, rather than failing per command. TRIGGER…
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: