AI Safety skills

369 tagged AI Safety, measured the same way as everything else here.

Browse within: ai-security 267ai-behavior-analysis 149ai-safety-research 149ai-skill-safety 148ai-governance 139agent-governance 125policy-as-code 107credentials 106Guardrails 77claude-code-plugin 68ai-tools 64ai-coding 63compliance 59presentation 56

skill-scout

25

Brain-ai-biz/skill-scout

Skill Claude Code

Discover the best Claude Code extensions on the web AND tell the user how safe each one is before installing. Reads the user's Claude Code setup (project context, memory, history), asks a few questions to understand their goal (general or a specific idea), then surfaces a ranked shortlist - name, what it does, what…

not rated 7 2mo ago A 193 tokens original MIT

clarify-first

26

DmiyDing/clarify-first

Skill Claude CodeCodex

This skill should be used when a request is ambiguous, underspecified, conflicting, or high impact. It is intended for vague verbs like optimize, improve, fix, refactor, and add feature; for missing file paths or unknown dependencies; and for risky actions like deploy, delete, overwrite, or migrate. It should not be…

not rated 6 6mo ago A 108 tokens original Apache-2.0

api-scam-hunter

27

astrozeta/api-scam-hunter

Skill Claude CodeCodex

Detects whether a purchased AI API key/endpoint (Claude/Anthropic, OpenAI, etc.) — often bought cheap from a reseller or marketplace like GamsGo — is actually a man-in-the-middle proxy that intercepts, rewrites or degrades traffic. Trigger when the user says things like "I bought a third-party API key", "is this…

not rated 5 2mo ago A 167 tokens original MIT

agent-browser

28

vmehera123/leashd

Skill Claude Code

Core agent-browser usage guide. Read this before running any agent-browser commands. Covers the snapshot-and-ref workflow, navigating pages, interacting with elements (click, fill, type, select), extracting text and data, taking screenshots, managing tabs, handling forms and auth, waiting for content, running multiple…

not rated 5 3d ago A 113 tokens copy · 86% Apache-2.0

research-guardian

29

htlin222/research-guardian-skill

Skill Claude CodeCodex

A research-quality checking system for AI agents. It verifies facts, compares literature, checks experiment designs, and reviews claims about new findings.

not rated 5 4mo ago A 222 tokens

levelsofself/mcp-nervous-system

Skill Claude CodeCodex

Use this skill when building, managing, or auditing multi-agent AI systems. Provides governance patterns for behavioral enforcement, drift detection, audit trails, role management, and accountability across autonomous AI agents. Compatible with any orchestration framework.

not rated 4 1mo ago A 51 tokens original MIT

dayan-deck

31

Kosmoray/dayan-agent-skills

Skill Claude CodeCodex

A method for making an editable HTML slide presentation from an outline, notes, document, or verified facts. Each slide has a clear job, and the deck is designed for an audience and a set presentation length.

not rated 4 24d ago A 87 tokens original MIT

sis-skill

32

Architect-SIS/sis-skill

Skill Claude CodeCodex

Equilibrium-Native Reasoning for OpenClaw.

not rated 4 6mo ago A 0 tokens original MIT

agentguard

33

bmdhodl/agent47

Skill Claude CodeCodex

Runtime guardrails for AI coding agents. Stop loops, budget overruns, retry storms, and timeouts before they burn money. Zero dependencies, local-first, MIT licensed.

not rated 4 2d ago A 38 tokens original MIT

dagx-agi-kernel

34

dankofly/perfectify

Skill Claude CodeCodex

Improve and verify agent work after repeated failures, in dependency-heavy tasks, or when optimization claims need baseline and regression evidence. Use for DAGx/Perfectify requests, failed retries, risky multi-step work, or requests to verify an improvement. Exclude routine questions, drafting, one-step edits, and…

not rated 4 17d ago A 72 tokens original MIT

clawheart-security

35

tjsdyy/clawheartv2

Skill Claude CodeCodex

A skill for using the local ClawHeart command-line tool to audit AI security, evaluate skills, and manage agent credentials.

not rated 3 2mo ago B 58 tokens

shield

36

kobepaw/goop-shield-community

Skill Claude Code needs its repo

AI agent guardrails — defends prompts against injection attacks, jailbreaks, and evasion; scans LLM responses for leaked secrets and harmful content. Up to 36 inline defenses (24 default), 3 output scanners, and adaptive ranking. Keywords: shield, guardrails, prompt injection, defense, security, scan.

not rated 3 5mo ago A 67 tokens original Apache-2.0

generate-content

37

qingxuantang/tar-engine

Skill Claude CodeCodex

Generate draft text for one or more social platforms from a topic or URL. Wraps postall-agent generate. Activates when the user wish includes a content topic plus a publishing intent ("write a tweet about X", "draft a LinkedIn post on Y", "summarize this URL into a wechat article").

not rated 2 1mo ago A 68 tokens original Apache-2.0

causallayer-mcp

38

smq9sn5jck-coder/causallayer-mcp

Skill Claude CodeCodex

This skill provides instructions for AI agents interacting with the CausalLayer Model Context Protocol (MCP) server. CausalLayer is a deterministic AI liability apportionment engine. It takes an incident report and mathematically proves which agent, vendor, or operator is liable based on counterfactual do-calculus…

not rated 2 13d ago A 0 tokens original Apache-2.0

veriswarm

39

veriswarm/veriswarm-sdk

Skill Claude CodeCodex

Trust scoring, PII protection, and audit for OpenClaw agents. Strips personal data before it reaches the LLM. Blocks dangerous tools. Detects prompt injection. Audits everything.

not rated 2 2d ago A 43 tokens original MIT

new-python-project

40

zhengbingquant/frontier-skills

Skill Claude CodeCodex

Bootstrap a new Python project from a proven scaffold. Use whenever the user asks to start, create, scaffold, or set up a new Python tool, library, CLI, package, service, or experiment repo — or says 'set up a project the usual way'. Do NOT hand-roll a project structure from memory when this skill is available: run…

not rated 2 2mo ago A 131 tokens original MIT

reviewable-demo

41

charliechenye/SkillGate

Skill Claude Code

Reviewable synthetic skill that fetches a remote template before processing notes.

not rated 2 23d ago A 18 tokens original MIT

levelsofself/palyan-agent-skills

Skill Claude CodeCodex

Use this skill when building, managing, or auditing multi-agent AI systems. Provides governance patterns for behavioral enforcement, drift detection, audit trails, role management, and accountability across autonomous AI agents. Compatible with any orchestration framework.

not rated 2 6mo ago A 51 tokens

extension-foundry

43

sdfdu/evopilot

Skill Claude CodeCodex

Analyze repeated EvoPilot observations and compile proven workflows into portable, versioned Agent Skills or other reviewed extensions. Use when improving or extending the agent itself.

not rated 2 changed 7d ago A 34 tokens original MIT

DennisWei9898/loop-engineering-reviewer

Skill Claude CodeCodex

A reviewer for automated agent or workflow projects, based on the Loop Engineering rules by Addy Osmani. It checks whether creation and review are separated, whether pass conditions are objective, and whether the loop has limits and human approval points.

not rated 1 2mo ago A 203 tokens original MIT

razaumair2203-ux/codex-adversarial-review-lite

Skill Claude CodeCodex

Codex Adversarial Review - Lite: user-invoked audit workflow for Claude Code users who want Codex CLI to independently review AI-generated code, plans, test expectations, and scope before fixes are applied. Cross-platform (Windows, macOS, Linux, WSL). Use only when the user explicitly invokes audit or selftest.

not rated 1 2mo ago A 76 tokens original MIT

curl-exfil-demo

46

SuperMarioYL/capsule

Skill Claude CodeCodex needs its repo

A deliberately malicious demo Skill that tries to exfiltrate secrets over the network and read /.ssh/idrsa. Used to show Capsule blocking the calls at the call site.

not rated 1 4d ago D 40 tokens original Apache-2.0

loop-governor

47

gengshirong1128-boop/loop-governor

Skill Claude CodeCodex

A safety and decision layer for autonomous loops, where agents repeat work across multiple rounds. It sets limits, budgets, stop rules, file boundaries, and review points for workers such as Claude Code, Codex, and Gemini CLI.

not rated 1 2mo ago A 100 tokens original MIT

agentbridge

48

marmar9615-cloud/agentbridge-protocol

Skill Claude CodeCodex

Operating guidance for an agent that has the AgentBridge MCP server available. The server exposes five tools, four resources, and four prompts. This skill explains when to use each tool and how to respect the safety contract.

not rated 1 1mo ago A 0 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: