agent safety skills

65 tagged agent safety, measured the same way as everything else here.

Browse within: ai-governance 35claude-plugins 35claude-code-marketplace 7approval-workflow 6audit-trail 6autonomous-agent 6

configure

01

SponsioLabs/Sponsio

Skill Claude CodeCodex

Use after /plugin install sponsio-claude-code to wire the runtime end-to-end. The plugin install only registers hooks + skills; the contract library and per-environment overrides are configured here. Bootstraps the per-plugin contract library tree at /.sponsio/plugins/, installs bundled starter libraries for popular…

438 4d ago E 263 tokens original Apache-2.0

sponsio

02

SponsioLabs/Sponsio

Skill Claude CodeCodex

Install, observe, tune, and enforce Sponsio: a runtime contract layer for LLM agents that blocks unsafe tool calls and scores output quality against declared rules. Use when the user wants to set up / add / install Sponsio, add guardrails or runtime safety to an LLM agent, generate or refine a sponsio.yaml, audit tool…

438 4d ago A 203 tokens original Apache-2.0

sponsio

03

SponsioLabs/Sponsio

Skill Claude CodeCodex

Install, observe, tune, and enforce Sponsio: a runtime contract layer for LLM agents that blocks unsafe tool calls and scores output quality against declared rules. Use when the user wants to set up / add / install Sponsio, add guardrails or runtime safety to an LLM agent, generate or refine a sponsio.yaml, audit tool…

438 4d ago A 203 tokens original Apache-2.0

Hyperion-GPU/ProofFlow-v0.1

Skill Claude CodeCodex

Use ProofFlow from Codex for Agent Work Ledger and maintainer workflows: wrap complex AI coding tasks in a Ledger, review diffs, export Proof Packets, triage issue text, and keep claims evidence-backed.

48 25d ago A 50 tokens original MIT

bridgeward

06

bridge-mind/BridgeWard

Skill Claude CodeCodex

Skeptical-reading and prompt-injection defense for AI agents. Activate whenever the agent reads externally-sourced or potentially-untrusted content — web pages, fetched URLs, search results, GitHub issues / PRs / comments / diffs, emails, Slack/Discord messages, RSS feeds, scraped HTML, MCP tool descriptions, MCP tool…

38 4mo ago F 213 tokens original MIT

injection-audit

07

bridge-mind/BridgeWard

Skill Claude CodeCodex

Audit a file, directory, web page, or piece of content for prompt-injection attempts. Use when reviewing untrusted content (scraped pages, downloaded files, third-party repos, MCP server tool descriptions, email archives, search-result corpora, RAG documents, code-review diffs) for hidden or visible attempts to…

38 4mo ago D 88 tokens original MIT

afu-brain

08

norika1207-lab/afu-brain

Skill Claude CodeCodex

Use Afu Brain as a MASL butler brain before executing OpenClaw or other agent tools. It turns private owner memory and human intent into safe decisions, skill routing, and approval gates.

23 3mo ago A 44 tokens

curiositech/windags-skills

Skill Claude CodeCodex

Safety, privacy, cost management, and frank advice for building always-on AI agents with episodic memory. Covers data hygiene, privacy risk surfaces, runaway cost prevention, scope creep, psychological effects of persistent AI companions, and responsible deployment patterns. This is the skill that tells you what can…

10 1mo ago A 166 tokens

agents-md-optimizer

10

MSApps-Mobile/claude-plugins

Skill Claude CodeCodex

Audit bloated agent-instruction files (CLAUDE.md, AGENTS.md, and their local/user-level variants) and rewrite them lean, or author a new one from scratch following the 200-line "recipe book" principle. Use this skill whenever the user mentions CLAUDE.md, AGENTS.md, agent instruction file, project memory file, context…

9 6d ago A 208 tokens original MIT

cowork-mem

11

MSApps-Mobile/claude-plugins

Skill Claude CodeCodex

Persistent memory across Cowork sessions. Use this skill at the START of every session to recall what happened before, and throughout any session to save important context — decisions, file changes, insights, errors, tool usage. Trigger whenever: the user says "remember this", "what did we do last time", "save this"…

9 6d ago A 157 tokens original MIT

MSApps-Mobile/claude-plugins

Skill Claude CodeCodex

Run a Google Cloud CLI (gcloud) health check — verify installation, authentication, active project, recommended region, enabled APIs (run, cloudbuild, artifactregistry, secretmanager, cloudresourcemanager, iam, plus gmail/calendar/monitoring/logging), API access token, billing, IAM bindings, service accounts, and…

9 6d ago A 175 tokens original MIT

polygraph

13

polygraphso/litmus

Skill Claude CodeCodex

Behavioral trust grades (A–F) for MCP servers. Use when an agent needs to check whether an MCP server is safe before using it, verify an onchain attestation before trusting or paying a server, look up a server's published grade, get a project graded, or understand why a server received a grade. Polygraph connects to…

8 1mo ago A 225 tokens original Apache-2.0

consolidate

14

choiyounggi/groundwork

Skill Claude CodeCodex

Periodically merge the memory index and long-tier memory files — deduplicate, resolve contradictions to the current truth, absolutize dates — and propose the result for your confirmation before any write. Use when MEMORY.md grows large (120+ lines) or long-tier memories accumulate duplicates, contradictions, or stale…

3 2d ago A 65 tokens original MIT

remember

15

choiyounggi/groundwork

Skill Claude CodeCodex

Save gate for persistent memories — confirm tier (long/short) and expiry with the user before writing, so hallucinated or transient facts never enter long-term memory. Use whenever saving a memory file, and when the expiry sweep reports archived memories.

3 2d ago A 51 tokens original MIT

setup

16

choiyounggi/groundwork

Skill Claude CodeCodex

First-time memory-loop setup — offer identity names, create HABITS.md from the template, write an initial config, and verify the hooks respond.

3 2d ago A 31 tokens original MIT

aishield

17

lm203688/aishield

Skill Claude CodeCodex

An AI and MCP security scanner aligned with the OWASP MCP Top 10, a list of common risks for tool-connected AI systems. It checks for issues such as prompt injection, command injection, poisoned tools, leaked secrets, unsafe permissions, supply-chain attacks, SSRF, and banned Chinese-language content.

2 2d ago A 101 tokens original MIT

ASER-ho/coding-agent-safety-gate

Skill Claude CodeCodex

Skill "coding-agent-safety-gate" from ASER-ho/coding-agent-safety-gate, covering coding-agent-safety-gate, skill name / 技能名称, purpose / 用途, suitable scenarios / 适用场景 and unsuitable scenarios / 不适用场景.

2 1mo ago A 0 tokens original MIT

omarkhandji-commits/midas

Skill Claude CodeCodex

When to use. The operator has identified an ICP (ideal customer profile) and wants a deliverability-aware sequence to test it.

1 1mo ago A 0 tokens original MIT

fiverr-gig

21

omarkhandji-commits/midas

Skill Claude CodeCodex

When to use. The operator wants a new Fiverr gig that ranks for buyer search terms in their category, or wants to refresh an underperforming one.

1 1mo ago A 0 tokens original MIT

reddit-launch-post

22

omarkhandji-commits/midas

Skill Claude CodeCodex

When to use. The operator wants to launch (or relaunch) a product on a relevant subreddit without getting mod-banned in the first hour.

1 1mo ago A 0 tokens original MIT

airlock

23

cjaston/airlock

Skill Claude CodeCodex

Use Airlock before installing or executing packages, running risky shell commands, committing generated code, or declaring an agent task complete. Airlock detects hallucinated/slopsquatted dependencies, destructive commands, leaked secrets, and suspicious test changes.

1 2mo ago E 50 tokens original MIT

AaronVick/AGENT_FABLES

Skill Claude CodeCodex

Retrieve revision-pinned software-agent failure evidence and unresolved verification gates without authorizing execution. Use before destructive, irreversible, privileged, broad-scope, production, infrastructure, database, filesystem, permission, MCP configuration, or supply-chain operations; when reviewing a…

0 16d ago A 79 tokens