AI Safety plugins

75 tagged AI Safety, measured the same way as everything else here.

Browse within: ai-security 33ai-behavior-analysis 29ai-safety-research 29ai-skill-safety 28claude-code-plugin 16ai-governance 8Guardrails 6

duda

26

DavidKim0326/DUDA

Plugin Claude Code

Bundles 1 skill, 1 hook · 137 tokens together

DUDA — Isolation Guardian. Prevents, diagnoses, and recovers isolation contamination in multi-layered architectures.

not rated 2 5mo ago A tokens not measured original Apache-2.0

canary

30

sonomoshq/Canary

Plugin Claude Code

Bundles 5 skills, 1 agent, 3 hooks · 208 tokens together

Sonomos Canary — persistent PII leak counter for Claude Code. 38 checksum-validated regex detectors (plus your own via rules.d) and Claude self-scan catch sensitive data you expose to AI, Canary Tokens give a CERTAIN alarm the instant a planted decoy secret reaches Claude, and Canary Wrapped turns your exposure…

not rated 2 1mo ago A tokens not measured original MIT

document-guard

31

davidmoneil/aifred-document-guard

Plugin Claude Code

Bundles 1 command, 1 hook · 14 tokens together

Prevents Claude from accidentally damaging important files. Intercepts Edit/Write operations and validates them against configurable protection rules including credential scanning, structural preservation, and semantic checks.

not rated 2 4mo ago A tokens not measured original MIT

ArdurAI/ardur

Plugin Claude Code

Bundles 5 hooks

Ardur runtime governance for Claude Code: mission-bound deny gating and signed Execution Receipts on tool calls.

not rated 2 3d ago A tokens not measured original MIT

fact-check

33

Nlai741533/EFC-Plugin

Plugin Claude Code

Bundles 1 skill · 60 tokens together

Systematically fact-check AI-generated research reports against primary sources. Catches the five recurring failure modes of LLM research: unit/scale errors, fabricated interpolation, source conflation, stale data, and attribution laundering.

not rated 2 3mo ago A tokens not measured original MIT

tone-police

35

zircote-plugins/tone-police

Plugin Claude Code

Bundles 1 command, 1 hook · 13 tokens together

Automatically filters angry, hostile, and profane language from user prompts before they reach Claude. Preserves intent while replacing hostility with constructive phrasing.

not rated 1 6mo ago A tokens not measured

auto-cc

36

SSHdotCodes/auto-cc

Plugin Claude Code

Bundles 1 hook

Automatic tool-call review for Claude Code. A local encoder scores every proposed tool call before it runs, approving safe work and blocking dangerous calls with an explanation the agent can act on.

not rated 1 25d ago A tokens not measured original Apache-2.0

shush

38

rjkaes/shush

Plugin Claude Code

Bundles 1 hook

Context-aware safety guard for Claude Code tool calls.

not rated 1 3mo ago A tokens not measured original Apache-2.0

vector

39

pharosone/vector-plugin

Plugin Claude Code

Bundles 4 skills, 1 MCP server · 267 tokens together

Red-team scanning for LLM agents — integrate Vector in your repo, harden against findings, generate agent profiles.

not rated 1 3mo ago A tokens not measured original MIT

vaporcheck

40

cdmx-in/vaporcheck

Plugin Claude Code

Bundles 1 hook, 1 MCP server

Stops AI coding assistants from using things that don't exist — blocks hallucinated/slop-squatted package installs and dead file paths (fail-closed hook + verifyidentifier MCP tool).

not rated 0 1mo ago A tokens not measured original Apache-2.0

ca-sandbox

41

arbiterForge/codeArbiter

Plugin Claude Code

Locally-hosted Codespace equivalent for codeArbiter. Pulls an untrusted repo into an ephemeral, isolated Docker container with no host-filesystem access and configurable egress, caches dependencies by content hash, then tears the box down. Requires Docker and nixpacks on PATH.

not rated 141 yesterday A tokens not measured AGPL-3.0

agile-v-skills

42

Agile-V/agile_v_skills

Plugin Claude Code

Complete AI-augmented engineering framework: Requirements traceability (REQ-XXXX), independent Build Agent + Red Team Verifier, multi-cycle lifecycle management, compliance-ready artifacts (ISO 9001, ISO 27001, GxP).

not rated 51 9d ago A tokens not measured CC-BY-SA-4.0

hs

43

frmoretto/hardstop

Plugin Claude Code

Pre-execution safety layer that blocks dangerous shell commands and credential file reads using pattern matching + LLM analysis. Fail-closed design.

not rated 32 4mo ago A tokens not measured

sodam-loop

46

sodam-ai/SoDam-Loop-Eng

Plugin Claude Code

비개발자용 한국어 안전 AI 반복 루프 (Phase 1a: repair). 위험작업은 항상 확인, 14개 가드레일 기본 ON.

not rated 3 yesterday A tokens not measured

rmanish2000-del/warrant-mcp

Plugin Claude Code

Writes warrant-mcp policies in plain English, shaped for the closed rule set so review has the best chance of accepting them first time: a short interview, sentences shaped for the closed rule set, and the failure shapes explained rather than avoided. The skill writes policy text only — enforcement is warrant-mcp's…

not rated 1 10d ago A tokens not measured MIT

nxtg-forge

48

nxtg-ai/forge-plugin

Plugin Claude Code

Governance-native AI development — 33 agents, 23 commands, 27 skills, 13 security hooks, drift detection, CRUCIBLE test auditing, and shared memory for Claude Code.

not rated 5 today A tokens not measured MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: