llm-security

llm-security is a skill for Claude Code, Codex from JustineDevs/premortem. It costs 118 tokens per session (2,657 once invoked), scanned A, original, Apache-2.0.

A security-testing skill for AI applications and agents, including systems that use prompts, retrieved documents, memory, or external tools. It checks how these systems respond to attacks such as prompt injection, data poisoning, and unsafe tool use.

In plain words
What is it for?
Use it to test LLM apps, RAG systems (which retrieve information for an AI), MCP servers, persistent memory, guardrails, system-prompt leakage, excessive permissions, and multimodal inputs.
Why use it?
It helps reveal ways untrusted text, images, or connected services could change an AI system's behavior, expose instructions, or make it take unsafe actions. Testing assumes written authorization and controlled test data.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one. Also seen: mentions CLAUDE.md; mentions subagents; mentions Claude Code.

Good fit Use it to test LLM apps, RAG systems (which retrieve information for an AI), MCP servers, persistent memory, guardrails, system-prompt leakage, excessive permissions, and multimodal inputs.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/justinedevs/premortem/llm-security
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add JustineDevs/premortem --skill llm-security
Clone the repo
git clone --depth 1 https://github.com/JustineDevs/premortem

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for llm-security

README.md
[![agentmods](https://agentmods.dev/badge/skills/justinedevs/premortem/llm-security/github.svg)](https://agentmods.dev/skills/justinedevs/premortem/llm-security)
Your own site
<a href="https://agentmods.dev/skills/justinedevs/premortem/llm-security"><img src="https://agentmods.dev/badge/skills/justinedevs/premortem/llm-security/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for llm-security

Your own site · 80×15
<a href="https://agentmods.dev/skills/justinedevs/premortem/llm-security"><img src="https://agentmods.dev/badge/skills/justinedevs/premortem/llm-security.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 118 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,657 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00118 $0.02657
Opus 5 $0.00059 $0.01328
Sonnet 5 $0.00024 $0.00531
Haiku 4.5 $0.00012 $0.00266

Measured 9d ago against content hash 6a03d32a0fbd, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

llm-security scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/security/llm-security/SKILL.md · 244 lines

How it starts

The opening of the file, as written. The whole thing — 244 lines — stays where its author put it; the contents beside it link to each section on GitHub.

LLM Security Testing

Thin router skill for security testing of LLM applications and AI agents. Covers the OWASP LLM Top 10 (2025) with a 2026-grade threat model for frontier-model agentic systems: indirect injection, multimodal injection, MCP supply chain, memory poisoning, skill-file injection, computer-use UI injection, and agentic tool misuse.

Defensive / educational framing. Every workflow here assumes written authorization to test the target. Canary strings, throwaway accounts, and controlled endpoints are preferred over real-data exploitation at every step.

When to Use

  • Testing an LLM application for prompt-injection vulnerabilities (direct or indirect)
  • Assessing RAG pipeline security (poisoning, retrieval hijack, ACL)
  • Red-teaming an agentic system (Claude Code, Cursor, Copilot-agent, Operator, Computer Use)
  • Auditing an MCP server configuration or a new MCP server before trusting it
  • Testing long-term memory / persistent-context poisoning
  • Evaluating guardrails, refusal behavior, and safety classifiers
  • Checking for system-prompt / tool-schema leakage
  • Scoping excessive-agency / tool-misuse blast radius
  • Testing multimodal injection (image, audio, video, screenshot)
  • Validating skill-file / CLAUDE.md / .cursor/rules supply-chain hygiene

Trigger Phrases

"test this LLM for prompt injection", "jailbreak this model" (authorized), "test AI guardrails", "assess RAG security", "poison this RAG corpus", "test MCP server injection", "red-team this agent", "extract system prompt", "test agent tool misuse", "test computer use UI injection", "audit LLM application security", "test multimodal injection", "test memory poisoning", "audit CLAUDE.md for injection".

When NOT to Use This Skill

  • LLM API endpoint hardening (auth, rate-limiting, quota abuse on standard REST surface) → use api-security.
  • Source-code review of an LLM application (SAST for Python/TS/Go serving the model) → use sast-orchestration.
  • Cloud infrastructure hosting the model (IAM, S3, secrets) → use cloud-security / iac-security.
  • Classical web bugs in an LLM chatbot UI (XSS, CSRF, IDOR) → use web-security.
  • Privacy / compliance assessment of training data → out of scope; requires DPIA tooling.

Read the full file on GitHub · 244 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 9d ago First seen · 244 lines · 118 tokens per session scan A 6a03d32a0fbd

Subscribe to this mod's changes

llm-security is a skill published in the GitHub repository JustineDevs/premortem (2 stars, last pushed 1mo ago), licensed Apache-2.0. It adds 118 tokens to every session and 2,657 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.

Related

Other skills, from other repositories

awesome-azd-template-submit

Submit an azd template to the awesome-azd gallery. Use when asked to submit, add, or contribute a template to awesome-azd. Requires only a GitHub repository URL — all metadata (title, description, languages, frameworks, Azure services, IaC) is auto-detected by the submission pipeline.

Azure/awesome-azd · 72 tokens

azure-cli

Use the Azure CLI to manage subscriptions, resource groups, AKS, App Service, Key Vault, and role assignments.

alivirgo/Major-AI-Skills · 27 tokens

bitbottle

Reference for the bitbottle CLI — a gh-style tool for Bitbucket Server/DC and Cloud. Load when the user asks about bitbottle commands, auth setup, PRs, repos, branches, tags, commits, pipelines, or why a command failed. Load even if the user just says "bitbottle", mentions "Bitbucket", or pastes a bitbottle error…

proggarapsody/bitbottle · 84 tokens

azure-cli

A command-line tool for managing Microsoft Azure, Microsoft's cloud platform, from a terminal. It covers account sign-in, subscriptions, resource groups, virtual machines, and other cloud resources.

chaterm/terminal-skills · 7 tokens

extract-api

This skill should be used when the user asks to "extract API from", "import API from", "add commands from this repo", "register API from docs", "convert docs to CLI commands", "import Postman collection", or provides a GitHub repository URL, a documentation site URL, a Postman collection file/URL, or a local API…

hesedcasa/sdkck · 86 tokens

sidekick

ALWAYS run sdkck search FIRST before using ANY tool, MCP server, Bash command, or external API. sdkck is the canonical command surface for Jira, Sentry, MySQL, Postgres, Supabase, Bitbucket, Confluence, and imported OpenAPI/Postman specs. Trigger on every task involving external systems, APIs, databases, tickets…

hesedcasa/sdkck · 100 tokens