list-tests

list-tests is a skill for Claude Code, Codex from VantaInc/vanta-mcp-plugin. It costs 16 tokens per session (731 once invoked), scanned A, original, MIT.

A report of failing Vanta compliance tests, sorted by how likely they are to be fixed from the current repository. Vanta is a tool that checks whether systems meet security and compliance requirements.

In plain words
What is it for?
Use it to review tests needing attention in AWS, Google Cloud, Azure, CloudFormation, or CDK projects and identify which ones can be addressed with repository code.
Why use it?
It helps separate repository changes from tests that need work in another cloud service or in Vanta itself.

Skill for Claude CodeCodex

Part of the vanta-mcp-plugin plugin — 3 skills shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/vantainc/vanta-mcp-plugin/list-tests
Any agent
npx skills add VantaInc/vanta-mcp-plugin --skill list-tests
Clone the repo
git clone --depth 1 https://github.com/VantaInc/vanta-mcp-plugin

Made for: Claude Code, Codex.

Or install vanta-mcp-plugin, the plugin that ships this one along with the rest of its 3 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for list-tests

README.md
[![agentmods](https://agentmods.dev/badge/skills/vantainc/vanta-mcp-plugin/list-tests.svg)](https://agentmods.dev/skills/vantainc/vanta-mcp-plugin/list-tests)
Your own site
<a href="https://agentmods.dev/skills/vantainc/vanta-mcp-plugin/list-tests"><img src="https://agentmods.dev/badge/skills/vantainc/vanta-mcp-plugin/list-tests.svg" alt="Measured on agentmods" height="20"></a>
Per session 16 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 731 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00016 $0.00731
Opus 5 $0.00008 $0.00365
Sonnet 5 $0.00003 $0.00146
Haiku 4.5 $0.00002 $0.00073

Measured 4d ago against content hash e725a61375c9, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

list-tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/list-tests/SKILL.md · 32 lines

How it starts

The opening of the file, as written. The whole thing — 32 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Show the user their failing Vanta tests, ranked by what the plugin can help with.

Steps

  1. Fetch failing tests. Call tests to get all tests with status NEEDS_ATTENTION.
  2. Categorize and rank tests. Group the failing tests into tiers: Ready to fix — Tests where:
    • The test's integration matches resources likely managed in this repo. Detect this by checking for deployment code: look for provider declarations (provider "aws" in .tf files for AWS, provider "google" for GCP, provider "azurerm" for Azure) and resource type prefixes (aws_, google_, azurerm_) in .tf files; or AWSTemplateFormatVersion in CloudFormation templates; or cdk.json for CDK projects. Use both signals — provider blocks are often absent in child modules or Terragrunt configs.
    • Present these first. These are one-command fixes with /vanta:fix-test <testId>. Fixable with guidance — Tests that are code-remediable but may not match this repo (different cloud provider, different integration). The user can still get remediation code, but may need to apply it elsewhere. Manual steps needed — Tests that require configuration changes in external tools, Vanta settings, or manual processes. The plugin can provide guidance but not generate code.
  3. Present the results. For each tier, show a table with columns:
    • Test name
    • Test ID
    • Number of failing entities
    • Integration (e.g., AWS, GitHub, Azure)
    • How long the test has been failing (from latestFlipDate)
    • For "Ready to fix" tests, show: Run /vanta:fix-test <testId> to generate a PR
  4. Highlight co-failure clusters. If multiple failing tests map to the same resource type or integration, note this. For example: "5 IAM tests are failing — fixing the password policy may resolve all of them at once."
  5. Keep it scannable. Use a table or bulleted list. Do not dump raw API responses. The user needs to quickly see what to fix first.

Edge cases

  • No failing tests: "All tests are passing. Nice work." Do not show an empty table.
  • User asks to filter (e.g., "show AWS tests"): Filter by integration name. If no failures match the filter, say so and show the full list: "No failing AWS tests found. Here's what is failing across other integrations:"
  • User asks to filter by framework (e.g., "SOC 2 gaps"): Filter by framework. "You have [N] failing tests mapped to SOC 2. Here are the ones I can help fix from this repo."
  • User asks "what should I fix first?": Rank by impact: IaC-fixable in this repo first, then highest entity count, then longest time failing. Highlight co-failure clusters as "biggest bang for the buck."
  • Very large number of failing tests: Group by integration and summarize counts rather than listing every test. Show the top 5-10 highest-impact items with a note: "[N] more tests failing. Want to see the full list or focus on [integration]?"

Read the full file on GitHub · 32 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 32 lines · 16 tokens per session scan A e725a61375c9

Subscribe to this mod's changes

list-tests is a skill published in the GitHub repository VantaInc/vanta-mcp-plugin (4 stars, last pushed 3mo ago), licensed MIT. It adds 16 tokens to every session and 731 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens

chat-perf

Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.

microsoft/vscode · 51 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens