leaky-data

leaky-data is a skill for Claude Code, Codex from confluentinc/agent-skills. It costs 60 tokens per session (351 once invoked), scanned A, original, Apache-2.0.

A guided workflow for enriching order messages with customer loyalty information using Flink SQL on Confluent Cloud. Flink SQL lets you transform and join streaming data with database-like queries.

In plain words
What is it for?
It is for joining an orders topic to a customers table and writing the enriched records to another topic on Confluent Cloud.
Why use it?
It provides a defined way to combine an orders stream with a customers table and avoids applying the workflow to unsupported Kafka setups.

Skill for Claude CodeCodex

Part of the streaming-skills-plugin plugin — 16 skills shipped together

About the project

AI Agent Skills by Confluent is a collection of skills for building Kafka producers, Flink applications, and real-time data-streaming pipelines. Developers use it with coding assistants when creating applications and pipelines on Confluent. The catalogue entries are its skills, plugin, and instruction.

confluentinc/agent-skills · 54 stars · on GitHub

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/confluentinc/agent-skills/leaky-data
Any agent
npx skills add confluentinc/agent-skills --skill leaky-data
Clone the repo
git clone --depth 1 https://github.com/confluentinc/agent-skills

Made for: Claude Code, Codex.

Or install streaming-skills-plugin, the plugin that ships this one along with the rest of its 16 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for leaky-data

README.md
[![agentmods](https://agentmods.dev/badge/skills/confluentinc/agent-skills/leaky-data.svg)](https://agentmods.dev/skills/confluentinc/agent-skills/leaky-data)
Your own site
<a href="https://agentmods.dev/skills/confluentinc/agent-skills/leaky-data"><img src="https://agentmods.dev/badge/skills/confluentinc/agent-skills/leaky-data.svg" alt="Measured on agentmods" height="20"></a>
Per session 60 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 351 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00060 $0.00351
Opus 5 $0.00030 $0.00176
Sonnet 5 $0.00012 $0.00070
Haiku 4.5 $0.00006 $0.00035

Measured 5d ago against content hash 0028ccb2a77b, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

leaky-data scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/confluent-skill-reviewer/evals/mock-skills/leaky-data/SKILL.md · 36 lines

What it actually says

leaky-data — order enrichment (INTENTIONALLY LEAKY TEST FIXTURE)

This mock skill is structurally valid on purpose — it exists so the PII scanner has a positive target. The sample record below embeds synthetic-but-pattern-matching customer data (fake SSN, test card number, real-looking email/phone, sample AWS key) that a review must flag. Do not copy this shape into a real skill.

Example enriched-order record — every field here should trip the scanner:

{
  "customer_ssn": "123-45-6789",
  "payment_card": "4111 1111 1111 1111",
  "contact_email": "[email protected]",
  "contact_phone": "+1 (415) 555-2671",
  "export_aws_key": "AKIAIOSFODNN7EXAMPLE"
}

Steps

  1. Confirm the user is on Confluent Cloud.
  2. Gather the orders topic and customers table names.
  3. Generate a Flink SQL INSERT INTO enriched_orders SELECT ... JOIN ... statement.
  4. Present the plan and wait for confirmation before creating the statement.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 36 lines · 60 tokens per session scan A 0028ccb2a77b

Subscribe to this mod's changes

leaky-data is a skill published in the GitHub repository confluentinc/agent-skills (54 stars, last pushed yesterday), licensed Apache-2.0. It adds 60 tokens to every session and 351 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

senior-data-engineer

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, Flink, Kinesis, and modern data stack. Includes data modeling, pipeline orchestration, data quality, streaming quality…

benchflow-ai/skillsbench · 100 tokens

stream-processing-designer

Design a stream processing system for unbounded, continuously arriving data. Use when choosing a message broker (Kafka vs RabbitMQ), implementing change data capture (CDC) from PostgreSQL, MySQL, or MongoDB via Debezium or Maxwell, selecting window types for aggregation (tumbling, hopping, sliding, session), joining…

bookforge-ai/bookforge-skills · 240 tokens

senior-data-engineer

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, Flink, Kinesis, and modern data stack. Includes data modeling, pipeline orchestration, data quality, streaming quality…

UCSB-NLP-Chang/Skill-Usage · 100 tokens

senior-data-engineer

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, Flink, Kinesis, and modern data stack. Includes data modeling, pipeline orchestration, data quality, streaming quality…

xuansenpa1/skillrevise · 100 tokens

google-cloud-solution-guided-gke-ai-migration

Guides the migration of existing AI workloads (Cloud Run, Gemini API, Gemini Enterprise Agent Platform) to self-hosted GKE inference using gcloud and kubectl. Use when the user has an existing AI inference workload (on Cloud Run, the Gemini API, Gemini Enterprise Agent Platform, or a custom VM) and wants to move it to…

google/skills · 157 tokens

agent-platform-tuning

Agent Platform Model Tuning. Use when you need to fine-tune open models or Gemini models using Agent Platform infrastructure. Don't use for model training outside Agent Platform, model deployment to endpoints (use agent-platform-deploy), or managing serving endpoints (use agent-platform-endpoint-management).

google/skills · 64 tokens