ingesting-into-data-lake

ingesting-into-data-lake is a skill for Claude Code, Codex from Kilo-Org/kilo-marketplace. It costs 228 tokens per session (2,724 once invoked), scanned A, a copy of ingesting-into-data-lake, Apache-2.0.

A guide for importing files or database data into an AWS data lake, a central store for data used by analytics and processing systems.

In plain words
What is it for?
Use it to ingest data into S3 Tables or Iceberg tables, including migrations from existing Glue Catalog tables.
Why use it?
It helps move data from sources such as S3, PostgreSQL, Snowflake, or BigQuery into queryable lake tables.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to ingest data into S3 Tables or Iceberg tables, including migrations from existing Glue Catalog tables.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/kilo-org/kilo-marketplace/ingesting-into-data-lake
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Kilo-Org/kilo-marketplace --skill ingesting-into-data-lake
Clone the repo
git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for ingesting-into-data-lake

README.md
[![agentmods](https://agentmods.dev/badge/skills/kilo-org/kilo-marketplace/ingesting-into-data-lake/github.svg)](https://agentmods.dev/skills/kilo-org/kilo-marketplace/ingesting-into-data-lake)
Your own site
<a href="https://agentmods.dev/skills/kilo-org/kilo-marketplace/ingesting-into-data-lake"><img src="https://agentmods.dev/badge/skills/kilo-org/kilo-marketplace/ingesting-into-data-lake/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for ingesting-into-data-lake

Your own site · 80×15
<a href="https://agentmods.dev/skills/kilo-org/kilo-marketplace/ingesting-into-data-lake"><img src="https://agentmods.dev/badge/skills/kilo-org/kilo-marketplace/ingesting-into-data-lake.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 228 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,724 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin 95% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00228 $0.02724
Opus 5 $0.00114 $0.01362
Sonnet 5 $0.00046 $0.00545
Haiku 4.5 $0.00023 $0.00272

Measured 7d ago against content hash aaa3265eec02, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade A, and why

ingesting-into-data-lake scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

This is a copy

95% identical to ingesting-into-data-lake — 39 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

skills/ingesting-into-data-lake/SKILL.md · 196 lines

How it starts

The opening of the file, as written. The whole thing — 196 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Ingest into Data Lake

Move data from a source into a queryable table in the data lake. This skill assumes the source connection (if one is needed) already exists. For Glue connection setup or troubleshooting, delegate to connecting-to-data-source.

Philosophy

Default to S3 Tables unless the environment says otherwise. S3 Tables is the recommended target for new data lake work. If the user's catalog inventory shows they haven't adopted S3 Tables, recommend standard Iceberg on their existing general-purpose bucket instead of forcing them to change posture.

Common Tasks

You MUST execute commands using AWS MCP server tools when connected -- they provide validation, sandboxed execution, and audit logging. Fall back to AWS CLI only if MCP is unavailable. You MUST explain each step before executing.

Workflow

1. Verify Dependencies and Context

  • You MUST check whether AWS MCP tools or AWS CLI are available and inform the user if missing
  • You MUST confirm target AWS region and verify credentials with aws sts get-caller-identity
  • For SageMaker Unified Studio project roles, note that target tables and connections may be scoped to the project. See the caller ARN detection pattern in querying-data-lake.

2. Classify the Source

User says... Source type Reference
"upload my file", "local CSV", "move to S3" Local file local-upload.md
"load from S3", "import CSV/JSON/Parquet from s3://" S3 files s3-files.md
"import from Oracle/Postgres/MySQL/SQL Server/Redshift/RDS/Aurora" JDBC jdbc-ingest.md
"pull from Snowflake", "Snowflake table to S3" Snowflake snowflake-ingest.md
"import from BigQuery", "GCP analytics to S3" BigQuery bigquery-ingest.md
"export DynamoDB", "DynamoDB to data lake" DynamoDB dynamodb-ingest.md
"migrate Glue table", "convert Hive to Iceberg" Catalog migration catalog-migration.md

Read the full file on GitHub · 196 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 7d ago First seen · 196 lines · 228 tokens per session scan A aaa3265eec02

Subscribe to this mod's changes

ingesting-into-data-lake is a skill published in the GitHub repository Kilo-Org/kilo-marketplace (175 stars, last pushed 20d ago), licensed Apache-2.0. It adds 228 tokens to every session and 2,724 once invoked, about $0.0011 per session on Opus 5. A static security scan graded it A with 0 findings. It is 95% identical to ingesting-into-data-lake, differing in 39 lines, and is treated as a copy.

Related

Other skills, from other repositories

cloud-databases-onboarding

Guides users through discovering their database requirements, recommends a Google Cloud database based on a recommendation matrix, and assists in database creation. Use when a user asks 'What database service should I use?', 'Help me pick a database', or when a user wants to create a new database on Google Cloud.…

google/skills · 83 tokens

google-cloud-storage-fuse

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when interacting with gcsfuse: decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size file, stat, and list caches, tune…

google/skills · 214 tokens

cloud-sql-basics

This file generates or explains Cloud SQL resources. Use this file when the user asks to create a Cloud SQL instance or database for MySQL, PostgreSQL, or SQL Server. Cloud SQL manages third-party MySQL, PostgreSQL, and SQL Server instances as resources in Cloud SQL. For example, when Cloud SQL creates an open-source…

google/skills · 108 tokens

azure-mgmt-mongodbatlas-dotnet

Manage MongoDB Atlas Organizations as Azure ARM resources using Azure.ResourceManager.MongoDBAtlas SDK. Use when creating, updating, listing, or deleting MongoDB Atlas organizations through Azure Marketplace integration. This SDK manages the Azure-side organization resource, not Atlas clusters/databases directly.

microsoft/skills · 64 tokens

azure-resource-manager-mysql-dotnet

Azure MySQL Flexible Server SDK for .NET. Database management for MySQL Flexible Server deployments. Use for creating servers, databases, firewall rules, configurations, backups, and high availability. Triggers: "MySQL", "MySqlFlexibleServer", "MySQL Flexible Server", "Azure Database for MySQL", "MySQL database…

microsoft/skills · 87 tokens

azure-resource-manager-redis-dotnet

Azure Resource Manager SDK for Redis in .NET. Use for MANAGEMENT PLANE operations: creating/managing Azure Cache for Redis instances, firewall rules, access keys, patch schedules, linked servers (geo-replication), and private endpoints via Azure Resource Manager. NOT for data plane operations (get/set keys, pub/sub) …

microsoft/skills · 114 tokens