ivanshamaev/de-agent-skills

Professional Data Engineering Agent Skills for developing AI Agentic Data Platforms.

This repository also configures its own agents. See what de-agent-skills tells them →

18Stars on the repository
161Mods indexed here, across every type
3mo agoLast push, which is what freshness is scored on
noneNo LICENSE: all rights reserved, so bodies are not copied

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

PostgreSQL for data engineering — declarative partitioning (RANGE/LIST/HASH/pgpartman), index types (B-Tree/BRIN/GIN/GIST/partial/covering), COPY bulk load, EXPLAIN ANALYZE plan reading, autovacuum tuning, window functions, JSONB, CTEs, LATERAL joins, and bulk-load patterns (UNLOGGED/pgbulkload).

not rated 18 3mo ago A 94 tokens

prefect-workflows

146

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Prefect 3.x workflows — flows, tasks, deployments, work pools, event-driven triggers, caching, retries, state hooks, Prefect Cloud/server, Python SDK, Docker/K8s infrastructure.

not rated 18 3mo ago B 45 tokens

pyspark_etl

147

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Use when designing, implementing, reviewing, or optimizing production PySpark ETL/DataFrame pipelines at GB-TB+ scale, including schemas, joins, partitioning, window functions, writes, UDF avoidance, and Spark performance diagnostics.

not rated 18 3mo ago A 53 tokens

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

PySpark Structured Streaming — streaming sources (Kafka/file/rate), output modes (append/complete/update), triggers (ProcessingTime/AvailableNow/Once/Continuous), watermarks, event-time windows (tumbling/sliding/session), stateful aggregations, deduplication, foreachBatch, stream-stream joins, checkpointing, fault…

not rated 18 3mo ago A 89 tokens

rag-data-pipeline

149

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Use when designing, building, or debugging a RAG (Retrieval-Augmented Generation) data pipeline — document ingestion, chunking strategies (fixed/recursive/semantic/structure-aware), embedding models (OpenAI/Cohere/sentence-transformers), vector stores (pgvector/Chroma/Qdrant/Weaviate), incremental refresh, hybrid…

not rated 18 3mo ago A 100 tokens

ray-data

150

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Ray Data distributed data processing — Dataset API, readparquet/readcsv/readjson/readdeltasharing, map/filter/flatmap/mapbatches, groupby/aggregations, Actors for stateful transforms, Ray remote functions, writeparquet/writeiceberg, streaming execution, GPU batch inference, integration with Spark and Airflow.

not rated 18 3mo ago A 69 tokens

redpanda

151

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Redpanda Kafka-compatible streaming — cluster setup, topic config, rpk CLI, producer/consumer tuning, tiered storage, Schema Registry, Kafka Connect compatibility, monitoring, Docker/Kubernetes deployment.

not rated 18 3mo ago B 43 tokens

soda-core

152

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Soda Core data quality — SodaCL checks (rowcount, missing, invalid, duplicate, freshness, schema, reference, custom SQL), configuration.yml for PostgreSQL/Spark/ClickHouse/BigQuery, soda scan CLI, Airflow integration, dbt integration, alerting.

not rated 18 3mo ago A 60 tokens

spark_sql

153

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Use when writing, reviewing, debugging, or optimizing production Spark SQL for Hive/lakehouse/HDFS tables, including CTE-heavy queries, joins, windows, partition pruning, Hive Metastore operations, insert/overwrite safety, query hints, statistics, EXPLAIN plans, AQE, skew, materialization, and SQL performance…

not rated 18 3mo ago A 71 tokens

sqlfluff

154

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

SQLFluff SQL linter — .sqlfluff config, dialect selection (ansi/bigquery/clickhouse/duckdb/hive/postgres/snowflake/sparksql/trino), rule sets, templating (Jinja/dbt), fix command, VS Code integration, pre-commit hook, CI/CD GitHub Actions, custom rules, noqa inline suppression.

not rated 18 3mo ago C 78 tokens

sqlmesh

155

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

SQLMesh data transformation framework — model kinds (FULL/INCREMENTALBYTIMERANGE/INCREMENTALBYUNIQUEKEY/SCDTYPE2/VIEW/SEED), plan/apply workflow, virtual environments, state-aware deploys, audits, unit tests, CI/CD, dbt migration.

not rated 18 3mo ago A 63 tokens

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Terraform for data infrastructure — S3/MinIO data lake buckets (versioning, lifecycle, SSE-KMS), IAM roles for Spark/Airflow (least-privilege, IRSA on EKS), MSK/Kafka clusters (awsmskcluster, encryption, custom broker config), Kubernetes data platform (helmrelease Airflow + Spark, namespace resource quotas), module…

not rated 18 3mo ago A 120 tokens

trino_iceberg

157

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Use when writing, optimizing, or maintaining Apache Iceberg tables with Trino — covering table DDL, all ALTER TABLE operations, partition transforms, bucketing, sorted tables, DML (INSERT/UPDATE/DELETE/MERGE), EXPLAIN plan reading, join optimization, statistics with ANALYZE, table maintenance…

not rated 18 3mo ago A 93 tokens

vertica

158

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Use when writing, reviewing, debugging, or optimizing SQL for Vertica — covering DDL (CREATE/ALTER/DROP TABLE, columns, projections, segmentation, partitions), DML (INSERT, UPDATE, DELETE, MERGE, TRUNCATE, COPY), CRUD patterns, and Vertica-specific performance guidance including encoding, segmentation keys, partition…

not rated 18 3mo ago A 77 tokens

ivanshamaev/de-agent-skills

Skill Claude CodeCodex

Use when optimizing, diagnosing, or reviewing Vertica 11.x SQL query performance — covering EXPLAIN plan reading, projection design for predicates/joins/GROUP BY/ORDER BY/analytic functions, segmentation strategies, column encoding, RLE, sort elimination, Top-K, INSERT-SELECT tuning, DELETE/UPDATE internals, and Data…

not rated 18 3mo ago A 77 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: