The all-in-one data stack for agents. Your data jobs deserve to run on the right tool. Whether you're bringing existing infra, or exploring our open source deployments, oleander empowers your entire data stack to work together as a team.
The all-in-one data stack for agents. Your data jobs deserve to run on the right tool. Whether you're bringing existing infra, or exploring our open source deployments, oleander empowers your entire data stack to work together as a team.
4 22d agoA
tokens not measured
copy · 91%Apache-2.0
Engine-agnostic oleander lake catalog conventions: catalog.namespace.table naming, hierarchy, and catalog-qualified reads/writes without raw storage paths. Use when naming tables, choosing namespaces, or referencing the lake catalog from Spark, Polars, SQL, or another engine.
Runs lake SQL through oleander's query router: queryrun for reads, querysubmit for writes, sparksqlsubmit for named Spark jobs. Use when querying oleander lake tables, exploring data, writing query results to a table, or handling engine routing and billing errors.
Runs Polars queries or scripts on oleander via the CLI, in local or distributed mode, and saves results to the lake catalog. Use when writing Polars jobs, choosing query vs script mode, using --save / --distributed, or wiring scan()/params/result contracts.
General Apache Spark best practices for scalable, maintainable DataFrame jobs: avoid driver materialization, reduce shuffle, join efficiently, and cache carefully. Use when optimizing Spark performance, reviewing PySpark jobs, or writing new DataFrame pipelines.
Spark patterns for reading and writing oleander lake catalog tables: spark.table(), append vs overwrite, and avoiding driver-side writes. Use when building Spark jobs that read or write Iceberg tables in the oleander catalog.
Preserves connected OpenLineage for oleander Spark jobs by avoiding collect()/toPandas() between read and write, and using env vars for runtime config. Use when lineage looks disconnected, jobs split after collect(), or rewriting Spark pipelines for continuous lineage.
Submits, monitors, aborts, and configures Spark jobs on oleander via MCP (sparkjobssubmit), the CLI, and the TypeScript SDK. Use when uploading artifacts, running PySpark jobs, polling run state, or automating Spark workflows.