
Data Lakehouse
- 50 installs
- 6 repo stars
- Updated March 13, 2026
- alphaonedev/openclaw-graph
data-lakehouse is a Claude skill for designing and implementing data lakehouse architectures (Delta Lake, Iceberg) for scalable big-data storage and analytics.
About
This skill designs and implements data lakehouse architectures that combine data lakes with warehouse features for scalable big-data storage and analytics on platforms like Delta Lake or Iceberg. A developer uses it for large-scale ingestion from S3 or Kafka, analytics workloads needing ACID and schema evolution, or Spark/Presto SQL over unstructured data. It provides CLI init/optimize commands, a REST API, and Spark code patterns.
- Lakehouse architecture design with Delta Lake or Iceberg
- Batch and streaming ingestion via Apache Spark
- Query optimization with caching, indexing, and Z-order clustering
Data Lakehouse by the numbers
- 50 all-time installs (skills.sh)
- Ranked #417 of 911 Databases skills by installs in the Skillselion catalog
- Data as of Jul 28, 2026 (Skillselion catalog sync)
data-lakehouse capabilities & compatibility
- Capabilities
- lakehouse design · data ingestion · query optimization
- Works with
- aws · kafka · databricks · snowflake
- Use cases
- database · data analysis
What data-lakehouse says it does
This skill enables the design and implementation of data lakehouse architectures, combining data lakes with warehouse features for scalable big data storage and analytics.
For analytics workloads needing real-time updates, such as ETL pipelines in e-commerce or IoT data processing.
npx skills add https://github.com/alphaonedev/openclaw-graph --skill data-lakehouseAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 50 |
|---|---|
| repo stars | ★ 6 |
| Last updated | March 13, 2026 |
| Repository | alphaonedev/openclaw-graph ↗ |
What it does
Design and implement data lakehouse architectures (Delta Lake, Iceberg) for scalable big-data storage and analytics.
Who is it for?
Building lakehouse storage and analytics on Delta Lake or Iceberg with Spark.
Skip if: Small datasets that do not need lakehouse-scale storage or ACID tables.
When should I use this skill?
You are building scalable big-data storage with ACID tables and SQL analytics.
What you get
A configured lakehouse (Delta/Iceberg tables) with ingestion and query-optimization set up.
By the numbers
- Supports 2 table formats: Delta Lake and Iceberg
Files
data-lakehouse
Purpose
This skill enables the design and implementation of data lakehouse architectures, combining data lakes with warehouse features for scalable big data storage and analytics. Use it to manage petabyte-scale data with ACID transactions, schema evolution, and optimized query performance on platforms like Delta Lake or Iceberg.
When to Use
- When handling large-scale data ingestion from sources like S3 or Kafka, requiring both raw storage and structured querying.
- For analytics workloads needing real-time updates, such as ETL pipelines in e-commerce or IoT data processing.
- If you're integrating with Spark or Presto for SQL analytics on unstructured data.
Key Capabilities
- Architecture Design: Generate blueprints for lakehouse setups, including partitioning strategies and metadata management (e.g., using Iceberg for table formats).
- Data Ingestion: Support for batch and streaming ingestion with tools like Apache Spark, handling formats like Parquet or ORC.
- Query Optimization: Implement caching and indexing for faster queries, such as creating Delta Lake tables with Z-order clustering.
- Scalability: Auto-scale storage and compute resources via cloud APIs, e.g., AWS Glue for ETL jobs.
- Security: Enforce row-level access controls using policies like AWS Lake Formation grants.
Usage Patterns
- Pattern 1: For new lakehouse setup, invoke the skill to generate a configuration file, then use it to initialize storage. Example: Create a Delta Lake table from CSV data.
- Pattern 2: In analytics workflows, use the skill to optimize queries by adding indexes, then execute via Spark SQL.
- Pattern 3: For maintenance, periodically run checks for schema evolution and merge operations on existing tables.
Common Commands/API
Use the OpenClaw CLI or API for this skill. Authentication requires setting $DATA_LAKEHOUSE_API_KEY as an environment variable.
- CLI Command: Initialize a lakehouse project:
openclaw skill data-lakehouse init --project my-lakehouse --storage s3://my-bucket --engine deltaThis creates a basic configuration file with S3 bucket and Delta Lake engine.
- API Endpoint: Create a table via POST request:
curl -H "Authorization: Bearer $DATA_LAKEHOUSE_API_KEY" \
-d '{"table_name": "sales_data", "format": "parquet", "partition_by": ["date"]}' \
https://api.openclaw.ai/data-lakehouse/tablesResponse includes table metadata for immediate use.
- Code Snippet: In Python, integrate with Spark:
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("lakehouse").getOrCreate()
df = spark.read.format("delta").load("s3://my-bucket/sales_data")
df.write.format("delta").mode("append").save("s3://my-bucket/sales_data")This appends data to a Delta table; ensure Spark is configured with AWS credentials.
- Config Format: Use JSON for lakehouse configs, e.g.:
{
"storage": "s3://my-bucket",
"engine": "iceberg",
"auth": {"key": "$DATA_LAKEHOUSE_API_KEY"}
}Load this via CLI: openclaw skill data-lakehouse apply --config path/to/config.json.
Integration Notes
- With Other Skills: Link to "big-data" skills by passing outputs, e.g., pipe data from a Kafka ingestion skill into this one using Spark streaming.
- External Tools: Integrate with AWS S3 by setting bucket policies; use
--aws-region us-west-2in CLI commands. For Spark, ensure dependencies likespark.deltaare in your environment. - API Integration: When calling from other services, handle retries for rate limits; example: Use the same
$DATA_LAKEHOUSE_API_KEYin chained API calls. - Dependency Management: Always specify versions, e.g., require Spark 3.0+ for Delta Lake compatibility.
Error Handling
- Common Errors: Handle authentication failures by checking if
$DATA_LAKEHOUSE_API_KEYis set; useos.environ.get('DATA_LAKEHOUSE_API_KEY')in scripts. - Prescriptive Steps: For storage access errors (e.g., S3 permissions), add try-except blocks:
try:
spark.read.format("delta").load("s3://my-bucket/data")
except Exception as e:
print(f"Error: {e}. Check bucket permissions and retry.")Retry transient errors like network issues with exponential backoff in API calls.
- Debugging: Use CLI flag
--verbosefor detailed logs, e.g.,openclaw skill data-lakehouse init --verbose. Validate configs withopenclaw skill data-lakehouse validate --file config.json.
Concrete Usage Examples
- Example 1: Building a sales analytics lakehouse:
First, initialize: openclaw skill data-lakehouse init --project sales-ana --storage s3://sales-data. Then, ingest data: Use the API to create a table, and append via Spark as shown above. Finally, query with: spark.sql("SELECT * FROM sales_data WHERE date > '2023-01-01'").
- Example 2: Optimizing an existing lakehouse for IoT data:
Run: openclaw skill data-lakehouse optimize --table iot_metrics --add-index date. This adds an index; verify with a query snippet: df = spark.read.format("iceberg").load("s3://iot-bucket/metrics").filter(df.date > current_date()). Monitor performance post-optimization.
Graph Relationships
- Related to: "big-data" (shares tags for data processing pipelines), "data-engineering" (cluster affiliation for ETL workflows).
- Connected via: Tags like "data-lakehouse" for cross-skill queries, and "data-engineering" cluster for sequential task flows.
Related skills
FAQ
Which table formats does data-lakehouse support?
Delta Lake and Iceberg, with ACID transactions and schema evolution.
How does it ingest data?
Batch and streaming ingestion with tools like Apache Spark, handling Parquet or ORC formats.