Skip to main content
dataskippr

What I do

Lakehouse data engineering

Reliable pipelines on Azure Databricks, from source system to the Gold layer.

The foundation of every data platform is a reliable Lakehouse. I build a medallion architecture on Azure Databricks with Delta Lake and Lakeflow Declarative Pipelines: idempotent ingestion, tested transformations and orchestration that lets you sleep at night. Performant and cost-conscious.

A pipeline that runs today is not the same thing as a pipeline that is still maintainable a year from now. That is why I build to a small set of fixed principles: idempotent processing, so a restart never produces duplicate data; expectations on every layer, so failures surface where they originate; and incremental loading with Auto Loader or CDC, so the bill doesn't grow with every full reprocess.

The medallion architecture (bronze, silver, gold) is the starting point, not a dogma. For a small platform with three source systems, a light variant often beats the full textbook version. What matters is that each layer has one clear responsibility, and that you can always rebuild from the raw data.

The result is a platform where the numbers are right, failures recover themselves, and the cost per pipeline is visible.

Delta LakeLakeflowAuto LoaderPhotonADLS Gen2

← All services

Ready to set course?

A no-obligation conversation about your Azure Databricks, Lakehouse or AI/BI challenge. I'm happy to think along.

Book a call