James Dixon, then at Pentaho, coined the term in 2010 with a comparison that still explains it: a data mart is bottled water, cleaned and packaged for easy consumption, and a data lake is the body of water in its natural state, which the contents stream into and which different users can examine, dive into or sample. The point was the questions. A mart answers the ones it was designed for; a lake keeps everything, including the fields nobody thought to model, so that the question that arrives in two years can still be answered from the original records.
The architecture is files on object storage, with a layer of tooling over them. Logs, exports, event streams, images and documents land as they are, usually in a raw zone, and are refined in stages, cleaned and standardised in a second zone, aggregated and modelled in a third, by pipelines that read the files and write new ones. Query engines read the files in place with SQL; data scientists read them with Python. Schema-on-read is the phrase: the structure is imposed by whoever reads, which gives flexibility and removes the guarantee a data warehouse offers, that a column means one thing.
| Zone | Contents | Who reads it |
|---|---|---|
| Raw | Files as they arrived, with arrival time | Pipelines; engineers replaying a load |
| Refined | Cleaned, typed, standardised records | Data scientists; pipelines feeding the warehouse |
| Modelled | Aggregated tables with agreed definitions | Analysts and reports |
Lake or swamp
The difference is a catalogue: what each dataset is, where it came from, who owns it, what each column means, how fresh it is and who may read it. Without one, the lake is an expense for storage that no one can use.
Since 2021 the two architectures have been converging under the name Armbrust and colleagues gave the pattern, the lakehouse: table formats over object storage that add transactions, schema enforcement and history to the files, so that warehouse-style tables and raw files live in one platform and one engine queries both. Microsoft's Fabric, Databricks and the cloud warehouses all now describe themselves this way, which is why the Fabric Data Engineer and Databricks Data Engineer exams test lake tables and warehouse modelling together, and why the data engineering courses treat the lake as the landing zone and the warehouse as the modelled layer on top rather than as rivals.
