Skip to content
Getting Digital

Data Lake

Also: data lakehouse, lakehouse, raw data store, object storage for analytics, data swamp

A data lake is a store that keeps data in its original form, files of any shape in cheap object storage, loaded before anyone has decided what questions will be asked of it, so that the structure is applied when the data is read rather than when it is written.

Assessment. A data lake preserves raw data for questions not yet asked, and that value depends on a catalogue. Raw files that cannot be found, whose columns cannot be explained and whose owner has left become unusable, and maintaining the catalogue is work that the advocates of the lake rarely wish to do. The curation should be budgeted together with the storage.

James Dixon, then at Pentaho, coined the term in 2010 with a comparison that still explains it: a data mart is bottled water, cleaned and packaged for easy consumption, and a data lake is the body of water in its natural state, which the contents stream into and which different users can examine, dive into or sample. The point was the questions. A mart answers the ones it was designed for; a lake keeps everything, including the fields nobody thought to model, so that the question that arrives in two years can still be answered from the original records.

The architecture is files on object storage, with a layer of tooling over them. Logs, exports, event streams, images and documents land as they are, usually in a raw zone, and are refined in stages, cleaned and standardised in a second zone, aggregated and modelled in a third, by pipelines that read the files and write new ones. Query engines read the files in place with SQL; data scientists read them with Python. Schema-on-read is the phrase: the structure is imposed by whoever reads, which gives flexibility and removes the guarantee a data warehouse offers, that a column means one thing.

ZoneContentsWho reads it
RawFiles as they arrived, with arrival timePipelines; engineers replaying a load
RefinedCleaned, typed, standardised recordsData scientists; pipelines feeding the warehouse
ModelledAggregated tables with agreed definitionsAnalysts and reports

Lake or swamp

The difference is a catalogue: what each dataset is, where it came from, who owns it, what each column means, how fresh it is and who may read it. Without one, the lake is an expense for storage that no one can use.

Since 2021 the two architectures have been converging under the name Armbrust and colleagues gave the pattern, the lakehouse: table formats over object storage that add transactions, schema enforcement and history to the files, so that warehouse-style tables and raw files live in one platform and one engine queries both. Microsoft's Fabric, Databricks and the cloud warehouses all now describe themselves this way, which is why the Fabric Data Engineer and Databricks Data Engineer exams test lake tables and warehouse modelling together, and why the data engineering courses treat the lake as the landing zone and the warehouse as the modelled layer on top rather than as rivals.

In practice

  • The landing: a logistics company streams every scan event from its depots into object storage as it happens, hundreds of millions of rows a month, with no schema beyond the event's own fields.
  • The warehouse answers: on-time delivery by region, from a modelled table refreshed nightly.
  • The lake answers: two years later, a question nobody modelled: do parcels scanned at a particular depot between two shifts go missing more often? The raw events still have the shift timestamp the warehouse model dropped.
  • The swamp version: the same events, with the depot code renamed three times and nobody who remembers which is which.

Often confused with

Data Warehouse
A warehouse holds modelled tables with agreed definitions for analysts and reports; a lake holds raw files for pipelines and data scientists. The lake usually feeds the warehouse, and lakehouse platforms now hold both.
Object Storage
Object storage is the service the files sit in; a data lake is the practice of organising, cataloguing and querying those files for analysis. A bucket full of exports is storage; a bucket with zones, a catalogue and a query engine is a lake.

Key takeaways

  • →Raw data kept in its original form and structured when read, for questions not yet asked.
  • →Zones from raw to modelled, pipelines between them, engines reading in place.
  • →The catalogue separates a lake from a swamp; the lakehouse merges the lake with the warehouse.

Related concepts

  • Raw files beside modelled tables; the lake usually feeds the warehouse.

  • Learn firstObject Storage

    The storage service the lake's files sit in.

Where this concept sits in the field

Certifications that test this

Vendor exams whose syllabus covers this concept: facts, cost and a preparation path on each page.

More courses from these categories

Courses from the categories where this concept is taught. Details, price and the provider link are on each course page.

FAQ

Does a data lake replace a data warehouse?
No; it precedes and feeds it. Analysts and reports need agreed definitions and fast tables, which is the warehouse; pipelines and data scientists need the raw records, which is the lake. Lakehouse platforms put both on one storage layer.
Who uses a data lake directly?
Data engineers building pipelines, data scientists training models on raw records, and anyone with a question the warehouse model did not anticipate. Business users read the modelled layer and rarely touch the files.
How much data justifies a lake?
Variety more than volume. A business with event streams, documents and images beside its transactional tables benefits at modest scale; one with three relational systems and no unstructured data needs a warehouse and a raw landing schema, not a lake.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 3 October 2026 · Getting Digital