Data lake
Data lake is a repository for structured, semi-structured and unstructured data, often retaining source history alongside curated datasets for analysis. Teams apply schemas, access controls and quality checks according to how the data will be used. A lake supports flexible analysis, but consistent reporting still requires agreed definitions and governance.
Also known as: enterprise data lake
Last reviewed
Why a data lake matters
Most organisations discover they need one when two reports disagree and both are defensible. The rules used to build each report live inside the extracts rather than anywhere a person can read, and there is no single durable copy of what actually happened.
A governed data lake can retain source history so teams can reprocess it when a metric changes. Agree ownership, lineage and business definitions too: storage alone does not reconcile contradictory reports. It also gives history somewhere to live that is not an operational database being slowed down by rows nobody queries.
The cost of that flexibility is discipline. A lake accepts anything, including datasets with no owner, no schema and no retention rule, and a lake full of those is a swamp.
How a data lake works
The pattern is consistent across cloud providers. Raw data lands in cheap object storage, a catalogue records what each dataset is, an access layer controls who can read it, and a query engine reads it in place.
On AWS that is typically:
- Amazon S3 for durable storage of raw and curated zones
- AWS Glue Data Catalog for table metadata and discovery
- AWS Lake Formation, with IAM and KMS, for permissions, encryption and audit
- Amazon Athena for SQL queries against the stored files
- Ingestion suited to the source: scheduled batch extracts, or event streams for telemetry
Azure and Google Cloud offer equivalents. The shape does not change: storage, catalogue, access control, query engine.
Data lake vs data warehouse
The difference is when structure is imposed. A data warehouse applies it on write, through a transformation pipeline you build and maintain, which makes queries fast and answers consistent because somebody already decided what a customer is. A data lake often retains raw data alongside structured, curated datasets. Flexible storage does not guarantee a lower total cost or consistent reporting.
Organisations may use either or both. A lakehouse combines lake storage with table management and analytical capabilities; it is related to a data lake but is not simply another name for one.
When you need one
You need a data lake when data arrives from several systems in shapes an operational database cannot hold, when you want to reprocess history rather than only read the current state, or when someone has asked you to prove how a number was produced. You do not need one when everything lives in a single system that already reports well, or when the real problem is two systems that do not talk to each other — that is an integration job. Governance is the harder half either way: Lake On Rails exists because deciding who owns each dataset is a people problem a platform does not solve.
Further reading
AWS: data lakes and warehouses explains the underlying approach.
Data lake: common questions
A data lake commonly retains raw and varied data for several uses; a warehouse organises data for analytical queries and reporting. Modern platforms overlap, and lakes can contain curated tables. Choose using reporting needs, data types, governance and total operating cost, rather than assuming a lake is always cheaper.
Missing ownership, unclear definitions, poor discovery and unmanaged copies make a lake difficult to trust. Assign an owner, document lineage and access, and agree quality checks and retention for each dataset.
Potentially, but the architecture does not establish compliance. Assess lawful basis, purpose limitation, minimisation, retention and appropriate security, along with other applicable obligations. The ICO's data protection principles are a starting point; get advice for your specific processing.
Want to talk about your project?
Tell us what you’re trying to achieve and we’ll map the fastest credible path.
