Scorchsoft
Guide

Data lake vs data warehouse: what does your business need?

For operations, finance and technology leads weighing up a data lake: how it differs from a data warehouse, the signals that say you are ready, what an AWS lake is made of, and how to start small.

Andrew Ward
Author
Andrew Ward
Managing Director
Last reviewed
Reading time
9 min
In this guide
  1. 01What a data lake actually is
  2. 02Data lake vs data warehouse: a business decision matrix
  3. 03Worked example: explaining margin by completed job
  4. 04Compare the full cost of each option
  5. 05Five signs your business is ready for one
  6. 06Four signs you are not ready
  7. 07What an AWS data lake is actually made of
  8. 08Governance is the part people skip
  9. 09How to start without a science project
  10. 10Where Scorchsoft fits

In summary

A data lake holds varied data and history for analysis; a data warehouse organises information for trusted reporting. Modern platforms overlap, so the choice depends on the business question, the source data and the team that will maintain it. This guide compares the options, works through a practical reporting example, explains cost and governance decisions, and shows how to start with one useful outcome.

Key takeaways

  1. A lake holds varied data and history; a warehouse organises information for reporting. Modern platforms overlap, and the workload should decide whether you need one or both.
  2. Disagreeing reports are a reason to investigate definitions, source data and ownership before choosing a new platform.
  3. On AWS a practical lake foundation is S3 for storage, Glue Data Catalog for metadata, Lake Formation with IAM and KMS for permissions, and Athena to query it in place.
  4. Compare total cost: ingestion, transformation, queries, access controls and maintenance as well as storage.
  5. Start with one question the business cannot answer today, name the owner of every dataset it needs, and build a single curated table — not a twelve-month ingestion programme.

What a data lake actually is

A data lake stores data for analysis, including raw records, history and prepared datasets. It can hold structured tables alongside files such as JSON events, documents and device readings. A data warehouse organises data for analytical queries and reporting, with agreed models and definitions. Your operational database serves the application's day-to-day transactions.

Those are different jobs, but modern platforms overlap. A lake can validate incoming data and maintain curated tables. A warehouse can support semi-structured data and transformations after loading. The useful distinction is the workload you need to support and who will maintain the trusted reporting layer. AWS's comparison sets out the main storage and reporting patterns.

Keeping raw history can let you revisit a calculation when a business rule changes. That does not mean keeping everything indefinitely. Decide which sources to retain, who may use them and how long each dataset remains useful.

Data lake vs data warehouse: a business decision matrix

Start with the question people need to answer, then compare the options against it.

  • Recurring management reports: a warehouse or reporting data mart is often a strong starting point. It gives analysts a curated model for measures such as revenue and margin. A lake can feed that model, but the report still needs agreed definitions and quality checks.
  • Raw history and reprocessing: a lake can preserve source records so you can rebuild a calculation, inspect events or prepare a new dataset. Specify identifiers, timestamps and source versions so the history is usable.
  • Mixed files and exploratory analysis: a lake is suited to retaining varied inputs while the team develops useful ways to analyse them. The flexibility creates work: someone must catalogue, secure and prepare those inputs.
  • A live operational screen: use the application database or an appropriate serving layer. A mobile app checking a booking needs predictable transactional behaviour, not an analyst's ad hoc query across raw files.
  • One missing connection: an integration may solve the problem directly. If finance needs approved invoices from the CRM, assess that connection before introducing another storage platform.
  • A small team: compare the skills and ongoing support each option needs. A managed warehouse may be easier to operate for a reporting-focused team; a lake may fit a team that already manages ingestion and analytical pipelines. Neither removes ownership work.

A lakehouse adds capabilities such as managed analytical tables to lake storage, bringing parts of the two approaches together. It is an option to evaluate, not a requirement for every business. You can also combine a lake for selected history with a warehouse for reporting. Avoid buying both simply because an architecture diagram contains both boxes.

Worked example: explaining margin by completed job

Imagine a field-service business whose directors want to understand margin by completed job. This is an illustrative scenario, not a claimed client result. The booking system holds jobs, the finance system holds invoices, and engineers upload timesheets and photographs.

For the first report, the essential inputs may be job identifiers, invoice lines, labour hours and material costs. If those are structured and accessible, an integration and a small reporting model might be enough. The owner of the margin definition must decide how to handle cancelled jobs, credit notes and work completed in a different month from its invoice.

A lake becomes more useful if the business also needs historical snapshots, source events that can be replayed, or new analysis of engineer documents. The photographs do not automatically belong in the financial reporting pipeline merely because storage is available. Give each source a purpose and an access policy.

The first acceptance test is concrete: choose a completed accounting period and reconcile the new report with agreed source totals. Investigate every unexplained difference. Record the refresh time and the treatment of late corrections. A technically successful import with unreconciled figures is not a trusted report.

Compare the full cost of each option

Ask for separate estimates for discovery, ingestion, transformation, storage, queries, permissions and ongoing support. Include the work of defining measures and reconciling them with existing reports. Low object-storage prices alone do not prove that a lake is cheaper than a warehouse for your workload.

Model an ordinary month and a demanding month. Consider data growth, report refresh frequency, concurrent users, large backfills and reprocessing after a rule changes. Include export or migration work if you later move platform. A useful proposal states these assumptions so you can compare alternatives on the same basis.

The data engineering service covers the pipelines and reporting foundations around these choices. For AWS-specific implementation, see AWS data lake development.

Five signs your business is ready for one

You are ready when the questions you want to answer have outgrown the systems holding the answers. Concretely:

  • Two reports disagree and both are right. Revenue in the sales dashboard does not match the finance pack because they were built from different extracts with different rules, and none of those rules is written down outside the files themselves.
  • You are deleting history to keep a database fast. Archiving rows out of an operational database because queries have slowed is the clearest signal that history needs a different home.
  • You want to reprocess the past. A new metric, a corrected calculation or a machine-learning experiment needs the raw events as they were, not the aggregated table someone built in 2019.
  • Data arrives in shapes your database cannot hold. Device telemetry, webhook payloads, documents, images, logs.
  • Someone has asked you to prove a number. A lender, an auditor or a customer's due-diligence team wants to see how a figure was produced and who approved access to the underlying data.

These are reasons to investigate a better data foundation. They do not by themselves prove that a lake is the right architecture; compare a targeted integration, a warehouse and a lake against the same business question.

Four signs you are not ready

A lake needs a clear use and someone to maintain it. These situations suggest starting with a smaller change.

  • Your data lives in one system and that system is fine. If everything is in one ERP or one CRM and the reporting inside it answers your questions, a lake adds a moving part and removes nothing.
  • Nobody owns a dataset. A lake with no owners becomes a swamp within a year. If you cannot name the person who decides what "active customer" means, fix that before you buy storage.
  • The real problem is a broken integration. "We cannot see our data" often means two systems do not talk to each other. That is an API and system integration job, and it is smaller, cheaper and faster than a lake.
  • You want a dashboard by Friday. A lake is infrastructure. It pays back over years. If the need is one urgent report, build the report.

What an AWS data lake is actually made of

On AWS the building blocks are well established, and the list is shorter than most proposals suggest. A practical baseline we build looks like this:

  • Amazon S3 as the durable storage layer, holding raw and curated datasets in separate zones.
  • AWS Glue Data Catalog for table metadata, so a dataset is discoverable rather than a filename somebody remembers.
  • AWS Lake Formation, with IAM and KMS, to enforce permissions, encryption and audit trails.
  • Amazon Athena to query the data in place with SQL, without standing up a cluster.
  • Ingestion that suits the source — scheduled batch extracts for business systems, event streams for telemetry.

We set this baseline out in more detail on our AWS data lake development page. Azure's equivalents are Data Lake Storage, Purview and Synapse or Fabric; Google's are Cloud Storage, Dataplex and BigQuery. The shape of the answer is the same on all three: cheap object storage, a catalogue, an access-control layer and a query engine.

Estimate storage alongside query engines, ingestion and the engineering time to maintain pipelines as source systems change. Use current AWS pricing and your expected workload rather than assuming a low storage bill means a low total cost.

Governance is the part people skip

The technical build is the easy half. The half that decides whether the lake is still useful in two years is the operating model around it: who owns each dataset, which number is trusted for which purpose, how long data is kept, and who approved access to it.

That is a people problem, and it does not get solved by an implementation partner configuring a platform. It gets solved by asking the department head who actually knows the answer, recording the answer against their name, and checking a year later whether it still holds.

We built Lake On Rails for exactly this gap. It sits above an existing AWS or Microsoft Fabric platform, routes each governance decision to the person who can answer it, and records it. It never moves or reads the contents of your data. If you are being pitched a platform and nobody has mentioned ownership, that is the question to ask.

Under UK GDPR you also need a defensible answer on retention and on lawful basis for holding personal data at rest. The ICO's guidance is the reference; storage being cheap is not a reason to keep personal data indefinitely.

How to start without a science project

The failure mode is a twelve-month programme that ingests everything and answers nothing. The alternative is narrow and boring, and it works.

  1. Pick one question the business genuinely cannot answer today, and one that would change a decision.
  2. Name the owner of each dataset that question needs, before you ingest it.
  3. Land the raw data in S3 with a catalogue entry, encryption and an access policy from day one.
  4. Build the one curated table that answers the question, and the one report that uses it.
  5. Write down the rule you used to build it, where someone other than the author can read it.
  6. Then add the second question, and reuse everything from step three.

Document how to recover a failed load, who approves changes and how a replacement team member can operate the pipeline. A second source can reuse that foundation, although its complexity still depends on the source and the quality of its data.

Where Scorchsoft fits

We are a UK application-layer software developer based in Birmingham, with more than 16 years of building apps, portals and integrations. That is where we are strongest on a data project.

What we do: build the applications, portals and integrations that produce and consume the data, and configure the AWS data lake storage, catalogue, permissions and query layer those applications need. We build governance software in Lake On Rails. We are comfortable on AWS, and we will tell you when a lake is not the answer.

We work on data projects of most shapes, often alongside an in-house data team or the supplier who runs the platform underneath. If what you need sits outside our usual remit, we will say so early and can recommend options or make an introduction.

Most mid-market organisations asking about a data lake want something specific: their systems talking to each other, a durable place to keep history, and somebody to decide who owns which number. That is work we do.

Frequently asked questions

A data lake can hold raw history, mixed files and curated datasets. A warehouse organises data around analytical queries and reporting. Modern platforms overlap, and both need quality checks, ownership and access controls. Choose based on the workload and operating model; using both is an option rather than a default requirement.

If the problem is one missing flow between systems, assess an integration first. A lake becomes more useful when you need retained history, replayable source data or varied inputs for several analytical uses. Agree the business question before choosing storage.

Total cost depends on storage, ingestion, transformation, query workload, access controls and ongoing support. Model expected data volumes and reporting frequency using current AWS prices, then add implementation and maintenance effort. Cheap storage alone does not establish a low-cost solution.

A lake can hold personal data, but its design must reflect the applicable data-protection requirements. Set a clear purpose, appropriate access controls and a justified retention policy, and arrange review and deletion. Use the ICO's current guidance when determining the requirements for your organisation.

The time depends on source access, data quality, permissions and the question the first release must answer. A scoped estimate should separate discovery and a first working pipeline from later sources and reporting features. Agree an acceptance test and operating owner before expanding.

Scorchsoft builds applications, integrations and data foundations, including the AWS storage, catalogue and query layer around them. Its data engineering service connects those pieces to business reporting needs, and Lake On Rails supports data ownership and governance decisions. The right scope can be agreed with your in-house team or platform supplier.

Terms used in this guide

Key topics covered

  • What a data lake is and is not
  • Data lake vs data warehouse vs operational database
  • Five signs you are ready
  • Four signs you are not
  • AWS lake building blocks
  • What drives the cost
  • Data ownership and UK GDPR retention
  • Starting with one question
  • How Scorchsoft can help

Sources referenced

Talk through your data lake options

Tell us what you are trying to answer and which systems hold the data. We will tell you whether a lake is the right answer, and what a first step would look like.

Book a Free Strategy Call
Andrew Ward

About the author

Andrew Ward

Managing Director

Andrew Ward is the founder and Managing Director of Scorchsoft and author of The Control Standard, Execute Your Tech Idea and The ChatGPT Guide for Business. With more than sixteen years of experience building software and running a business, he writes about practical ways to apply technology, use AI and lead teams that deliver.

Andrew holds a first-class degree in Computer Science with Business Management from the University of Birmingham and has represented Great Britain in bench press, winning world championship bronze in 2023.

Want to talk about your project?

Tell us what you’re trying to achieve and we’ll map the fastest credible path.