Data Lakes Are Graveyards: From Storage to Signal
There is a well-known saying in the data world: a data lake is where data goes to die. The phrase has the ring of a joke, and like most good jokes, it is a precise description of reality. Companies spend fortunes building vast storage systems, fill them with everything they have, and then discover that the lake has become a swamp: unmanaged, unsearchable, and useless for the decisions it was supposed to inform.
The tragedy is not the technology. The technology works. The tragedy is the assumption behind it: that storing data is the same as having data. It is not. Storage is the cheapest part of the data lifecycle. Signal, the ability to turn raw data into a decision, is the expensive part, and it is where almost everyone underinvests.
1. The Collect-Everything Fallacy
The promise of the data lake was seductive: store everything now, sort it out later. The promise was made in an era when storage was expensive and analysis was cheap. The reality is the opposite: storage is nearly free, and analysis is where the cost lives. Collecting everything without a plan is not strategy, it is hoarding.
The cost of hoarding is hidden but real. Every terabyte of junk makes the useful data harder to find. Every duplicate schema confuses the analysts. Every forgotten dataset is a future lawsuit or a future embarrassment. The discipline is the inverse of the fashion: collect what you need, know why you need it, and have a plan for what it becomes. The lake that is curated beats the lake that is vast.
2. The Schema of the Future
The deepest flaw of the lake is the promise that "we will figure out the structure later". Later never comes, because retrofitting structure onto years of unstructured chaos is more expensive than building it upfront. The data that arrives with a schema, a clear definition of what each field means, is data that can be used. The data that arrives without one is a liability.
The practical rule is simple: every dataset entering the lake carries its schema, its owner, and its freshness. Not a complex governance regime, just the basics: what is this, who maintains it, how old is it allowed to get. The basics are what turn a dump into an asset. Without them, every future analysis starts with archaeology, and archaeology is not analysis.
3. The Three Kinds of Data
The lake becomes manageable when you sort it into three kinds. The first is source data: the raw events, logs, and transactions as they happened, immutable and untouched, the ground truth of the business. The second is refined data: cleaned, joined, and shaped for analysis, produced by known pipelines on a schedule. The third is derived data: metrics, models, and reports, the outputs that people actually look at.
The three kinds have different rules. Source data is protected and never edited. Refined data has owners and quality bars. Derived data has consumers and freshness expectations. The sorting is not bureaucracy, it is hygiene. The teams that know which kind of data they are touching make far fewer mistakes than the teams swimming in an undifferentiated sea.
4. The Pipeline Is the Product
The data lake is not the product. The pipeline is: the chain of transformations that turns raw events into a trustworthy number. The number is only as good as the weakest link in the chain, and most chains are held together by undocumented scripts that nobody fully understands and nobody dares to touch.
The professional approach treats pipelines like software: versioned, tested, documented, and owned. A pipeline with tests that run, a schedule that is monitored, and an owner who is accountable, is an asset. A pipeline that is a mystery is a time bomb. The investment in pipeline quality is the highest-return investment in the entire data stack, because every downstream consumer pays for pipeline quality every single day.
5. Data Quality Is a Feature
Every data team has the same conversation, in the same order. First: we need more data. Then: we have the data, but it is wrong. Then: we fixed it, but now it is slow. The first two phases are a trap: more data with bad quality just means more wrong answers with higher confidence.
The fix is to treat quality as a feature with explicit requirements, not a hope. Every critical dataset has a quality bar: how complete, how timely, how accurate it must be, measured continuously, with alerts when it slips. The bar is not a one-time validation, it is an ongoing contract. The teams that treat quality as a feature spend less time firefighting and more time analysing, which is the whole point of having data at all.
6. The Metrics Layer
The final layer of the stack is where data becomes decisions: the metrics that the business actually runs on. The metrics layer is where the biggest messes live, because everyone defines everything differently. Revenue, churn, active user, each means something different in each report, and nobody notices until two reports disagree loudly in a meeting.
The metrics layer imposes a single definition for every important number, stored once, computed one way, documented plainly. The definition is the contract that ends the arguments. The teams that invest in a metrics layer do not eliminate disagreement, but they move it to where it belongs: about what the number should be, not about what the number means.
7. The Cost of Not Knowing
The graveyard metaphor is not about storage, it is about opportunity. Every decision made without data is a guess, and the guesses compound. The company that cannot answer "how much does it cost us to serve this customer" or "which features actually keep people" is flying blind, and the blindness is a choice, not a fate.
The cost of not knowing is invisible, which is why it is so easy to ignore. There is no line item for the decision that would have been better with data. But the compounding effect is the difference between companies that iterate and companies that gamble. The data function is not a cost centre, it is the difference between informed and uninformed bets, and the market pays for that difference.
8. From Lake to Signal
The journey from lake to signal is not a technology project, it is a cultural one. It starts with the decision that data is a product with users, owners, and quality bars. It continues with the boring work of schemas, pipelines, and metrics. It ends with the organisation where a question can be answered in minutes, and the answer can be trusted.
The companies that make this journey do not have better technology than their competitors. They have better hygiene. They treat their data the way they treat their money: with ownership, with discipline, and with a clear idea of what it is for. The lake is not the destination. The signal is. And the signal is built, not collected.
Tags
#technology #business
Comments
No comments yet. Be the first!
Leave a comment