SKUtrak

Why does a demand signal repository cost more than a data lake?

Digital cube and water surface background, 3d rendering. Digital drawing.

Pound for pound, terabyte for terabyte, it’s true that a demand signal repository (DSR) will cost more than a data lake to implement – because a DSR is a specialised form of data lake (or sometimes data warehouse) that adds structure, semantics and meaning to data such that downstream activities are vastly simplified. To understand the differences and the cost implications, let us:

  1. Outline the key features of a data lake;
  2. Introduce the additional features required for a DSR; and,
  3. Explore some of the benefits of a DSR, over and above a general-purpose data lake.


What are the key features of a data lake?

A data lake is a highly versatile, general-purpose store for any type of data an organisation wants to collect, manage, and utilise—potentially in many different ways. Data lakes are ideal for storing data in its original form, i.e., the way you collect or receive it, no matter the source or type. From CSVs to video content, a data lake can handle it all.

A few alternative definitions, from some of those who would know, might help here:

  1. Amazon Web Services
    “A data lake is a centralised repository that allows you to store all your structured and unstructured data at any scale.”
  2. Databricks
    “A data lake is a central location that holds a large amount of data in its native, raw format.”
  3. Google
    “A data lake is a centralised repository designed to store, process, and secure large amounts of structured, semi-structured, and unstructured data.”
  4. Microsoft
    “A data lake is a centralised repository that ingests and stores large volumes of data in its original form.”
  5. Snowflake
    “A data lake is essentially a highly scalable storage repository that holds large volumes of raw data in its native format until needed for various purposes.”


So here, at least, the major data platform providers are in complete agreement; data lakes deal with raw data, in its original form, at an unlimited scale, and thus suit all forms of structured and unstructured data. Data lakes emerged during the hype cycle for “Big Data” in the 2010s and are now firmly established as the standard way for organisations of all sizes to ingest and store raw data from all sources. Many enterprises have implemented at least one, some more, waves of data lake technology and all major cloud data platforms provide data lake capabilities within their offer.


How does a DSR differ from a data lake?

A DSR is different. Whilst a data lake is a general-purpose data store, a demand signal repository is – as its name suggests – designed to deal with a subset of an organisation’s data; demand signals.

What is a demand signal?

A demand signal is any data that indicates demand for products or services and encompasses a large set of measures that show both met and unmet demand. Understanding the principles of “met and unmet demand” is important because true demand can never be measured directly, only inferred from met (or achieved) demand plus an estimate of unmet (missed) demand. Let’s explore this with an example:

Your local bakery makes 100 croissants in the morning before the store opens and by the end of the day has only sold 80, it knows shopper demand that day was precisely 80 croissants. If the following day, however, the bakery makes 80 – in line with demand on the previous day – and by 10.30 am has sold all of them, then it has established demand for at least 80 croissants, but perhaps there was some unmet demand? If the bakery were to record requests from every shopper who couldn’t find a croissant after 10.30, it would have an additional measure of unmet demand, but the bakery wouldn’t know about demand from those shoppers who didn’t ask, despite wanting a croissant.


In this example, the record of shopper requests is a powerful demand signal; it increases information about the true demand for croissants on the second day, although it’s still imperfect as the bakery can assume that some shoppers didn’t ask and some may have bought more than one.

A DSR is a repository (data store) for demand signals (information about demand), so it’s more useful than a general-purpose data lake for demand intelligence. However, it is unable to deal with accounting, personnel, or legal datasets (except for those parts that relate specifically to demand).


What are the pros and cons of a DSR compared to a data lake?

Think of a DSR as a highly-tuned ‘special edition’ of a data lake; let’s try a few analogies:

  • Family SUV vs race-car
    A family SUV/4WD will get multiple people from A to B, over many types of terrain, in any conditions, with minimal fuss and relative comfort… but it won’t win any races. A race car is optimised for specific conditions, particular terrain, little or no comfort and an absolute focus on competing and winning.
  • Handyman vs plasterer
    A general handyman can paint a wall, tidy a garden or fix a kitchen sink – very useful, general-purpose jobs around the home. If, however, you ask your handyman to plaster a wall, you’re likely to wish you paid the premium for a specialist plasterer who will provide a high-quality finish the first time, every time… just don’t ask her to fix the dishwasher!
  • Decathlete vs sprinter
    Decathletes are extraordinary—competing in ten events at every athletic meet and performing at a very high level in all of them. If one ran in the 100m final of the Olympics, it’s certain they would come last… but ask a 100m winner to pole vault or throw a javelin, and you will quickly realise the limit of their talents.


In these examples, the data lake is the family SUV, handyman, or decathlete, and the DSR is the race car, plasterer, or sprinter. The DSR is the specialist, and the data lake is the generalist.

Exit mobile version