We understand and sympathise with your frustration. It’s certainly true that you can experience a delay between a data refresh on your retailer customer portal and that same information in SKUtrak. We work hard to minimise this but it is inevitable to some degree – let’s explain why…
UK grocery retailer systems weren’t built for mass data extraction
At the time of writing (summer 2024) no grocery retailer data system (“portal”) is designed for fast, efficient data supply at scale. All of the retailer portals are, at their heart, user-oriented reporting systems and are compromised in different ways when it comes to serving data at scale, in near real-time.
Each day SKUtrak plans and prepares thousands of reports that it will run against the different grocery portals. Over the years, the Atheon team behind SKUtrak has amassed a huge body of operating data on the portals: the reports available, typical run times, failure rates, service downtime, data delays, data errors etc. SKUtrak uses all of this information to plan an intelligent schedule optimised for the various system constraints. Our goals are to:
- Protect all SKUtrak customers from portal downtime, data errors and potential account lockouts;
- Respect the integrity and reliability of the portal systems ie. don’t overload them or risk increasing their failure rate;
- Treat all SKUtrak Premium customers the same – no favourites; and,
- Capture data for SKUtrak Free on a reasonable efforts basis but at a lower priority than SKUtrak Premium
SKUtrak starts collecting data as soon as possible; depending on the retailer portal (it’s typical timing, reliability and any recent issues) this can be as early as 05:30. Throughout the morning, we monitor the availability of retailer portals, assessing the accuracy and completeness of data retrieved, and SKUtrak is constantly adapting its data collection schedule based on the performance we experience.
If portal X is behaving erratically, we will reduce the data collection load and monitor the results achieved. As results improve, data collection frequency and load will be increased. This is a real-time adaptation that occurs every 5 minutes against every data source, every day.
There will be times when you, or any other supplier with portal access, can get a specific report from a portal before SKUtrak collects it…. but SKUtrak keeps going 24/7, always gathering daily data, even when portals are down for days. Even when reports run early in the morning complete and download but are missing data or reporting incorrect values. SKUtrak goes further and collects 3-7 days at a time – because retailer portals often restate recent history e.g. yesterday’s stock positions, or sales from 3 days ago. (These restatements are something we have written about before – Why EPOS data is always wrong)
There’s more to collecting the data… than just collecting the data
All data collected by SKUtrak is processed to meet the stringent needs of true demand intelligence. Every report collected is subject to the same process:
- Check the contents
- Identify and re-collect partial files, corrupted downloads etc.
- Check values against tolerances and tag with warnings (continue processing) or errors (halt processing and recollect)
- Load the data into the SKUtrak catalogue; creates a permanent record of every successful download, for audit and diagnostic purposes
- Transform raw data into SKUtrak DSR format; creates a unified dataset across ALL retailer data sources, ready for reporting and analysis with SKUtrak Insight and/or live query through SKUtrak Share
- Load and aggregate transformed data; loads multiple days of data, applying retailer source data changes to the recent past and recalculating all aggregates e.g. weekly summaries
- Cross-match all product and location codes to SKUtrak Data Management reference data, ensuring that all downstream analysis can be performed in retailer or supplier coding and classification
- Release validated data to SKUtrak Share and SKUtrak Insight
- Run SKUtrak KPI alerts and fire notification emails
Stages 1 to 6 tend to take between 10 and 15 minutes to complete but can run for longer if stage 1 detects problems that require automated recollection of report data. In such cases, it’s typically 20-60 minutes between original report collection and data available within SKUtrak Share and SKUtrak Insight.
Some retailer portals experience repeated issues, including:
- Delays in processing retailer data; overnight batch processes fail and may need to be re-run the following morning
- Large-scale errors in data for short periods e.g. zero sales reported between 08:00 and 09:30
- Restatement of data from the previous day(s)
- One or more reports not loading correctly i.e. partial data collection possible but not the full demand intelligence stack (e.g. stock files, ranging/ distribution etc.)
- In rare cases, limited or no access for hours/ days at a time
What are we doing to address this?
Firstly, SKUtrak reports ALL source system issues as incidents on the SKUtrak status page. Here you can track SKUtrak performance, and the performance of underlying data sources (retailer portals and SKUtrak Direct sources), over the past 90 days. We also expose full details of all SKUtrak system issues – yes, we do experience issues with our SKUtrak technology from time to time and we believe in full transparency on these. You can subscribe to SKUtrak status notifications by email, if you want to be notified of every system incident.
In addition, we continue to work on optimisations to:
- Our data collection ‘agents’ – increasing their intelligence so that we detect issues earlier, collect data more efficiently and use inter-agent communication to increase/ decrease loads as portals improve/ worsen in performance
- Our scheduling patterns – learning from observed success and failure so that we can start collection as early as possible
- Our end-to-end process – finding ways to reduce the latency of data processing from the moment we retrieve a report to the point where it’s available within SKUtrak
- Our reporting and system visibility – we want to share as much information with you as possible/ is useful, without drowning you in tech speak
Closing comments
Certainty and accuracy come at a cost.
Until retailers choose to make machine-readable data available at-scale, on-demand, it will be necessary for CPGs to collect data via reports on retailer portal platforms. These platforms are limited in the number of concurrent sessions, and the volume of data, they can serve; collecting data this way will always be hostage to capacity, crowding and system reliability.
SKUtrak prioritises completeness and accuracy, ensuring that the data it holds is as close as possible to a perfect record of that shared by the retailer. This, and the responsible approach we take when collecting data – so as not to overload retailer portals – means there will always be a slight lag between retailer portal updates and SKUtrak data availability.
We will continue to reduce the lag over time, but some degree of latency—even a few minutes—is inevitable.
Further reading:
Why EPOS data is *always* wrong
