Corporate Business Alliance

Technology · 24 September 2026

Data quality is decided upstream

Analysts spend much of their time cleaning data that was entered wrongly at source, and they clean the same errors every month. Lasting data quality comes from fixing capture, ownership and definitions where data is created.

In many organisations, data quality is treated as a job for the analytics team. Reports arrive with duplicates, missing values, inconsistent codes and dates in three formats, and analysts clean them — by hand, in spreadsheets, before every run. The work is repeated each period because nothing upstream has changed. The cleaned data is only as good as the latest analyst's rules, and two analysts cleaning the same data will often produce different results.

The underlying problem is that the quality of data is decided where the data is created, and the people who create it are rarely the people who suffer from its errors.

Define the data before measuring it

Many apparent quality problems are definition problems. "Customer" may mean an account in the billing system, a legal entity in the credit system and a contact in the sales system. "Revenue" may be recognised, invoiced or received. Two reports built on different definitions will disagree, and both may be correct.

A short, agreed set of definitions for the terms that matter most — owned by the business, not by IT — resolves more disputes than any amount of cleaning. Each definition should say what the term means, where the authoritative value is held and who is responsible for it.

Give data an owner

Data without an owner degrades. Assign each important dataset to an owner in the part of the business that creates it, with responsibility for its accuracy and completeness. The owner does not have to fix every error personally; they have to ensure the processes that create the data are designed to get it right, and to act when the measures show they are not.

Validate at the point of entry

The cheapest place to prevent an error is at the moment it is entered. Systems and forms can enforce much of this directly:

  • mandatory fields for the information that downstream processes rely on;
  • selections from controlled lists instead of free text for codes, categories and countries;
  • format and range checks on dates, amounts and identifiers;
  • duplicate checks before a new customer, supplier or product is created.

Each of these removes a class of error permanently, rather than requiring it to be found and corrected every month.

Measure quality and show it

What is not measured is rarely improved. A small set of data quality measures — completeness of key fields, the rate of duplicates, the number of records failing validation, the time taken to correct them — reported to the data owners makes the state of the data visible to the people who can change it.

Keep a record of what cleaning does

Some cleaning will always be needed downstream. When it is, it should be done by documented, repeatable rules rather than by hand, so that the result is the same each time and anyone can see what was changed. Where the same correction is made every period, that is a signal to fix the source.

Better analysis starts before analysis. Organisations that invest in definitions, ownership and validation at the point of capture spend less time cleaning and more time finding things out.

All insights