Skip to content
Getting Digital

Data Quality

Also: data quality dimensions, data accuracy, data completeness, data validation, data cleansing, dirty data

Data quality is the degree to which a dataset is fit for the purpose it is used for, assessed along named dimensions, whether the records are complete, unique, consistent, timely, valid and accurate, and measured against the requirements of the decision the data supports rather than against an absolute standard.

Assessment. Quality is relative to the use, and most quality programmes fail by forgetting it. A customer table that is adequate for a marketing mailing is inadequate for a regulatory return, and an attempt to make every dataset perfect for every purpose exhausts the budget on fields that are not used. The practical method is to state the purpose, choose the dimensions that purpose depends on, measure them at the source and publish the measure beside the data.

The UK government's data quality framework, written for public-sector statisticians and applicable to any organisation, defines quality as fitness for purpose and names six dimensions along which it is assessed. Completeness: the required records and values are present. Uniqueness: each entity appears once. Consistency: values describing the same entity do not contradict each other across datasets. Timeliness: the data represents the period it claims to and arrives when it is needed. Validity: values are in the correct format and within the expected range. Accuracy: the data describes the real-world entity correctly. A dataset can score well on five and fail on the sixth; a complete, valid, consistent list of addresses that is two years out of date fails on timeliness and accuracy for a delivery service and passes for a historical analysis.

DimensionQuestion it answersTypical check
CompletenessIs anything missing?Share of required fields populated; expected row count against the source
UniquenessIs anything duplicated?Duplicate keys; near-duplicate names and addresses
ConsistencyDo the sources agree?The same customer's status in the CRM and the billing system
TimelinessIs it current enough?Age of the newest record; arrival time against the deadline
ValidityIs the format right?Dates that parse; postcodes that match the pattern; values within range
AccuracyIs it true?Sample compared with the source document or the real world

Where the checks run determines what they cost. A check at the point of entry, a form that refuses an invalid postcode, prevents the defect; a check in the pipeline catches it before the warehouse; a check in the report finds it after a decision was taken on it. The cost of a defect rises at each stage, which is why the engineering preference is for validation at the boundary and quarantine of the rows that fail, with the count of quarantined rows published as the quality measure. Analysts still meet the defects that passed, and data wrangling is in large part the work of handling them: the Power BI analyst exam lists resolving inconsistencies, unexpected or null values and data quality issues as a tested skill in its data-preparation domain.

Publishing the measure

A dataset with a known completeness of 94 per cent and a documented reason for the gap is usable; the same dataset with an unknown completeness is not. Quality measures belong beside the data in the catalogue, with the date they were taken, so that every reader can judge fitness for their own purpose.

In practice

  • Purpose: a monthly statutory return of customer counts by region.
  • Dimensions that matter: uniqueness (a customer with two accounts is counted once), consistency (region agrees between the CRM and billing), timeliness (the month's closing position, not the mid-month one).
  • Dimensions that do not: accuracy of the telephone number, completeness of the marketing-consent field.
  • Measure: duplicate rate below an agreed threshold and region agreement above one, both reported with the return; the thresholds are set by the regulator's tolerance, not by an internal ideal.

Often confused with

Data Wrangling
Data wrangling is the work an analyst does to reshape and repair one dataset for one analysis; data quality is the measured condition of the data and the programme that keeps it acceptable at the source. Wrangling treats the symptom in one place; a quality programme treats the cause for everyone.
Data Governance
Governance decides which dimensions are measured, who owns the result and what happens when it falls; quality is the measured result. A governance register without measures is a list of owners; measures without owners are reports nobody acts on.

Key takeaways

  • →Fitness for purpose, assessed on six dimensions: completeness, uniqueness, consistency, timeliness, validity, accuracy.
  • →Choose the dimensions the decision depends on; measure them at the source; publish the measure beside the data.
  • →The cost of a defect rises at every stage it passes; validate at the boundary and quarantine what fails.

Related concepts

  • Governance decides what is measured and who owns the result; quality is the measured result.

  • Validation at the pipeline boundary is where most defects are caught.

  • Wrangling handles the defects that passed the checks, one analysis at a time.

Where this concept sits in the field

Certifications that test this

Vendor exams whose syllabus covers this concept: facts, cost and a preparation path on each page.

More courses from these categories

Courses from the categories where this concept is taught. Details, price and the provider link are on each course page.

FAQ

How much data quality is enough?
Enough for the decision the data supports, stated as a threshold on each relevant dimension. A marketing list tolerates a few per cent of outdated addresses; a payment file tolerates none. The threshold comes from the consequence of the error, which only the owner of the decision can state.
Can data quality be automated?
The measurement can, and the modern pipeline tools run validity, completeness and uniqueness checks on every load. Accuracy cannot be fully automated, because it requires comparison with reality; it is checked by sampling against source documents or by reconciling totals with an independent system.
Who is responsible for data quality?
The owner of the dataset, for the standard; the team that enters or generates the data, for meeting it; the engineering team, for measuring and reporting it. Programmes that place the whole responsibility on the analysts who consume the data fix the same defects every month.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 3 October 2026 · Getting Digital