
PIM · DATA QUALITY
Measuring Product Data Quality: The Key PIM KPIs & How to Steer Them
“Our product data is actually pretty good.” You hear that sentence in almost every project — and it is worthless. “Actually pretty good” cannot be steered, cannot be prioritised and certainly cannot be argued in front of senior management. As long as data quality remains a gut feeling, it gets cut when budgets get tight.
Hence our thesis: data quality is not a feeling but four measurable KPIs — Completeness, Accuracy, Consistency and Timeliness. Anyone who defines these four metrics cleanly and measures them regularly turns a vague quality promise into a number that can be tracked, improved and reported. And these four KPIs now decide far more than whether you have a pretty product page: they decide whether AI systems can understand, surface and recommend your products at all.
In this article you get a precise definition for each of the four KPIs, a formula you can recalculate, realistic reference values and the concrete levers you use to steer each one in your PIM.
Why data quality has to be measurable
Without a metric, the same thing always happens in data quality projects: people argue about individual cases. Somebody finds a product without an image, somebody else finds a wrong unit of measure — and out of that comes either alarmism or a shrug. Neither leads to a decision.
Measurable KPIs solve exactly three problems:
- Prioritisation. If you know that the “Tools” category sits at 62% completeness and “Consumables” at 94%, you know where your team goes next. Without a number, the loudest opinion wins.
- Progress. Data maintenance is invisible work. A KPI makes it visible — and with it the effect of any investment in PIM, processes or people.
- Commitment. A target value (“95% Completeness for all shop-ready items by Q4”) is a promise you can work towards. “We’ll make the data better” is not.
The prerequisite for all of this is a system in which product data actually sits in a structured, central place. Why data is the real centrepiece of every PIM system — and not the software around it — is something we described in Data: the centrepiece of every PIM system. In an Excel landscape, these KPIs simply cannot be captured in any serious way.
The 4 core KPIs of product data quality — with formulas
The following four metrics are the core: each is defined as a percentage, each can be calculated and steered automatically in the PIM, and together the four give you a robust overall picture. Alongside them, two further dimensions play a role — Validity (format conformity), which we deliberately treat here under Accuracy because in the PIM it is checked by the same set of rules, and Uniqueness (freedom from duplicates), which is sensibly captured separately per data set: it is more of a clean-up KPI than an ongoing steering KPI, particularly after assortment acquisitions and supplier imports.
One important point up front: all four KPIs always relate to a defined scope — that is, to a set of products (e.g. all shop-ready items) and a set of attributes (e.g. all mandatory fields for the “online shop” channel). Without that scope, every number is arbitrary.
1. Completeness
Definition: Completeness measures the share of mandatory fields actually filled in, out of all mandatory fields required for a channel.
Formula:
Completeness (%) = (filled mandatory fields / required mandatory fields) × 100
Example: Your online shop requires 20 mandatory attributes per item. With 1,000 items, that makes 20,000 required fields. 17,400 are filled. → Completeness = 17,400 / 20,000 × 100 = 87%.
What to watch out for: “Filled” does not mean “meaningfully filled”. A field containing “tbd”, “-” or “see data sheet” counts as filled technically, but not in substance. So define minimum criteria (e.g. description ≥ 200 characters, no placeholder text) and factor them into the count. A second view is also worthwhile: item-level Completeness — that is, the share of items that are 100% complete. It is typically well below the field-level figure and is more honest, because an item with 19 out of 20 fields still cannot be published.
2. Accuracy
Definition: Accuracy measures the share of attribute values that are factually correct and comply with the defined rules — in other words, values that pass validation against a reference or a rule set.
Formula:
Accuracy (%) = (correct attribute values / checked attribute values) × 100
Example: Out of 50,000 checked values, 1,150 fail validation (wrong unit, implausible value, invalid EAN check digit). → Accuracy = 48,850 / 50,000 × 100 = 97.7%.
What to watch out for: Accuracy is the only one of the four KPIs that needs an external truth — a reference to check against. Consistency only compares sources with one another and establishes that they diverge, not which one is right. Only Accuracy answers that. This reference can be automated (check-digit logic for GTIN/EAN, value ranges, permitted value lists, unit plausibility) or manual on a sample basis (comparison against a data sheet or manufacturer specification). Keep the two cleanly separated: automated rule violations are fully measurable, whereas the factual correctness of free text can only be assessed on a sample basis. In your reporting, always state which share was checked automatically and which by sampling.
3. Consistency
Definition: Consistency measures the share of records that are free of contradictions across all systems, channels and languages — same meaning, same spelling, same structure.
Formula:
Consistency (%) = (contradiction-free records / checked records) × 100
Example: Out of 1,000 items, 60 differ between PIM and shop (deviating unit of measure, different colour designation, missing translation of an attribute value). → Consistency = 940 / 1,000 × 100 = 94%.
What to watch out for: Consistency has three dimensions that you should measure separately, because they have different causes:
- System consistency: Do PIM, shop, marketplace and catalogue match?
- Language consistency: Are all language variants at the same status and the same version?
- Semantic consistency: Does “colour: anthracite” mean the same thing everywhere, or does it appear elsewhere as “dark grey”, “RAL 7016” and “Anthrazit”?
The third point is the underrated one. It is also the one AI systems stumble over — more on that below. Cleanly maintained value lists and a consistent classification of product data are the most effective lever here: where values come from a controlled list, semantic inconsistency cannot arise in the first place.
4. Timeliness
Definition: Timeliness measures the share of records that were updated or published within a defined time window — in other words, that are not out of date.
Formula:
Timeliness (%) = (records within the timeliness window / all records) × 100
Example: Your timeliness window for price and availability data is 24 hours. Out of 1,000 items, 970 were updated within that window. → Timeliness = 970 / 1,000 × 100 = 97%.
What to watch out for: The timeliness window is not a universal figure; it depends on the type of data. Prices and stock levels need hours, marketing copy and images more like months, regulatory information often a fixed review interval. So define the window per attribute group and calculate Timeliness separately for each. A second useful metric alongside it is time-to-market: the time from item creation to channel-ready release. It does not measure the data itself but the process speed behind it — and often explains why Timeliness collapses.
Target values & benchmarks — to be read as reference values
Now the question that always comes up: which value is good enough?
An important qualification up front: the values below are experience-based reference values from product data projects, not universally valid study results. They depend heavily on industry, assortment depth, channel mix and maturity. Use them as a starting point for defining your own targets — not as a standard you have to measure yourself against.
| KPI | Critical | Solid | Target corridor | Typical trigger for problems |
|---|---|---|---|---|
| Completeness (field level, mandatory fields) | < 80% | 90–95% | ≥ 95% | New channels with additional mandatory fields; assortment acquisitions |
| Completeness (item level, 100% complete) | < 60% | 75–85% | ≥ 90% | Individual permanently empty fields across the whole assortment |
| Accuracy (automatically checkable) | < 95% | 97–99% | ≥ 99% | Supplier data without inbound validation |
| Consistency (system & language comparison) | < 90% | 93–97% | ≥ 98% | Manual maintenance in parallel in the shop; free text instead of value lists |
| Timeliness (price/stock, 24-hour window) | < 90% | 95–98% | ≥ 99% | Batch runs instead of event triggers; interface errors without monitoring |
| Timeliness (content, 12-month window) | < 70% | 80–90% | ≥ 90% | No review cycle for existing items |
Two notes on reading this table. First, 100% is not a sensible target — the last percentage points cost disproportionately much, and part of them lies outside your control (supplier data, translation runs). Second, the development over time is more meaningful than the absolute value. An assortment that has climbed from 71% to 88% Completeness is in better shape than one that has been stuck at 91% for two years and is slowly slipping.
So define your target values in a differentiated way: by channel (marketplaces demand more than your own shop), by assortment relevance (A-items stricter than slow movers) and by data type.
How to steer the KPIs in your PIM
Measuring alone improves nothing. The real value emerges where the KPIs are coupled to concrete mechanisms in the PIM. Four levers cover the bulk of it:
1. Mandatory field definitions per channel. The starting point for Completeness. Define per output channel which attributes are compulsory — the shop needs different ones from Amazon, Amazon different ones from the print catalogue. Only this makes completeness calculable at all, because you then have a defined denominator. A good PIM shows this value live as a progress indicator per item and per channel — so the editorial team sees immediately, while maintaining data, what is still missing.
2. Validation rules and plausibility checks. The lever for Accuracy. Value ranges, unit checks, mandatory formats, check-digit logic, dependencies between fields (“if category X, then attribute Y is mandatory”). What matters is that these rules take effect at import and at data entry, not only in a downstream report — because by then you are correcting errors that have already been published. How automatic plausibility checks are built in practice is something we show in PIM validation: automatic plausibility checks.
3. Controlled value lists and classification. The lever for Consistency. Everywhere an attribute value comes from a maintained list instead of a free text field, an entire class of errors disappears. The same applies to a clean classification structure: it ensures that similar products also get similar attribute sets — the prerequisite for completeness and consistency being comparable at all.
4. Workflows, approvals and ownership. The lever for Timeliness. An item sitting in the system without an approval step and without a responsible role ages unnoticed. Concretely: a status model (draft → in enrichment → approved → published), assigned owners per attribute group, reminders when the timeliness window is exceeded, and a fixed review cycle for existing items.
On top of that comes the organisational frame: a KPI nobody looks at regularly is decoration. Set a fixed reporting interval (monthly is usually enough), show the values per category and per responsible team, and treat deviations like any other metric — with an action, not a discussion.
The bridge to GEO and agent readiness
Until recently, data quality was above all an argument about conversion, returns and service costs. That still holds — but it has gained a new and considerably harder dimension.
When an AI system — a language model in a product search, a shopping agent, a generative answer engine — processes your products, it does exactly what a human does not: it reads the data structurally and literally. It does not interpret generously, it does not fill gaps from experience, and it does not know that for you “anthracite” and “dark grey” mean the same thing. For a system like that, the following applies very directly:
- Missing attributes means: the product cannot be selected in an attribute-based query. If you do not state the installation height, you do not show up for “fits a 60 cm recess” — not ranked lower, but absent entirely.
- Incorrect values means: the system recommends wrongly — and that comes back to you later as a return or a complaint.
- Inconsistent values means: the system cannot reliably compare or group products and, in case of doubt, leaves them out.
- Out-of-date data means: recommendations run on prices or availabilities that no longer exist.
Put differently: the four KPIs are simultaneously the metrics for your AI visibility. Completeness determines whether you get into the selection at all. Accuracy and Consistency determine whether you are represented correctly and comparably. Timeliness determines whether the recommendation still holds. What this means concretely for preparing your data is something we go into in agent-ready product data — that article is about visibility itself, this one about measuring and steering it.
The practical benefit of this connection: you do not need a separate initiative or a second metrics system for “AI readiness”. If you have your four data quality KPIs under control, you are already working on your discoverability in generative systems. It is the same work — just with an additional, very topical argument to put in front of senior management.
Conclusion: four numbers instead of a gut feeling
Product data quality becomes manageable as soon as it has a number. Completeness, Accuracy, Consistency and Timeliness together form a complete and practicable measurement system: clearly defined, automatically calculable in the PIM, coupled to concrete steering levers — and at the same time the best indicator of whether your products appear in AI-supported channels at all.
Start small: one channel, one defined set of mandatory fields, one monthly value. The first percentage you measure is almost always sobering — and that is exactly why it is the most valuable point in the whole project.
FAQ: measuring product data quality
What is product data quality?
Product data quality describes how well product data fulfils its purpose — that is, whether it is complete, correct, consistent and current enough to be published without errors across all sales channels. It is measured via four core KPIs: Completeness, Accuracy, Consistency and Timeliness.
How do you measure the data quality of product data?
Via four percentage metrics, each relating to a defined scope of products and attributes: Completeness = filled mandatory fields / required mandatory fields × 100. Accuracy = correct attribute values / checked attribute values × 100. Consistency = contradiction-free records / checked records × 100. Timeliness = records within the timeliness window / all records × 100. A PIM system calculates these values automatically per channel and category.
Which KPIs are there for product data quality?
The four core KPIs are Completeness, Accuracy, Consistency and Timeliness. Additionally useful are item-level Completeness (the share of items that are 100% complete) and time-to-market (the time from item creation to channel-ready release), which makes visible the process speed behind timeliness.
How do you calculate the completeness of product data?
Completeness (%) = (filled mandatory fields / required mandatory fields) × 100. Example: 20 mandatory attributes across 1,000 items make 20,000 required fields; if 17,400 are filled, Completeness is 87%. It is important not to count placeholders such as “tbd” or “-” as filled, and to look at item-level completeness as well.
Which target value is good for product data quality?
As experience-based reference values — not a universally valid standard and not the result of a study: Completeness from around 95% at field level, Accuracy from around 99% for automatically checkable rules, Consistency from around 98%, and Timeliness from around 99% for price and stock data within a 24-hour window. That time frame is an assumption, not an industry norm; the same applies to the 12-month window for content. All these values depend on industry and assortment, and the development over time is more meaningful than the absolute figure.
What does data quality have to do with AI visibility?
AI systems read product data structurally and literally: missing attributes exclude a product from attribute-based queries entirely, inconsistent values prevent comparability, incorrect values lead to wrong recommendations, and out-of-date data leads to invalid offers. This makes Completeness, Accuracy, Consistency and Timeliness the metrics for visibility in AI-supported search and shopping systems as well.
How do you steer data quality KPIs in a PIM?
Via four mechanisms: channel-specific mandatory field definitions (steers Completeness), validation and plausibility rules at import and at data entry (Accuracy), controlled value lists and a clean classification (Consistency), and workflows with a status model, ownership and review cycles (Timeliness). This is complemented by a fixed reporting interval per category and per responsible team.
Want to see how these KPIs are calculated and steered live in a PIM?
In a personal demo we’ll show you how mandatory field logic, validation and workflows interact in Online Media Net (OMN).