Kindly fill up the following to try out our sandbox experience. We will get back to you at the earliest.
Data Quality Management: How to Build a System That Works
Data quality management explained: the six dimensions, a scoring formula with thresholds, and how monitors are counted per table or column for pricing.

Key Takeaways
- Data quality management is measurement plus a response. Measuring quality and reporting it is not management. If a failing check does not reach a named owner with a deadline attached, you have a dashboard rather than a system.
- Score against six dimensions and weight by consumption. Accuracy, completeness, consistency, timeliness, validity and uniqueness. Weight the score by what a table feeds, not by how many tables you own.
- Set the threshold before you set the alert. A table feeding a regulatory return needs a different response from an exploratory table, and the policy has to say which is which in advance, in writing.
- A monitor is one check against one table or one column. That is the definition Decube publishes on its pricing page, which makes the count arithmetic rather than a negotiation with a sales team.
- You do not need a monitor per column. Two checks per table for freshness and volume, plus column checks only on the columns that are joined on, aggregated, filtered on or published outside the company.
- Decube bills per user, not per monitor. Starter includes 1,000 monitors and Growth includes 3,000, with anything beyond the cap at 0.59 USD per monitor, so widening coverage does not reprice the contract.
What Data Quality Management Is
Data quality management is the practice of defining what good data means for a specific use, measuring live data against that definition on a schedule, and routing every failure to somebody who is accountable for fixing it. All three parts are load bearing. A definition without measurement is a policy. Measurement without an owner is a dashboard. An owner without a definition is a person guessing.
It is worth separating two terms that get used as if they were one. Data quality is a property of a dataset at a point in time: this table is 98 percent complete today. Data quality management is the system that keeps that property inside an agreed range over time and tells you the moment it stops. Most organizations can produce the first number when asked. Far fewer can say what happened the last time it dropped, who was told, and how long it took to recover.
The reason this matters more each year is that data has stopped being something a small analytics team looks at and started being something that acts. A price is set from it, a claim is assessed against it, a regulatory return is filed from it, and increasingly a model is trained on it. Every one of those consumers inherits whatever was wrong upstream, and none of them will tell you. The bad number simply arrives somewhere it matters, dressed as a fact.
The Six Dimensions of Data Quality and How to Measure Each One
Six dimensions are the common language of this field, and most guides list them. Fewer say how to measure one, which is the only part that turns a dimension into something you can act on. The table below pairs each dimension with the question it answers and with a specific check that produces a number.
| Dimension | The question it answers | A check that produces a number |
|---|---|---|
| Accuracy | Does the value match the real world thing it describes? | Compare a sample against a trusted reference, such as a bank statement, a physical count or a source system of record, and report the percentage that agree. |
| Completeness | Is everything that should be here actually here? | Count nulls and empty strings in the columns your definition marks as required, and count expected rows against a control total from the source. |
| Consistency | Does the same fact agree wherever it appears? | Reconcile the same measure across two systems, such as revenue in the warehouse against revenue in the finance ledger, and report the variance. |
| Timeliness | Is the data recent enough for the decision it feeds? | Measure the gap between the newest row and now, and compare it against the freshness the consumer needs, not against the pipeline schedule. |
| Validity | Does the value conform to the rules of its field? | Test format, range and permitted values: dates that parse, currency codes on the approved list, quantities above zero, identifiers matching their pattern. |
| Uniqueness | Is anything counted twice? | Count duplicate keys, and count near duplicates on the business key rather than the surrogate key, which is where the real duplicates hide. |
You cannot write those checks for a table you have never looked at, which is why the first pass over a new source is always data profiling. Profiling tells you what the data currently looks like, and the definition of good is written against what you find, not against what the documentation claims.
How to Score Data Quality and Where to Set the Threshold
A dimension becomes usable when it becomes a number, and a set of numbers becomes usable when it becomes one number per table. The arithmetic is deliberately simple, because a score nobody can recompute by hand is a score nobody trusts.
- Score one dimension. Take the checks you wrote for that dimension on that table and divide the ones that passed by the total. Six of seven completeness checks passing is 86 percent.
- Score one table. Average the dimension scores for that table. Give a dimension a heavier weight only when you can say why, for example weighting accuracy above uniqueness on a table that feeds a financial report.
- Score a data product. Weight each table by what it feeds rather than treating all tables equally. A staging table nobody reads and the table behind the customer statement should never carry the same weight in the same average.
- Report the trend, not the level. A score of 94 percent means nothing on its own. A score that was 99 percent last week means a great deal, which is why the score has to be stored every run rather than recalculated on demand.
The threshold is where most programs stall, because the honest answer is that it depends on the consumer, and teams read that as permission to skip the decision. It is not. Write the tiers down, assign every table to one, and let the tier decide both the target and the response. The policy below is a defensible starting point that you tighten once you have a few months of trend data.
| Tier | What belongs in it | Target score | What happens when it fails |
|---|---|---|---|
| Tier 1 | Feeds a regulatory return, a customer facing number, a payment, or a model that makes decisions about people. | 99 percent or above | Page the on call owner, hold the downstream publish, and record the incident. Failing quietly is not an option at this tier. |
| Tier 2 | Feeds internal dashboards, planning and reporting that people act on within days. | 95 percent or above | Alert the owning team in their channel with the failing check named, and mark the affected dashboard as suspect until it clears. |
| Tier 3 | Exploratory, experimental or archival data that nobody makes a decision from yet. | No target | Monitor freshness and volume only, so you notice if it silently dies, and raise no alert. |
Two rules keep this honest. Nothing enters tier 1 without a named owner, because a threshold with nobody behind it is decoration. And a table only moves up a tier when a consumer moves it, never because the data team thinks it looks important, which is what stops every table drifting into tier 1 over eighteen months.
Building a Data Quality Management System That Actually Works
The cornerstone of a working data quality management system is a defined framework: clear policies, clear procedures and agreed metrics. The success of it lies not only in implementing that framework but in reviewing and updating it so it stays relevant. The five steps below are the order that works, and skipping the first two is the most common reason a program produces alerts nobody acts on.
- Define the data quality requirements. Decide which data actually matters to the organization, what standard is expected of it, and which measures will be used to judge it. This is where the six dimensions become specific checks against named tables rather than an abstraction.
- Put data governance in place. Governance is what makes the rest enforceable: policies and procedures for managing data, and defined roles and responsibilities so that every tier 1 table has a person attached to it. Without this step the alerts fire into an empty room.
- Implement the data quality controls. Once requirements and ownership exist, put the controls in: validation rules at the point of entry, standards for how data is captured, and cleaning and enrichment where the source cannot be fixed. DATAVERSITY has a useful walkthrough of how to implement a data quality framework if you are starting from nothing.
- Monitor data quality continuously. Run the checks on a schedule against live data rather than auditing a sample once a quarter. Continuous data quality means the gap between a value going wrong and somebody knowing is measured in minutes, not in the time until the next review meeting.
- Move towards continuous improvement. Building the system is not a single event. Review which checks fire most, which fire and are always ignored, and which incidents nobody had a check for at all. The last group is the one that tells you what to build next. There is more on running that loop in our guide to data quality management best practices.
How Monitors Are Counted, and Whether You Need One Per Column
This is the question buyers ask before they ask about anything else, and it is the one most vendor pages avoid. The direct answer is no: you do not need a monitor on every column, and a platform that pushes you towards one monitor per column is selling you volume rather than coverage. A monitor is a single check run against a table or a column, so the number you need is the number of things that can go wrong in a way somebody would care about.
Here is the counting rule. For each table, start with two table level checks, one for freshness and one for volume, because those two catch the failures that break everything downstream at once: the load did not run, or it ran and brought back a fraction of the rows. Then add column checks only on the columns that earn one.
- Columns that are joined on. A null or a duplicate in a join key does not produce an error, it produces a wrong answer that looks right.
- Columns that are aggregated. Anything that gets summed, averaged or counted into a reported figure.
- Columns that are filtered on. Status flags, dates and category fields, because an unexpected value silently drops rows out of every report built on that filter.
- Columns published outside the company. Anything a customer, a partner or a regulator will see, where the cost of being wrong is not an internal conversation.
On a typical warehouse table of forty columns, that rule produces two table checks and a handful of column checks rather than forty. Multiply the result across the tables in your tier 1 and tier 2 products and you have a monitor count you can defend in a budget conversation, arrived at by arithmetic rather than by a guess.
On counting for pricing, Decube publishes both the definition and the allowances, which means you can do the sum before you talk to anyone. Its pricing page states that a monitor is a single data quality or observability check run against a table or column. The plan allowances and the overage rate are below, taken from that page on 4 September 2026.
| Decube plan | Data sources | Monitors included | List price | Annual minimum |
|---|---|---|---|---|
| Starter | Up to 3 | 1,000 | 175 USD per user per month | From 21,000 USD a year, minimum 10 users |
| Growth | Up to 10 | 3,000 | 225 USD per user per month | From 54,000 USD a year, minimum 20 users |
| Enterprise | Unlimited | Unlimited | Custom, volume pricing available | Quoted for large teams |
| Additional monitor add on | Any plan | Beyond the plan cap | 0.59 USD per monitor | No minimum commitment, available on all plans |
The practical consequence is that the monitor count is not what drives the bill. Decube prices per user and bills annually, with additional users beyond the plan minimum charged at the same per user rate, so a team that doubles its coverage from 800 checks to 1,600 does not double its contract. That is the opposite of a consumption model, where every new check is a reason to hesitate, and it is the reason coverage tends to widen rather than stall.
Counting is only half of it. The checks themselves have to be cheap enough to create that nobody rations them, which is what machine learning based monitoring is for: a data observability platform learns the normal pattern for freshness and volume on each table and flags the anomalies automatically, so the arithmetic above applies to the checks you write by hand rather than to the baseline coverage.
Where a Data Catalog and Data Observability Fit
The steps above establish the groundwork, but two capabilities decide whether the system holds together in practice, and both are worth understanding on their own terms.
- Data catalog. A centralized record of an organization's data assets, giving an overview of the sources, the definitions and the relationships between elements. With a complete picture of what exists, teams can find the gaps quickly instead of rediscovering the same table three times a year. Our guide to data catalog concepts covers what belongs in one.
- Data observability. Observability is what lets you see how data moves through your systems, spot issues as they develop and act before the number reaches a dashboard. It is the difference between knowing your quality score and knowing why it moved.
Combining the two is what turns a set of checks into a system. The catalog tells you what exists and who owns it, so a failing check has somewhere to go. Observability tells you what changed, so the owner can act on it rather than starting an investigation. Run either one alone and you get half a loop: an inventory nobody monitors, or alerts nobody can route.
Decube is built around that combination rather than around one half of it. The platform carries the catalog, column level lineage, and quality and observability monitoring in one product, which is why a failing check can be traced back to the column it came from and forward to the reports that depend on it without leaving the tool.
The Other Processes a Data Quality Management System Needs
A working system needs two more things that sit outside the monitoring layer and are routinely underinvested in.
- Data cleaning and enrichment. Cleaning removes the errors, inconsistencies and inaccuracies already in the data. Enrichment adds the information that is missing or improves what is there, usually from a reference source. The two go together: cleaning tells you what you can trust, enrichment fills what is absent, and neither is a substitute for fixing the source when the source can be fixed.
- Training and ownership. Standards hold when the people entering and using data understand why they exist. That means regular sessions rather than an onboarding slide, and a culture where a team treats the quality of the data it produces as part of its job rather than as the data team's problem. Most quality issues are created upstream by people who never see the consequence, and a well trained team fixes more of them than any check ever will.
Do You Need Data Quality Management Services or a Platform?
These are different purchases and they solve different problems, which is worth saying plainly because the same search returns both. A service is people doing the work for a defined period. A platform is software that keeps doing it after they leave. Most teams that get stuck bought one when they needed the other.
| Your situation | What to buy | Why |
|---|---|---|
| Nobody can say what data you hold or what state it is in | A service first | A discovery and profiling engagement produces the inventory and the baseline. Buying a platform before this means configuring monitors against unknown tables. |
| A known backlog of dirty historical data to correct once | A service | Remediation is finite work with an end date. It does not need a subscription. |
| Quality is fine on the day it is fixed and degrades within weeks | A platform | That pattern is the definition of a monitoring gap. No amount of remediation fixes it, because the problem is time rather than the data. |
| You need to evidence quality to an auditor or a regulator on a recurring basis | A platform | Evidence means a dated record of checks that ran and what they found. A consultant report is a point in time, and the next audit needs a new one. |
| A small team with no capacity to run a program | Both, in order | Use a service to design the framework and set the initial thresholds, then run it on a platform so the running cost is software rather than headcount. |
How to Evaluate a Data Quality Management Platform
Evaluation criteria are easy to write and hard to make discriminating, because every vendor answers yes to a feature list. These six are worth asking because the answers genuinely differ, and because each one has a follow up question a demo cannot dodge.
| Criterion | What to ask in the demo | What a good answer looks like |
|---|---|---|
| How monitors are counted and priced | What counts as one monitor, and what happens when we go past the allowance? | A published definition and a published overage rate, so you can size the cost yourself before the call. If the answer is "it depends on your usage", you cannot budget it. |
| Coverage without configuration | What is monitored on day one, before we write a single check? | Freshness, volume and schema monitored automatically on every connected table, with anomalies learned from the data rather than from thresholds you have to guess at. |
| Column level lineage | Show me this reported figure traced back to the source columns it came from. | A lineage graph resolved to the column, produced from the query history rather than from a diagram somebody maintains by hand. |
| Where the alert goes | Who gets told when this check fails, and how is that person decided? | Ownership resolved from the catalog and routed to a channel the team already reads, not an email to a shared inbox. |
| Custom checks in SQL | Our business rule cannot be expressed as a template. Write it here, now. | A custom SQL monitor written and running inside the demo. Business logic that only your organization has is where template only tools stop. |
| Evidence for an audit | Export the last quarter of check results for these tables. | A dated, exportable record of what ran and what it found. If it only lives in a dashboard, it is not evidence. |
The Future of Data Quality: Embrace the Change
Two forces are changing this field now, and it is worth being specific about what each one actually changes rather than noting that things are moving.
Machine learning has changed what a check can be. Historically a monitor was a rule somebody wrote and maintained: this column must not be null, this count must be above this number. Models learn the normal pattern for freshness and volume on a given table and flag departures from it, which catches the failures nobody thought to write a rule for and removes the maintenance burden that used to cap how many tables a team could cover. The rules do not go away, they get reserved for the business logic only your organization knows.
Regulation has changed who has to be able to prove it. Open banking and PSD2 pushed financial data across organizational boundaries, and once a number leaves your building the standard shifts from being right to being demonstrably right on a stated date. That same shift is now arriving through AI regulation, where the obligation is to evidence what data a system used. In practice both of them convert data quality from an internal quality bar into a record you have to be able to produce, which is a different engineering problem and a much better argument for funding one.
The more useful way to think about the direction of travel is that data quality is moving from something a team reviews to something a system proves continuously. If your program cannot yet produce a dated record of what was checked and what it found, that is the gap to close first, ahead of any new capability.
Data Quality Is the Key to Successful Data Management
Everything above reduces to one loop: define what good means, measure it continuously, route the failures to somebody accountable, and review what you learn. Organizations that run that loop can act on their numbers without checking them first, which is the whole point and is much rarer than it sounds.
If you are starting from nothing, the sequence below gets you to a working first version inside a month. It is deliberately narrow, because the programs that fail are the ones that try to cover everything in the first quarter.
| Week | What you do | What you have at the end of it |
|---|---|---|
| Week 1 | Pick the five tables that feed your most important reported numbers, and name an owner for each one. | A tier 1 list short enough to actually finish, with a person against every row. |
| Week 2 | Profile those five tables and write the definition of good for each, dimension by dimension. | A specific, checkable standard rather than an aspiration. |
| Week 3 | Turn each definition into checks: freshness and volume per table, plus column checks on the columns that earn one. | A monitor count you calculated rather than guessed, and a baseline score per table. |
| Week 4 | Set the thresholds, wire the alerts to the owners, and agree what happens on a failure. | A loop that closes without anybody watching a dashboard. |
From there the work is repetition rather than invention: add the next five tables, watch which checks earn their alerts, and retire the ones that only ever produce noise. If you want to see what that looks like running against your own warehouse, you can request a demo and walk through it on your tables rather than on a sample dataset.
Frequently Asked Questions
What is data quality management?
Data quality management is the practice of defining what good data means for a specific use, measuring live data against that definition on a schedule, and routing every failure to a named owner who is accountable for fixing it. It is usually measured across six dimensions: accuracy, completeness, consistency, timeliness, validity and uniqueness. The distinction that matters is between data quality, which is a property of a dataset at a point in time, and data quality management, which is the system that keeps that property inside an agreed range and tells you the moment it stops.
Is data quality monitored with a monitor per column, and how are monitors counted for pricing?
No, you do not need a monitor on every column. The rule that works is two table level checks per table, one for freshness and one for volume, plus column level checks only on the columns that are joined on, aggregated, filtered on or published outside the company. On a forty column table that produces a handful of checks rather than forty. For counting, Decube publishes the definition on its pricing page: a monitor is a single data quality or observability check run against a table or column. The Starter plan includes 1,000 monitors and the Growth plan includes 3,000, Enterprise is unlimited, and anything beyond the plan cap is charged at 0.59 USD per monitor with no minimum commitment. Because Decube prices per user rather than per monitor, at 175 USD per user per month on Starter and 225 USD on Growth, widening your monitor coverage does not reprice the contract.
What is a data quality management system?
A data quality management system is the combination of a framework, the tooling that runs it and the people accountable for it. The framework sets the policies, procedures and metrics. The tooling runs the checks continuously against live data and raises the failures. The people are the named owners each monitored table is assigned to. A system that has the first two but not the third produces alerts nobody acts on, which is the most common way these programs stall.
What is data quality control, and how is it different from data quality management?
Data quality control, sometimes shortened to data QC, is the checking step: testing data against defined rules and flagging what fails. Data quality management is the wider system that decides which rules exist, who owns each dataset, what score is acceptable, what happens when a check fails and how the whole thing improves over time. Control is one activity inside management, and doing control alone is why some teams have thorough testing and still cannot say whether their data is getting better or worse.
What is continuous data quality?
Continuous data quality means the checks run on a schedule against live production data rather than as a periodic audit of a sample. The practical test is how long it takes for somebody to find out after a value goes wrong. Under a periodic model the answer is however long remains until the next review. Under a continuous model it is the interval between check runs, which is typically minutes or hours, and the failure is caught before the number reaches a dashboard or a customer.
What should I look for when evaluating a data quality management platform?
Six criteria separate platforms in practice. Ask how monitors are counted and what the overage rate is, so you can size the cost yourself. Ask what is monitored on day one before you write any checks. Ask to see a reported figure traced back through column level lineage to its source columns. Ask how the platform decides who gets alerted when a check fails. Ask to have one of your own business rules written as a custom SQL monitor during the demo. And ask for an export of a quarter of check results, because if the history only lives in a dashboard it will not serve as audit evidence.
Do I need data quality management services or a data quality platform?
Buy a service when the work is finite: an initial discovery and profiling engagement, or a one off remediation of a known backlog of historical data. Buy a platform when the problem is time rather than data, meaning quality is acceptable on the day it is fixed and degrades within weeks, or when you have to evidence quality to an auditor or a regulator on a recurring basis. Small teams often need both in order: a service to design the framework and set the initial thresholds, then a platform so that running it costs software rather than headcount.
What is customer data quality management?
Customer data quality management is data quality management applied to customer records, where uniqueness and accuracy carry more weight than they do elsewhere. The dominant failure is duplication, because the same customer arrives through several channels and creates several records, so the deduplication check has to run on the business key rather than the surrogate key. Accuracy matters more than usual too, since contact and address fields decay steadily without anything in the pipeline going wrong, which is a case where enrichment against a reference source does more good than tighter validation at entry.














.webp)