Kindly fill up the following to try out our sandbox experience. We will get back to you at the earliest.
Data Quality Monitoring: Metrics, Monitors and Alerts
How to run data quality monitoring in practice: which metrics to track, how to set thresholds, how to control alert volume, and what to do when a check fails.

Key Takeaways
- A monitor is five decisions, not one. A metric, a check, a threshold, an alert route and a named owner. Teams usually pick the metric and leave the other four to chance, which is why the alerts get muted.
- Six categories of check cover almost everything worth catching. Freshness, volume, schema, distribution, referential integrity and business rules. Each one catches a different class of failure and each one has a blind spot the others cover.
- Set thresholds from history, not from opinion. Fourteen days of observed values and a percentile beats a round number. For seasonal data, compare each day against the same weekday rather than against yesterday.
- Decide your alert budget before you build monitors. A team that can genuinely triage ten alerts a week should not run a monitor set that produces sixty. Work backwards from the budget to the monitor count.
- An alert nobody owns is not a control. Every monitor needs a named human, a severity, a route and a definition of what closure means, or the failure is discovered by a customer.
- Monitoring cost is driven by column count and check frequency. Monitoring every column hourly is the expensive default. Tiering columns by whether anyone downstream depends on them usually removes most of the bill without removing much of the coverage.
What Data Quality Monitoring Involves in Practice
Data quality monitoring is the practice of running automated checks against your data on a schedule, comparing the result to an expected range, and raising an alert to a named owner when the result falls outside it. Everything difficult about it sits in the words "expected range" and "named owner".
A single monitor is five separate decisions, and most teams make only the first one deliberately. The metric says what you measure, for example the number of null values in a column. The check says how you evaluate it, for example whether that count exceeds a limit. The threshold says where the limit sits. The route says who hears about it and by what channel. The owner says whose job it is to act. A monitor missing any of the last four still produces alerts, but nothing happens when it fires.
This distinction matters because monitoring programmes rarely fail on coverage. They fail because a large monitor set was built quickly, thresholds were guessed, alerts landed in a shared channel with no owner, and within two months the channel was muted. The rest of this article is about avoiding that specific outcome.
Why Data Quality Checks Are Worth Running
The case for checks rests on timing rather than on the existence of bad data. Bad data is discovered late, and the cost of a data problem scales with how long it goes unnoticed. A null rate that jumps at 06:00 and is caught at 06:30 costs a rerun. The same jump caught three weeks later has already been read by an executive dashboard, exported to a regulator and used to train a model, and now every one of those has to be corrected and explained.
The second reason has grown quickly. Data that once ended in a dashboard a human could sanity check now feeds retrieval systems and agents that will answer confidently from whatever they are given. Automated checks are the only practical way to notice a problem before an automated consumer acts on it.
The third reason is regulatory. Supervisors in the markets Decube customers operate in, including OJK in Indonesia, APRA in Australia, MAS in Singapore and the NAIC in United States insurance, increasingly ask organisations to evidence control over the data behind reported figures. A monitoring history with dated results and closed incidents is that evidence. An assurance that the team watches the numbers is not.
The Six Categories of Check That Matter
Nearly every useful data quality check belongs to one of six families. They are worth learning as a set, because each one has a blind spot that one of the others covers, and a monitor set drawn from only two families will feel busy while missing whole classes of failure.
1. Freshness
A freshness check asks when the table was last updated and compares that to how often it should be. It is the highest value check per unit of effort, because a stale table is both common and completely invisible in the data itself: every row looks correct.
A concrete example: alert when the maximum value of updated_at in the orders table is more than 90 minutes old during business hours. What it catches is the whole class of silent pipeline failures, jobs that died, credentials that expired, upstream systems that stopped sending. What it misses is anything about the content of the data. A table can be updated perfectly on time and be full of nonsense. Our explainer on data freshness covers how to pick the interval for tables that do not run on a neat schedule.
2. Volume
A volume check counts rows and compares the count to what is normal for that table at that point in the cycle. It is the natural partner to freshness: freshness tells you the pipeline ran, volume tells you it ran properly.
A concrete example: alert when the daily row count for transactions falls below 60 percent or rises above 160 percent of the median for the same weekday over the last four weeks. What it catches is partial loads, duplicated loads, a filter someone changed, and a source system that quietly dropped a region. What it misses is any change that keeps the row count intact, which includes most corruption of values. It is also the check most likely to produce false alarms on genuinely variable data, which is why it needs the seasonal handling described in the next section.
3. Schema
A schema check watches the shape of the table: which columns exist, what type each holds, and whether anything was added, removed or retyped since the last run.
A concrete example: alert on any column removal or type change in a table that has downstream dependencies, and log without alerting on column additions. What it catches is the upstream change nobody announced, which is one of the most common causes of a broken model that ran green. What it misses is everything about the values. A schema check is also the one most worth wiring to ownership rather than to a channel, because the fix almost always sits with a team other than the one receiving the alert.
4. Distribution
A distribution check looks at the statistical shape of a column: null rate, distinct count, minimum and maximum, mean, or the share of rows falling into each category.
A concrete example: alert when the null rate for customer_email exceeds 2 percent, or when the share of orders with status "pending" moves by more than 15 percentage points against the trailing 14 day average. What it catches is the quiet corruption that freshness and volume checks cannot see, including a source field that started arriving empty and an enumeration that gained a new value. What it misses is anything about relationships between tables, and it is the family most sensitive to threshold choice, because a distribution check with a badly chosen limit is a noise generator.
5. Referential integrity
A referential integrity check verifies that the relationships between tables hold: that every foreign key has a parent, that a join does not multiply rows, that a supposedly unique key is unique.
A concrete example: alert when the count of rows in order_items whose order_id has no match in orders is greater than zero, and alert when the row count after a join exceeds the row count of the left table. What it catches is the failure mode that silently doubles revenue in a report, plus records that vanish from analyses because their join partner never arrived. What it misses is the correctness of the values themselves, and it is expensive to run across large tables, so it belongs on the joins that feed reporting rather than on every relationship in the warehouse.
6. Business rules
A business rule check tests a statement that is true about your business and only about your business. Nobody can write these for you, which is exactly why they catch what generic monitoring never will.
Concrete examples: no order may have a delivery date earlier than its order date, the sum of ledger entries per account must reconcile to the account balance, a policy in the active state must have a premium greater than zero. What they catch is logically impossible data, which is usually a defect in an application rather than in a pipeline, and which nothing else will ever flag. What they miss is anything you did not think to write down. They also decay: a rule written for an older product model quietly becomes wrong, so business rules need an owner and a review date more than any other family.
Check Type, What It Catches and a Realistic Alerting Rule
The table below is the version of the six families you can act on. The alerting rules are starting points for a daily warehouse table with real downstream consumers, and they assume you will tune them after two weeks of watching what fires.
| Check type | What it catches | A realistic alerting rule |
|---|---|---|
| Freshness | Dead pipelines, expired credentials, silent upstream stoppages | Alert when the last update is older than 1.5 times the normal interval, during business hours only for tables that do not run overnight |
| Volume | Partial loads, duplicate loads, a dropped source or region | Alert outside 60 to 160 percent of the median for the same weekday over four weeks, evaluated once the load window has closed |
| Schema | Unannounced upstream changes that break models without erroring | Alert on any column drop or type change; log column additions without alerting |
| Distribution | Quiet corruption of values, new enumeration values, fields that stopped arriving | Alert when a null rate or category share moves beyond the 1st to 99th percentile of the trailing 14 days, and only on columns something downstream depends on |
| Referential integrity | Orphan records, join explosions, broken uniqueness | Alert when orphan count is greater than zero on reporting joins; run daily rather than hourly because it is expensive |
| Business rules | Logically impossible records that every generic check passes | Alert on any violation, because a business rule breach is either a real defect or a rule that needs retiring; review the rule set quarterly |
How to Set a Threshold Without Guessing
The threshold is where most monitoring programmes go wrong, and it goes wrong in a predictable way. Somebody picks a round number, 5 percent nulls or a 20 percent volume swing, because it sounds reasonable. The number has no relationship to how the data actually behaves, so it either fires constantly or never fires at all.
The method that works is simple and takes about an hour per table. Run the metric against fourteen days of history without alerting on it. Look at the range of values you observe. Set the threshold at a percentile of that observed range rather than at a round number, and start wide. A first threshold that fires twice in two weeks is useful. One that fires twenty times gets muted before you have a chance to tune it.
Two adjustments matter. First, weight the threshold by consequence, not by how bad the number looks. A 3 percent null rate in a column feeding a regulatory return deserves a tighter limit than a 30 percent null rate in a column nobody queries. Second, revisit thresholds after any deliberate change to the pipeline, because a threshold set before a backfill will misbehave after it.
Seasonal data needs a different comparison rather than a different number. If the volume triples every Monday and collapses every public holiday, comparing today against yesterday guarantees noise. Compare each day against the same weekday over the last four weeks, and maintain an exception list for known events such as holidays, sale periods and month end closes so those days are evaluated against their own history. Where the pattern is more complex than a weekly cycle, flexible thresholds that adapt to the observed pattern are worth more than any static number you could pick.
| How the data behaves | How to set the threshold | Worked example |
|---|---|---|
| Stable, low variance | Percentile of a 14 day observed range, tightened over time | Null rate on customer_id sits between 0 and 0.3 percent for 14 days, so alert above 0.5 percent |
| Weekly seasonality | Compare against the same weekday, not against yesterday | Monday volume is 3 times Wednesday volume, so the Monday bound is built from the last four Mondays |
| Known calendar events | Exception list evaluated against its own history | Month end close triples ledger row count, so the last day of each month is compared to previous month ends |
| Trending up or down | Threshold on the rate of change, not on the absolute value | A table growing 4 percent a week alerts when weekly growth leaves the 0 to 10 percent band |
| New table with no history | Log only for two weeks, then set the threshold | No alert configured until 14 days of observed values exist, so the limit is derived rather than guessed |
Alert Fatigue and the Arithmetic of an Alert Budget
Alert fatigue is easy to file under team morale. Treat it instead as the mechanism by which a monitoring programme stops working, because it has arithmetic you can do before you build anything.
Work it backwards. Decide how much time the team will genuinely spend triaging data alerts each week. Say it is four hours across the rota. A real triage, which means reading the alert, checking the data, deciding whether it matters and either acting or dismissing it with a note, takes about twenty minutes on average once you include the ones that turn out to be real. Four hours therefore buys roughly twelve triaged alerts a week. That is your alert budget, and it is a hard number.
Now compare it to what your monitor set will produce. If you configure 200 monitors and each one fires on average once a month, that is about 46 alerts a week, roughly four times the budget. The team will not work four times harder. It will start skimming, then batching, then muting, and the monitors that mattered will be muted alongside the ones that did not. Nobody decides to ignore alerts. The budget decides for them.
Four moves close the gap, in the order they are worth doing. Cut coverage to columns and tables something downstream actually depends on, which typically removes more than half the monitor set with almost no loss of protection. Widen thresholds on anything that fires more than once a fortnight without a real cause. Group related alerts, so one upstream failure that trips fourteen downstream tables arrives as one incident rather than fourteen notifications. Then split the routes, so only the severities that need a human now go to a channel a human watches now, and the rest go to a daily digest.
One rule keeps the budget honest over time: any monitor that has fired more than three times without anyone acting on it is either wrongly thresholded or watching something nobody cares about. Fix it or delete it. A monitor set that only shrinks when someone complains will always drift back towards noise.
What a Good Data Quality Incident Looks Like
An alert becomes useful at the moment it turns into an incident with a name attached. Four things have to be true for that to happen, and a monitoring setup missing any of them produces notifications rather than control.
- One named owner, not a channel. A specific human is accountable for the incident from the moment it opens. Shared ownership across a team reliably becomes no ownership, and the fastest way to see this is to ask who closed the last three data incidents. If the answer is "the data team", nobody owned them.
- A severity that decides the route. Severity records who gets woken up and how quickly, rather than how bad the data looks. Assign it from downstream consequence, which means a table feeding a regulatory report outranks a table feeding an internal exploration, regardless of how large the anomaly is.
- A root cause traced to a source, not a symptom. "Null rate spiked" is a symptom. "The upstream CRM export changed a field name on 3 August" is a cause. Getting from one to the other is a lineage question, which is why column level data lineage shortens incident time more than any additional monitor does: it tells you what changed upstream and which reports downstream already consumed the bad values.
- A definition of closure that includes downstream. Fixing the data does not close the incident. Closure comes when the corrected data has flowed through, the consumers who read the bad version have been told, and the check that caught it has been reviewed to see whether it should have caught it sooner.
| Severity | When it applies | Route | Expected first response |
|---|---|---|---|
| S1 | Regulatory, financial or customer facing data is wrong or missing | Page the on call owner | Within 15 minutes, at any hour |
| S2 | A business critical dashboard or model is affected but the exposure is internal | Direct message to the named owner plus the team channel | Within the business day |
| S3 | A quality metric moved beyond its threshold with no confirmed downstream impact | Team channel | Next working day |
| S4 | Informational, including expected changes and known exceptions | Daily digest, no notification | Reviewed weekly |
The pattern above is worth writing down before the first incident rather than after the third, and it needs to live where the alerts arrive. Our guide to incident workflows for data quality covers how the states, assignment and closure records fit together once this becomes routine.
What Data Quality Monitoring Costs
Cost is the question buyers ask earliest and most published material answers last, so it is worth being direct about what drives it. Three things do: how many columns you monitor, how often each check runs, and how much data each check has to scan to produce its answer.
The multiplication is unforgiving. Monitoring 40 columns hourly is 960 checks a day. Monitoring 4,000 columns hourly is 96,000 checks a day, and each one is a query against your warehouse that you pay for on top of whatever the monitoring tool charges. This is why per column pricing feels reasonable at pilot scale and alarming at production scale, and why the honest answer to "what does it cost" starts with "how many columns actually need watching".
That question has a good answer. Most warehouses contain far more columns than anything downstream reads. Tiering by dependency rather than monitoring everything uniformly is the single largest cost decision available, and it usually costs very little coverage.
| Tier | What sits in it | What to monitor | Typical share of columns |
|---|---|---|---|
| Tier 1, critical | Columns feeding regulatory reports, financial statements, customer facing figures or production models | All six check families, run at the frequency of the pipeline | Small, often under 5 percent |
| Tier 2, important | Columns feeding widely used dashboards and internal decisions | Freshness, volume and schema always; distribution on the columns that vary | Perhaps 15 to 25 percent |
| Tier 3, everything else | Staging tables, raw landing zones, columns nothing reads | Table level freshness and volume only, no column checks | The majority |
| Tier 0, unused | Columns with no queries against them in 90 days | Nothing. Consider deprecating them instead | Larger than most teams expect |
Two further levers matter once the tiering is in place. Match check frequency to how often the data changes, because running an hourly check on a table that loads once a day pays for 23 answers you already knew. And prefer checks that read metadata over checks that scan rows where the two would tell you the same thing, since a freshness check against table metadata costs a fraction of a full column profile.
The number worth carrying into a vendor conversation is the price for your tier 1 and tier 2 column count at your actual pipeline frequency, plus the warehouse compute the checks will consume. Vendors quote the price per column. The compute is the part that surprises people.
Where Decube Fits
Decube runs the six check families as scheduled monitors and routes what they find into incidents with owners and severities rather than into a notification channel. The data observability platform covers the metric, check and threshold layer, including thresholds that adapt to seasonal patterns instead of holding a fixed number, and it connects each failure to column level data lineage so the root cause and the affected downstream tables arrive with the alert rather than after an hour of searching.
The part worth deciding yourself is the tiering. Which columns are tier 1 is a business question, not a platform question, and getting it right is what keeps both the alert volume and the bill in proportion. If you want to see how the checks and incident routing work against your own tables, request a demo and bring a list of the ten columns you would be most embarrassed to get wrong.
Frequently Asked Questions
How do I test data quality?
Run automated checks on a schedule and compare each result to a range derived from history. Six families cover most failures: freshness, volume, schema, distribution, referential integrity and business rules. Start with freshness and volume on the tables something downstream depends on, because they catch the most common failures for the least effort, then add distribution and business rule checks on the columns that matter most.
How do I check data quality for a process improvement project?
Measure before you change anything. Profile the tables in scope for null rates, distinct counts, duplicate keys and orphan records, and record the numbers with a date. That baseline is what lets you show the improvement later. Then convert the worst findings into standing monitors so the gains do not quietly reverse once the project closes.
How do I maintain data quality over time?
Attach an owner to every monitor, keep the alert volume inside what the team can genuinely triage, and review the monitor set on a schedule. The failure mode is never a lack of checks. It is a growing pile of alerts nobody acts on, at which point the checks exist but no longer control anything.
What is data quality degradation detection?
It is the practice of noticing that a quality metric is drifting in the wrong direction before it crosses a hard limit. Rather than alerting only when the null rate exceeds 5 percent, you track the trend and alert when the rate of change leaves its normal band. It catches slow problems, such as a source system gradually sending more empty fields, that a fixed limit only reveals once the damage is done.
What is consistency in data quality?
Consistency means the same fact holds the same value everywhere it appears. A customer whose status is active in the billing system and closed in the warehouse is a consistency failure even though both records look valid on their own. In monitoring terms it is checked with cross system comparisons and referential integrity checks rather than with column profiling.
How do I monitor data freshness?
Compare the most recent timestamp in the table against how often the table is supposed to update, and alert when the gap exceeds about 1.5 times the normal interval. Use the load timestamp rather than a business date, restrict the check to the hours the pipeline is meant to run, and keep an exception list for known non loading days so weekends and holidays do not generate false alarms.
Can data contracts act as data quality checks at the source?
Yes, and they catch problems earlier than warehouse monitoring can, because a contract rejects bad data at the point of production rather than detecting it downstream. They do not replace monitoring. A contract enforces the shape and the rules you agreed on, while monitoring catches the drift nobody agreed on, such as a valid field whose distribution changed.
Can agentic AI run data quality controls?
It can already do parts of the work well: proposing checks by profiling a table, drafting thresholds from observed history, grouping related alerts into one incident and suggesting a likely root cause from lineage. What it should not do unsupervised is decide that an anomaly is acceptable and close it, because that is a business judgement with a named accountable human behind it.
What are the challenges of using agentic AI for data quality?
Three recur. An agent that both raises and resolves incidents removes the audit trail a regulator expects, so the accountable human has to stay in the loop. An agent proposing thresholds from history will encode existing bad data as normal unless someone reviews the baseline. And agents depend on the same lineage and metadata that most organisations have not finished building, so the agent is only as good as the context it is given.
Configuring Freshness, Volume and Schema Drift Monitors in Decube
This article has argued that a check only becomes a control once it runs on a schedule with a threshold and an owner behind it. The 55 second walkthrough shows that step inside Decube: schema drift and job failure monitors switch on by themselves as soon as a source is connected, while freshness, volume, field health and custom SQL monitors are configured by hand against a chosen dataset and incident level. Watch it to see the freshness monitor learn each table's own update pattern instead of holding a fixed interval, which is the threshold problem from the section above solved in the product.














.webp)