GDPR and HIPAA Data Compliance Tooling: The 9 Capabilities Each Requires

The 9 capabilities GDPR and HIPAA each require of a data platform, cited to the article and CFR section that sets them, plus what no vendor can do for you.

By

Jatin S

Updated on

September 9, 2026

Key Takeaways

  • Both regimes ask your tooling for the same nine things, in different words. An inventory of where regulated data sits, classification down to the column, a recorded purpose, access bound to a role, a log of who read what, lineage, executable retention, breach scoping, and documentation that outlives its author.
  • The overlap is bigger than the difference, so build once. Inventory, classification, least access, logging and retained written documentation satisfy GDPR Article 30, Article 25(2), Article 5(2) and Article 32 and HIPAA 164.514(d)(2), 164.312(a), 164.312(b) and 164.316(b) at the same time.
  • The real difference is erasure against retention. GDPR Article 17 gives an individual the right to have data erased. HIPAA gives no such right and 45 CFR 164.316(b)(2)(i) requires documentation to be kept for six years. A platform that can only do one of those will fail the other audit.
  • Consent is not the GDPR default and is not the HIPAA model at all. Article 6(1) lists six lawful bases and consent is only one of them. HIPAA works from permitted uses under 164.502(a) plus the minimum necessary limit at 164.502(b), and asks for an authorization only outside that set.
  • An auditor accepts exports, not screenshots. The catalog produces the Article 30 record and the 164.514(d)(2) access map, lineage answers Article 19 and the 164.528 accounting of disclosures, and access control produces the grant list and its approval history.
  • No vendor can be your controller. Article 28(3)(a) says a processor acts only on documented instructions, Article 24(1) puts the duty to demonstrate compliance on the controller, and 164.308(a)(1)(ii)(A) makes the risk analysis a required task for you.

What GDPR and HIPAA Each Require of Your Data Tooling

GDPR requires your tooling to prove that you know what personal data you hold, why you hold it, who can reach it, where it has flowed and when it will be deleted. HIPAA requires your tooling to prove that access to protected health information is limited to what each role needs, that every touch of it is recorded and reviewed, and that the documentation of all of this has been kept.

Read those two sentences next to each other and the shape of the work becomes obvious. GDPR is written around the individual and their rights over their own data, so it pushes you toward knowing every copy and being able to act on it. HIPAA is written around the custodian and their duty of care, so it pushes you toward limiting access and proving what happened. The engineering that satisfies one covers most of the other.

The 9 Capabilities Your Data Tooling Must Have

Read the table first. The nine sections after it explain what each capability has to do in practice and what an auditor will ask you to produce for it.

CapabilityWhat GDPR requires, and whereWhat HIPAA requires, and whereEvidence an auditor accepts
1. Inventory of regulated dataArticle 30(1) record of processing activities, naming the categories of data subjects and of personal data, the recipients, and the time limits for erasure. Article 30(4) makes it available to the supervisory authority on request.164.514(d)(2)(i)(B) requires you to identify, for each person or class of person, the category or categories of protected health information to which access is needed.A current export of every asset with its owner, its data categories and its retention period, dated.
2. Column level classificationArticle 9(1) prohibits processing data concerning health unless an Article 9(2) condition applies, so you have to know which columns are health data.164.514(b)(2) lists the eighteen identifier types that must be removed before information counts as de-identified under the safe harbor method.The classification policy list, the rules that apply it, and the set of columns each policy currently covers.
3. Recorded purpose and lawful basisArticle 5(1)(b) purpose limitation and Article 6(1), which requires at least one of six lawful bases. Article 30(1)(b) puts the purposes in the record.164.502(a) permits use and disclosure only for a listed purpose, and 164.502(b) limits it to the minimum necessary for that purpose.The purpose recorded as a field on the asset itself, not in a separate spreadsheet nobody updates.
4. Access bound to role and categoryArticle 25(2) says that by default only the personal data necessary for each purpose are processed, and that the obligation covers their accessibility.164.312(a)(1) access control, with unique user identification required at 164.312(a)(2)(i), and 164.308(a)(4) information access management.The grant list per classification policy, plus who approved each grant and when.
5. A log of who read whatArticle 5(2) accountability: the controller must be able to demonstrate compliance. Article 24(1) repeats the duty.164.312(b) audit controls, and 164.308(a)(1)(ii)(D) information system activity review, which is a required specification, not an addressable one.Access logs at asset level, plus proof that somebody actually reviews them on a schedule.
6. Column level lineageArticle 15(1)(c) entitles the individual to the recipients or categories of recipient. Article 19 requires you to pass an erasure or a rectification on to each of them.164.528(a)(1) gives an individual the right to an accounting of disclosures made in the six years before the request, subject to the exceptions listed there.A lineage graph you can query by column, showing every downstream copy and consumer.
7. Retention and erasure you can executeArticle 17(1) right to erasure, Article 5(1)(e) storage limitation, and Article 12(3), which gives you one month to respond, extendable by two further months.The opposite instinct. No right to erasure, and 164.316(b)(2)(i) requires documentation to be retained for six years from creation or from when it last was in effect.A retention rule attached to the classification, and a record of what was deleted, when, and everywhere it existed.
8. Breach scopingArticle 33(1) gives 72 hours to notify the supervisory authority, and Article 33(3)(a) wants the categories and approximate number of data subjects and records concerned.164.404(b) requires notification to individuals without unreasonable delay and no later than 60 calendar days after discovery.A query that answers which records, which categories and how many, in hours rather than weeks.
9. Documentation that outlives its authorArticle 24(2) implementation of data protection policies, and Article 30(3), which requires the record to be in writing, including electronic form.164.316(a) and (b)(1) require written policies and procedures, retained under 164.316(b)(2)(i) and reviewed and updated under (b)(2)(iii).Versioned policy documents attached to the assets they govern, with an approval history.

1. An inventory of where regulated data actually sits

Article 30 of GDPR is the provision that most directly describes a data catalog, even though it never uses the word. It requires a record of processing activities containing the purposes of the processing, a description of the categories of data subjects and of the categories of personal data, the categories of recipients, the envisaged time limits for erasure of the different categories of data, and a general description of the security measures. Article 30(4) requires you to make that record available to the supervisory authority on request.

HIPAA arrives at the same place from the access side. 45 CFR 164.514(d)(2)(i) requires a covered entity to identify the persons or classes of persons in its workforce who need access to protected health information, and, for each of them, the category or categories of protected health information to which access is needed. You cannot write that down for a system you have not inventoried.

The tooling test is simple. Can it discover every source, including the ones with no tidy connector, and can it keep the inventory current without somebody remembering to update it. An inventory that is a document rather than a live index is out of date the week after it is signed off.

2. Classification down to the column

Article 9(1) prohibits the processing of data concerning health, along with genetic and biometric data, unless one of the conditions in Article 9(2) applies. That prohibition is unenforceable inside your own estate unless you know which columns hold that data. Classification is not a labeling exercise for its own sake; it is the thing that makes every other control addressable.

HIPAA is unusually concrete here. The safe harbor method at 45 CFR 164.514(b)(2) lists eighteen identifier types that must be removed before information stops being individually identifiable, and the list is worth reading in full because it is longer than most people assume. It includes names, all geographic subdivisions smaller than a state, all elements of dates except the year, telephone and fax numbers, email addresses, social security numbers, medical record numbers, health plan beneficiary numbers, account numbers, certificate and license numbers, vehicle and device identifiers, web addresses, IP addresses, biometric identifiers, full face photographic images, and any other unique identifying number, characteristic or code.

That list is a classification specification you can implement directly. A tool that can hold a policy per identifier type, apply it by pattern or keyword as new columns arrive, and show you every column currently under each policy, is doing the work that 164.514(b) assumes you have done.

Classification is also where governance stops being a document and starts being enforceable, which is the connection between this section and the broader data governance program it sits inside.

3. A recorded purpose, and a lawful basis under it

Article 5(1)(b) requires personal data to be collected for specified, explicit and legitimate purposes and not further processed in a manner incompatible with them. Article 6(1) then requires at least one of six lawful bases for the processing to be lawful at all: consent, performance of a contract, a legal obligation, vital interests, a task in the public interest, or legitimate interests. Consent is one of six, and for a hospital or an insurer it is frequently not the one that applies.

HIPAA does not use the language of lawful basis. It works the other way round: 164.502(a) sets out the uses and disclosures a covered entity is permitted or required to make, and anything outside that set needs an authorization under 164.508. Treatment, payment and health care operations sit inside the permitted set and do not need one.

For tooling, both regimes reduce to the same requirement: the purpose has to be a field on the asset, visible to whoever queries it, and it has to be reviewed when the asset is reused for something new. A purpose recorded once in a project document at procurement time and never seen again is not a control.

4. Access bound to a role and a data category, not to a table name

Article 25(2) is the sharpest sentence in GDPR on this point. It requires the controller to ensure that, by default, only personal data which are necessary for each specific purpose are processed, and it says explicitly that the obligation applies to the amount collected, the extent of processing, the period of storage and their accessibility.

HIPAA states the same idea as a standard with named implementation specifications. 164.312(a)(1) requires technical policies and procedures that allow access only to those persons or software programs granted access rights under 164.308(a)(4). Unique user identification at 164.312(a)(2)(i) and an emergency access procedure at (a)(2)(ii) are marked Required. Automatic logoff and encryption at (a)(2)(iii) and (iv) are marked Addressable, which is not the same as optional: 164.306(d)(3) requires you to assess each addressable specification and, if you do not implement it, to document why it would not be reasonable and appropriate and to implement an equivalent alternative measure if one is.

The engineering consequence is that grants have to be expressed against classifications rather than against object names. A rule that says this role may not read any column classified as PHI keeps working when a new table lands. A list of table names does not.

HIPAA states what that identification has to produce:

A covered entity must identify: (A) Those persons or classes of persons, as appropriate, in its workforce who need access to protected health information to carry out their duties; and (B) For each such person or class of persons, the category or categories of protected health information to which access is needed and any conditions appropriate to such access. (45 CFR 164.514(d)(2)(i))

5. A log of who read what, and evidence that somebody reads the log

This is the requirement most often half implemented. HIPAA sets it twice. 164.312(b) is the audit controls standard, and 164.308(a)(1)(ii)(D) requires procedures to regularly review records of information system activity such as audit logs, access reports and security incident tracking reports. That second one is marked Required, so the review itself is the obligation, not just the existence of the log.

Standard: Audit controls. Implement hardware, software, and/or procedural mechanisms that record and examine activity in information systems that contain or use electronic protected health information. (45 CFR 164.312(b))

Note the two verbs. Record and examine. A platform that writes logs nobody opens satisfies half of a standard.

GDPR does not prescribe logging in the same detail, but Article 5(2) makes the controller responsible for being able to demonstrate compliance with all six principles in Article 5(1), and Article 24(1) repeats that duty. In practice the only way to demonstrate that access was limited is to show a record of the access that occurred.

6. Column level lineage, because both regimes ask where the data went

Article 15(1)(c) entitles a data subject to know the recipients or categories of recipient to whom their personal data have been or will be disclosed. Article 19 then goes further and puts a positive duty on you.

The controller shall communicate any rectification or erasure of personal data or restriction of processing carried out in accordance with Article 16, Article 17(1) and Article 18 to each recipient to whom the personal data have been disclosed, unless this proves impossible or involves disproportionate effort. (Regulation (EU) 2016/679, Article 19)

You cannot communicate an erasure to each recipient if you do not know who the recipients are. In a modern warehouse the recipients are usually internal: the twelve downstream tables, three dashboards and one feature store that copied the column.

HIPAA asks a version of the same question at 45 CFR 164.528(a)(1), which gives an individual the right to an accounting of disclosures made in the six years before the request. The section lists nine categories of disclosure that are excluded from the accounting, including treatment, payment and health care operations, so the scope is narrower than it first appears. It is still a six year lookback.

Both of those are lineage questions, and both need it at column level rather than table level, because the obligation attaches to a person's data and not to a dataset. Table level lineage tells you a table was copied. Column level lineage tells you which patient identifier ended up in the feature store.

7. Retention and erasure you can actually execute

Article 17(1) gives an individual the right to obtain erasure without undue delay on six listed grounds, including that the data are no longer necessary for the purpose, that consent has been withdrawn with no other ground available, and that the data were unlawfully processed. Article 5(1)(e) sets the background rule that data are kept in a form permitting identification for no longer than is necessary. Article 12(3) gives you one month to respond, extendable by two further months where the request is complex, provided you tell the individual inside the first month.

HIPAA has no equivalent right. It runs the other way. 45 CFR 164.316(b)(2)(i) requires documentation to be retained for six years from the date of its creation or the date when it last was in effect, whichever is later. An organization subject to both regimes therefore has to be able to delete a person's data on request under one law while preserving records under the other, which is a policy decision per data category and not a switch you flip on a platform.

The tooling requirement is that retention lives with the classification. A policy that says data classified this way is deleted after this period, applied automatically to every column that carries the classification, is the only version of this that survives growth. It also depends entirely on capability 6, because an erasure you cannot trace to every copy is an erasure you cannot certify.

8. Breach scoping, against two very different clocks

Article 33(1) requires the controller to notify the competent supervisory authority without undue delay and, where feasible, not later than 72 hours after becoming aware of a personal data breach, unless it is unlikely to result in a risk to the rights and freedoms of natural persons. Where the 72 hours are missed, the notification has to carry the reasons for the delay. Article 33(3)(a) says the notification must describe the categories and approximate number of data subjects concerned and the categories and approximate number of records concerned. Article 33(5) requires you to document every breach, its effects and the remedial action, so the authority can verify compliance.

HIPAA is more generous on time and narrower on scope. 45 CFR 164.404(b) requires notification to affected individuals without unreasonable delay and in no case later than 60 calendar days after discovery of a breach of unsecured protected health information.

The 72 hour clock is what makes this a tooling question rather than a legal one. Article 33(3)(a) is asking which categories and how many, and that is a catalog and lineage query. Teams that cannot answer it in hours end up notifying with an estimate and correcting it later, which is exactly the position nobody wants to be in.

9. Documentation that outlives the person who wrote it

45 CFR 164.316(a) requires reasonable and appropriate policies and procedures, and 164.316(b)(1) requires them to be maintained in written form, which may be electronic. The implementation specifications are all marked Required: retain for six years under (b)(2)(i), make the documentation available to the people who have to follow it under (b)(2)(ii), and review and update it in response to environmental or operational changes under (b)(2)(iii).

GDPR reaches the same result through Article 24(2), on implementing data protection policies where proportionate, and Article 30(3), which requires the record of processing activities to be in writing, including in electronic form.

Policies do get written. What goes wrong is that they are written somewhere the data is not, so nobody reading a table ever sees the rule that governs it. Attaching the policy to the asset is what metadata governance is for, and it is the difference between a policy that is followed and a policy that exists.

Where GDPR and HIPAA Overlap, So You Build It Once

Five things satisfy both regimes at the same time. If you are standing up a program for the first time, these are the five to build first, because every hour spent on them counts twice.

  • Inventory and classification. Article 30(1)(c) wants the categories of personal data. 164.514(d)(2)(i)(B) wants the categories of protected health information per role. One catalog with column level classification answers both.
  • Access limited to what the role needs. Article 25(2) makes it the default. 164.312(a)(1) makes it a standard and 164.502(b) makes it the minimum necessary rule. One access model built on classifications satisfies all three.
  • Technical security measures. Article 32(1)(a) names pseudonymisation and encryption explicitly. 164.312(a)(2)(iv) and 164.312(e)(2)(ii) name encryption as addressable specifications for stored and transmitted data. The same controls answer both, and Article 32(1)(d) adds a duty to test them regularly.
  • A logged, reviewable record of activity. Article 5(2) requires you to be able to demonstrate compliance. 164.312(b) and 164.308(a)(1)(ii)(D) require the log and the review of it. One audit log serves both.
  • Written documentation, retained and reviewed. Article 30(3) and (4) require a written record available to the supervisory authority. 164.316(b) requires written policies retained for six years and reviewed as things change.

The quality of what is in the catalog decides whether any of this holds. A classification applied to a column that is stale or duplicated somewhere else does not protect anything, which is why this work and data quality management are the same program rather than two.

Where They Genuinely Differ

The differences below are the ones that change what you build, not the ones that change the vocabulary. Each row names the provision on both sides.

QuestionGDPRHIPAA
What makes holding the data lawfulArticle 6(1): one of six lawful bases, and for health data an Article 9(2) condition on top.164.502(a): the use or disclosure must fall inside the permitted or required set, or carry an authorization under 164.508.
The role of consentOne lawful basis among six. Article 7(3) lets the individual withdraw at any time and says it must be as easy to withdraw as to give.Not the operating model. Treatment, payment and health care operations need no authorization; an authorization is required only outside the permitted set.
The default limiting ruleArticle 5(1)(c) data minimisation: adequate, relevant and limited to what is necessary.164.502(b) minimum necessary, with the exceptions listed at 164.502(b)(2), including disclosures to a provider for treatment.
DeletionArticle 17(1) right to erasure on six listed grounds, with Article 19 requiring you to pass it on to each recipient.No right to erasure. 164.316(b)(2)(i) requires documentation to be retained for six years.
The clock on an individual requestArticle 12(3): one month, extendable by two further months if you notify inside the first month.164.524(b)(2): 30 days, with one extension of no more than 30 days and a written statement of the reasons.
The clock on a breachArticle 33(1): 72 hours to the supervisory authority where feasible, with reasons required if you miss it.164.404(b): no later than 60 calendar days after discovery, to the affected individuals.
Telling people where their data wentArticle 15(1)(c): the recipients or categories of recipient. Article 19: notify each of them of an erasure or rectification.164.528(a)(1): an accounting of disclosures for the six years before the request, with nine listed exclusions.
The vendor relationshipArticle 28(3): a written contract with the eight listed terms, including deletion or return of the data at the end of the service.164.504(e)(2): a business associate contract setting the permitted uses and the safeguards, with 164.504(e)(1)(ii) requiring you to act on a known pattern of breach by the associate.
The ceiling on a penaltyArticle 83(4): up to 10 000 000 EUR or 2 percent of total worldwide annual turnover, whichever is higher. Article 83(5): up to 20 000 000 EUR or 4 percent, for infringements of the basic principles and of data subject rights.160.404(b)(2) sets four tiers by culpability, each with a per violation minimum, a per violation maximum and a calendar year cap for identical violations. The dollar amounts are adjusted for inflation every year and published in the table at 45 CFR 102.3.

Note the asymmetry in that last row, because it is misquoted constantly. GDPR sets its two ceilings in the text of Article 83 itself, so those two numbers are stable and citable. HIPAA does not: 45 CFR 160.404(a) says the amounts are adjusted annually and published at 45 CFR part 102, so any fixed dollar figure you read in an article is only as current as the year it was written. If you need the number, read the table, not a blog.

What a Catalog, Lineage and Access Control Each Contribute to an Audit

An audit is a request for artifacts. The useful way to think about tooling is to ask which artifact each component produces, because that is what you will be asked to hand over.

The catalog produces the record and the access map

The Article 30 record of processing activities and the 164.514(d)(2)(i) statement of who needs access to which categories are both, in substance, an export from a catalog: every asset, its owner, its data categories, its purpose, its recipients and its retention. The reason to hold this in a catalog rather than a document is Article 30(4), which requires you to make the record available to the supervisory authority on request. A record assembled by hand each time it is asked for will be out of date the day it is produced.

The same export is what tells you whether your inventory is complete, which is why the coverage of a data governance platform across sources with no native connector matters more in a regulated estate than any individual feature does.

Lineage answers the questions about movement

Three obligations are lineage questions in disguise. Article 19 asks you to tell each recipient about an erasure. 164.528 asks you to account for disclosures over six years. Article 33(3)(a) asks, inside 72 hours, which categories and approximately how many records were affected. All three are answerable in minutes with column level lineage and answerable in weeks without it.

Classification and access control produce the grant list

Under 164.312(a)(1) an auditor asks for the list rather than for a description of your access model: which roles can read which classified data today, who approved that, and when it was last reviewed. Classification policies are what make that list expressible in a sentence rather than in ten thousand table names, and an approval step on policy changes is what turns the list into evidence rather than a snapshot.

The four minute walkthrough above shows the mechanics: a policy with its own label and color so the classification is visible in the catalog, rules that tag columns by keyword or regular expression and pick up new tables as they arrive, a managed asset view listing every object under the policy with a CSV export for the audit, and a reviewer approval step on every policy change. The CSV export is the artifact. The approval history is what makes it credible.

What a Vendor Cannot Do for You

Every platform in this category, Decube included, will tell you it helps with GDPR and HIPAA. That is true and it is also incomplete. Five duties stay with you no matter what you buy, and each of them is put there by the text of the law rather than by a vendor's reluctance.

  • It cannot be your controller. Article 28(3)(a) says a processor processes personal data only on documented instructions from the controller, and Article 24(1) puts the duty to demonstrate compliance on the controller. A signed processing agreement moves obligations, not accountability.
  • It cannot choose your purpose or your lawful basis. Article 5(1)(b) and Article 6(1) are decisions about your business, made before any tool is configured. A platform can record the answer and enforce it; it cannot supply it.
  • It cannot decide what minimum necessary means in your organization. 164.514(d)(2)(i)(A) puts the identification of who needs access on the covered entity. Only you know that a billing analyst does not need a clinical note.
  • It cannot do your risk analysis. 164.308(a)(1)(ii)(A) makes an accurate and thorough risk analysis a Required implementation specification, and 164.306(b)(2) makes the whole security decision a judgment about your own size, technical infrastructure, costs and risks. GDPR adds Article 35, which requires a data protection impact assessment before processing that is likely to result in a high risk.
  • It cannot rescue a business associate contract you do not monitor. 164.504(e)(1)(ii) says a covered entity is not in compliance if it knew of a pattern of activity by a business associate that amounted to a material breach and did not take reasonable steps to cure it, and terminate the contract if that failed.

There is a sixth that is easy to miss. Several HIPAA specifications are marked Addressable rather than Required, and 164.306(d)(3) says that where you do not implement one, you must document why it would not be reasonable and appropriate and implement an equivalent alternative measure if one is. The documenting is yours. A vendor cannot write your justification for turning something off.

The Other Regulators a Healthcare or Insurance Data Team Answers To

GDPR and HIPAA are the two most searched, but they are rarely the only two in scope. Decube customers in regulated markets also answer to OJK in Indonesia, APRA in Australia, MAS in Singapore, and the NAIC framework for insurance in the United States. Almost no vendor content covers those, which is why buyers in those markets get generic answers when they ask.

The useful point is that the nine capabilities above do not change. Inventory, classification, purpose, access, logging, lineage, retention, breach scoping and documentation are what every one of those supervisors asks for, in their own vocabulary and with their own reporting formats. What changes is who receives the export and how often, not what has to exist behind it.

If you are also deploying AI on top of that data, the EU AI Act adds its own timetable, and it is worth naming which obligation you mean because they land on different dates. According to the European Commission's own guidelines for providers of general purpose AI models, read on 6 September 2026, obligations for providers of GPAI models entered into application on 2 August 2025, the Commission's enforcement powers enter into application from 2 August 2026, and providers of GPAI models placed on the market before 2 August 2025 must comply by 2 August 2027. That is three dates for one set of obligations, and any article that gives you a single "EU AI Act deadline" has collapsed them. The Commission guidance is the source to check against.

Where to Start if You Are Standing This Up Now

Start with capability 1 and capability 2, in that order, and do not start anywhere else. Every other item on the list depends on knowing where the regulated data is and what it is. Access control without classification degenerates into a list of table names. Lineage without an inventory covers the sources you happened to connect. Retention without classification is a policy nobody can apply.

Then work down the list in order, because the order is not arbitrary. Purpose depends on inventory, access depends on classification, logging depends on access, lineage depends on the catalog, retention depends on lineage, and breach scoping depends on all of them.

The test to apply to any platform you are evaluating is whether it can produce, on demand and without a project, the four artifacts an auditor asks for: the asset inventory with categories and retention, the grant list per classification with its approval history, the access log with evidence of review, and the column level lineage for a named identifier. The word HIPAA appearing on a website proves nothing. Ask for those four on your own data during the trial.

Frequently Asked Questions

What is the difference between GDPR and HIPAA?

GDPR is a European regulation covering all personal data about anyone in the EU, and it is built around the individual's rights over that data. HIPAA is a United States rule covering protected health information held by covered entities and their business associates, and it is built around the custodian's duty of care. The practical differences that change what you build are erasure, where GDPR Article 17 gives a right to deletion and HIPAA gives none while 45 CFR 164.316(b)(2)(i) requires six years of retained documentation; the breach clock, which is 72 hours to the supervisory authority under GDPR Article 33(1) and 60 calendar days to individuals under 45 CFR 164.404(b); and consent, which is one of six lawful bases under GDPR Article 6(1) but is not the HIPAA operating model at all.

Does HIPAA have a right to erasure like GDPR?

No. There is no right to erasure anywhere in HIPAA. The Privacy Rule gives individuals a right of access under 45 CFR 164.524, a right to request amendment, and a right to an accounting of disclosures under 164.528, but not a right to have their record deleted. HIPAA pushes in the opposite direction: 45 CFR 164.316(b)(2)(i) requires documentation to be retained for six years from creation or from when it last was in effect, whichever is later. An organization subject to both regimes has to decide, per data category, which obligation governs.

What is HIPAA data governance?

It is the set of controls that let a covered entity or business associate show that access to protected health information is limited, recorded and reviewed. In tooling terms it is four things: an inventory naming which systems and columns hold PHI, a statement of which roles need which categories of PHI as required by 45 CFR 164.514(d)(2)(i), access control implementing 164.312(a)(1) with unique user identification, and audit controls under 164.312(b) plus the regular review of activity records required by 164.308(a)(1)(ii)(D). The written policies behind those controls have to be kept for six years under 164.316(b)(2)(i).

What does GDPR require of a data catalog?

GDPR never uses the phrase, but Article 30 describes one. The record of processing activities must contain the purposes of the processing, the categories of data subjects and of personal data, the categories of recipients, the envisaged time limits for erasure of the different categories of data, and a general description of the security measures under Article 32(1). Article 30(3) requires it in writing, including electronic form, and Article 30(4) requires you to make it available to the supervisory authority on request. A catalog that holds owner, classification, purpose, recipients and retention against every asset produces that record as an export.

Can one data platform cover both GDPR and HIPAA?

It can cover the overlap, which is most of the engineering. Five things satisfy both at once: inventory and classification, access limited to what a role needs, technical security measures including encryption and pseudonymisation, a logged and reviewed record of activity, and written documentation retained and reviewed. What no single platform settles for you is the conflict between GDPR Article 17 erasure and the HIPAA six year retention rule at 164.316(b)(2)(i). That is a policy decision per data category, taken by you, and then configured.

What evidence does an auditor actually accept?

Exports, not screenshots. Four artifacts cover most of what is asked for. The asset inventory with owner, data categories, purpose and retention, which is the Article 30 record. The grant list per classification with who approved each grant and when, which is what 45 CFR 164.312(a)(1) is asking you to demonstrate. The access log with evidence that somebody reviews it, because 164.312(b) says record and examine and 164.308(a)(1)(ii)(D) makes the review a required specification. And column level lineage for a named identifier, which is what answers GDPR Article 19 and the 164.528 accounting of disclosures.

Which data compliance tools suit a healthcare company?

Judge them on the four artifacts an audit asks for rather than on the logos on the website. The tool has to discover every source that holds PHI including the ones with no native connector, classify down to the column against something like the eighteen identifier types listed at 45 CFR 164.514(b)(2), bind access to those classifications rather than to table names, log and expose who read what, carry column level lineage so an erasure or an accounting of disclosures is answerable, and export all of it. Decube covers those in one platform, which is what lets the inventory, the grant list, the log and the lineage be read together instead of assembled from four tools at audit time. Ask any vendor to produce the four artifacts on your own data during the trial.

How long do you have to report a data breach under GDPR and HIPAA?

Under GDPR Article 33(1) the controller notifies the competent supervisory authority without undue delay and, where feasible, not later than 72 hours after becoming aware of the breach, unless it is unlikely to result in a risk to the rights and freedoms of natural persons; if the 72 hours are missed the notification must carry the reasons for the delay. Under 45 CFR 164.404(b) a covered entity notifies affected individuals without unreasonable delay and in no case later than 60 calendar days after discovery of a breach of unsecured protected health information. GDPR Article 33(3)(a) also requires you to state the categories and approximate number of data subjects and records concerned, which is the part that needs a catalog.

What are the penalties under GDPR and HIPAA?

GDPR sets its ceilings in the text. Article 83(4) allows fines up to 10 000 000 EUR or 2 percent of total worldwide annual turnover of the preceding financial year, whichever is higher, for infringements of controller and processor obligations. Article 83(5) allows up to 20 000 000 EUR or 4 percent, whichever is higher, for infringements of the basic principles, the conditions for consent and data subject rights. HIPAA works differently: 45 CFR 160.404(b)(2) sets four tiers by culpability, each with a per violation minimum, a per violation maximum and a calendar year cap for identical violations, and 160.404(a) says the amounts are adjusted for inflation annually and published in the table at 45 CFR 102.3. Read that table for the current figures rather than trusting a number quoted in an article.

Does the EU AI Act change what healthcare data tooling must do?

It adds obligations on the AI side rather than changing the data controls, and its dates land differently depending on which obligation you mean. According to the European Commission's guidelines for providers of general purpose AI models, obligations for providers of GPAI models entered into application on 2 August 2025, the Commission's enforcement powers enter into application from 2 August 2026, and providers of GPAI models placed on the market before 2 August 2025 must comply by 2 August 2027. Never treat "the EU AI Act deadline" as one date. For a data team the practical effect is that the classification, lineage and access records you already keep for GDPR and HIPAA are the same records that evidence what an AI system was trained on and what it can reach.

Is Atlan worth it?
Atlan is worth it if your primary need is a modern data catalog with strong column-level lineage and cloud-native integrations (Snowflake, dbt, Databricks). It is harder to justify if you also need data observability and quality coverage across a heterogeneous stack — those capabilities require separate vendors, adding cost and complexity.
What is the best Atlan alternative
Decube is purpose-built for regulated financial services, with native observability, approval-gated lineage, PII auto-classification, and an AI layer (TrustyAI) that does not route metadata to a public LLM. These map directly to regulatory frameworks supervised by MAS, OJK, BNM, and APRA. Atlan AI's OpenAI dependency is often a procurement blocker in these environments.
How does Atlan compare to Alation?
Both are catalog-first platforms with strong discovery. Alation pioneered search-first data culture and analyst adoption. Atlan is stronger on column-level lineage and cloud integrations. Both require external tooling for observability and broad data quality coverage.
How long does it take to migrate from Atlan to another platform?
Migration time depends on estate size and the number of active integrations. SaaS-native platforms like Decube deploy in 2–6 weeks without professional services. The longer task is typically re-establishing business glossaries, data ownership, and custom attributes — that effort is roughly the same regardless of which platform you move to.
What is the difference between a context layer and a semantic layer?
A semantic layer standardizes how metrics are defined and calculated so every analyst and BI tool uses the same numbers. A context layer encodes governance rules, data lineage, quality signals, and organizational knowledge so AI agents can make safe, autonomous decisions. The semantic layer is for human-facing analytics. The context layer is for AI-facing autonomy.
Can I use a semantic layer without a context layer?
Yes - and most organizations do today. If your primary consumers are human analysts using BI tools, a semantic layer alone is sufficient. The context layer becomes essential when you introduce AI agents that need to understand not just what a metric means but whether and how they are allowed to use it.
Is a context layer the same as a data catalog?
No. A data catalog is a component of a context layer. The catalog inventories data assets and stores metadata. The context layer activates that metadata by delivering it to AI agents at query time through APIs and MCP connections. Modern platforms like Atlan extend catalog functionality into full context layer infrastructure.
Which tool implements a context layer?
Purpose-built context layer platforms include Decube, which combines catalog, lineage, quality, and governance into a metadata layer that delivers context to AI agents via MCP. You can also build a context layer on custom infrastructure using a vector database (for semantic search), a knowledge graph
How long does it take to implement a context layer?
Most enterprise context layer implementations take 8–16 weeks when using a purpose-built platform like Atlan. Building from scratch on custom infrastructure typically takes 6–12 months. The timeline depends heavily on how much governance metadata already exists and how many data sources need to be connected.
What is Data Context?
Data Context is the information that explains what data means, where it comes from, how it is transformed, whether it can be trusted, and how it should be used. It combines metadata, lineage, data quality, and governance so people and systems can confidently use data for analytics, reporting, and AI.
How is Data Context different from metadata?
Metadata describes data, while Data Context makes data usable and trustworthy. Metadata provides definitions, ownership, and technical details. Data Context extends this by adding lineage, quality signals, and governance rules, creating a complete, operational understanding of data.
Why is Data Context important for AI?
AI systems require Data Context to interpret data correctly, safely, and reliably. Without context, AI models may misunderstand metrics, use stale or incorrect data, or expose sensitive information. Data Context ensures AI uses trusted, well-defined, and policy-compliant data.
How does data lineage contribute to Data Context?
Data lineage provides visibility into how data flows and transforms across systems. It shows upstream sources, downstream dependencies, and transformation logic, enabling impact analysis, root-cause investigation, and confidence in reported numbers.
How do organizations build Data Context in practice?
Organizations build Data Context by unifying metadata, lineage, observability, and governance into a single operational layer. This includes defining business meaning, capturing end-to-end lineage, monitoring data quality, and enforcing usage policies directly within data workflows.
What is Context Engineering?
Context Engineering is the practice of designing and operationalizing business meaning, data lineage, quality signals, ownership, and policy constraints so that both humans and AI systems can reliably understand and act on enterprise data. Unlike traditional metadata management, Context Engineering focuses on decision-grade context that can be consumed programmatically by AI agents in real time.
How is Context Engineering different from prompt engineering?
Prompt engineering focuses on how questions are phrased for an AI model, while Context Engineering focuses on what the AI system already knows before a question is asked. In enterprise environments, context includes data definitions, lineage, quality, and usage constraints—making Context Engineering foundational for trustworthy and scalable Agentic AI.
Why is Context Engineering critical for Agentic AI?
Agentic AI systems reason, decide, and act autonomously across multiple systems. Without engineered context—such as trusted data meaning, lineage, and real-time quality signals—agents cannot assess risk or impact correctly. Context Engineering ensures AI agents act safely, explain decisions, and know when to pause or escalate.
What are the core components of Context Engineering?
The four core components of Context Engineering are: Semantic context (business meaning and definitions) Lineage context (end-to-end data flow and dependencies) Operational context (data quality and reliability signals) Policy context (privacy, compliance, and usage constraints) Together, these form a unified context layer that supports enterprise decision-making and AI automation
How should enterprises prepare for Context Engineering?
Enterprises should follow a phased approach: Inventory critical data and trust gaps Unify metadata, lineage, quality, and policy into a single context layer Expose context through APIs for AI agent consumption By 2026, this foundation will be essential for deploying Agentic AI at scale with confidence and auditability.
How do you measure the ROI of a data catalog?
ROI is measured by comparing the quantifiable benefits (such as reduced data search time, fewer data quality issues, and lower compliance effort) against the total costs (implementation, licensing, and support). Typical metrics include time savings, productivity gains, and compliance cost reduction.
What is a data catalog and why is it important for ROI?
A data catalog is a centralized inventory of data assets enriched with metadata that helps users find, understand, and trust data across an organization. It improves data discovery, reduces search time, and enhances collaboration — all of which contribute to measurable ROI by cutting operational costs and accelerating insights.
How quickly can businesses see ROI after implementing a data catalog?
Time-to-value varies with deployment and adoption, but many organizations begin seeing measurable improvements in days to months, especially through faster data discovery and reduced compliance effort. Early wins in these areas can quickly justify the investment.
What factors should you include when calculating the ROI of a data catalog?
When calculating ROI, include: Implementation and training costs Recurring maintenance and licensing fees Savings from reduced data search and rework Compliance cost reductions Productivity and decision-making improvements This ensures a holistic view of both costs and benefits.
How does a data catalog support data governance and compliance ROI?
A data catalog enhances governance by classifying data, enforcing rules, and providing transparency. This reduces regulatory risk and compliance effort, leading to direct cost savings and stronger data trust.
What is data lineage?
Data lineage shows where data comes from, how it moves, and how it changes across systems. It helps teams understand the full journey of data—from source to final reports or AI models.
Why is data lineage important for modern data teams?
Data lineage builds trust in data by making it transparent and explainable. It helps teams troubleshoot issues faster, assess impact before changes, meet compliance requirements, and confidently use data for analytics and AI.
What are the different types of data lineage?
Common types of data lineage include: Technical lineage – Tracks data movement at table and column level. Business lineage – Connects data to business definitions and metrics. Operational lineage – Shows how pipelines and jobs process data. End-to-end lineage – Combines all of the above across systems.
Is data lineage only useful for compliance?
No. While data lineage is critical for audits and regulatory compliance, it is equally valuable for debugging data issues, impact analysis, cost optimization, and AI readiness.
How does data lineage help with data quality?
Data lineage helps identify where data quality issues originate and which reports or dashboards are affected. This reduces time spent on root-cause analysis and improves accountability across data teams.
What is Metadata Management?
Metadata management involves the management and organization of data about data to enhance data governance, data asset quality, and compliance.
What are the key points of Metadata Management?
Metadata management involves defining a metadata strategy, establishing roles and policies, choosing the right metadata management tool, and maintaining an ongoing program.
How does Metadata Management work?
Metadata management is essential for improving data quality and relevance, utilizing metadata management tools, and driving digital transformation.
Why is Metadata Management important for businesses?
Metadata management is important for better data quality, usability, data insights, compliance adherence, and improved accuracy in data cataloging.
How should companies evolve their approach to Metadata Management?
Companies should manage all types of metadata across different environments, leverage intelligent methods, and follow best practices to maximize data investments.
What is a data definition example?
A data definition example could be: “Customer: a person or entity that has made at least one purchase within the past year.” It clearly sets business meaning and inclusion criteria.
Why is data definition important in data governance?
It ensures everyone interprets data consistently, reducing ambiguity and improving compliance, reporting, and collaboration.
Who should own data definitions?
Ownership should be shared between business domain experts (for context) and data stewards (for technical accuracy).
How often should data definitions be reviewed?
Ideally quarterly or whenever there’s a structural change in business logic, data models, or product offerings.
What’s the difference between data definition and data catalog?
A data catalog inventories data assets; data definition explains what those assets mean. Combined, they create full visibility and trust.
Why is Data Lineage important for businesses?
Data Lineage provides transparency and trust in your data ecosystem. It helps organizations ensure data accuracy, simplify root-cause analysis during data quality issues, and maintain compliance with regulations like GDPR or SOX. By understanding data flows, teams can make faster, more reliable decisions and improve overall data governance.
What are the key components of Data Lineage?
The main components of Data Lineage include: Data Sources: Where the data originates (databases, APIs, files). Transformations: How data is processed or modified. Data Pipelines: The tools or systems that move data. Destinations: Where the data is stored or consumed (dashboards, reports, models). Metadata: The contextual details that describe each step in the data’s lifecycle.
How does Data Lineage support Data Governance and AI readiness?
Data Lineage acts as the foundation for strong data governance by providing visibility into data ownership, transformation logic, and usage. For AI initiatives, lineage ensures that models are trained on accurate and traceable data, making AI outputs more explainable and trustworthy. Platforms like Decube’s Data Trust Platform unify lineage with data quality and metadata management to help enterprises achieve AI readiness.
What tools are commonly used for Data Lineage?
Several tools help automate and visualize data lineage, such as Decube, Atlan, Alation, Collibra, and OpenLineage. These tools connect to data warehouses, ETL pipelines, and BI tools to automatically map relationships between datasets — saving time and reducing manual effort.
What is Data Lineage?
Data Lineage is the process of tracking how data moves and transforms across an organization — from its origin to its final destination. It shows where data comes from, how it changes through different systems or pipelines, and where it ends up being used. In short, data lineage helps you visualize the journey of your data.
What does “data context” mean?
Data context refers to the semantic, structural, and business information that surrounds raw data. It explains what data means, where it comes from, who owns it, and how it should be used.
What is a centralized LLM framework?
It’s an enterprise-wide system where all departments access AI through a shared platform, equipped with guardrails, context layers, and multimodal capabilities.
What are guardrails in AI?
Guardrails are controls—policies, access restrictions, and compliance checks—that ensure AI outputs are secure, ethical, and aligned with enterprise goals.
How does data context affect ROI in AI?
Models trained or prompted with contextualized data deliver outputs that are relevant, trustworthy, and actionable—leading to faster adoption and higher business value.
What is MCP (Model Context Protocol) and why does it matter?
MCP defines how models interact with external tools and data sources. Feeding it with strong context ensures the AI agent can act accurately and responsibly.
What is a Data Trust Platform in financial services?
A Data Trust Platform is a unified framework that combines data observability, governance, lineage, and cataloging to ensure financial institutions have accurate, secure, and compliant data. In banking, it enables faster regulatory reporting, safer AI adoption, and new revenue opportunities from data products and APIs.
Why do AI initiatives fail in Latin American banks and fintechs?
Most AI initiatives in LATAM fail due to poor data quality, fragmented architectures, and lack of governance. When AI models are fed stale or incomplete data, predictions become inaccurate and untrustworthy. Establishing a Data Trust Strategy ensures models receive fresh, auditable, and high-quality data, significantly reducing failure rates.
What are the biggest data challenges for financial institutions in LATAM?
Key challenges include: Data silos and fragmentation across legacy and cloud systems. Stale and inconsistent data, leading to poor decision-making. Complex compliance requirements from regulators like CNBV, BCB, and SFC. Security and privacy risks in rapidly digitizing markets. AI adoption bottlenecks due to ungoverned data pipelines.
How can banks and fintechs monetize trusted data?
Once data is governed and AI-ready, institutions can: Reduce OPEX with predictive intelligence. Offer hyper-personalized products like ESG loans or SME financing. Launch data-as-a-product (DaaP) initiatives with anonymized, compliant data. Build API-driven ecosystems with partners and B2B customers.
What is data dictionary example?
A data dictionary is a centralized repository that provides detailed information about the data within an organization. It defines each data element—such as tables, columns, fields, metrics, and relationships—along with its meaning, format, source, and usage rules. Think of it as the “glossary” of your data landscape. By documenting metadata in a structured way, a data dictionary helps ensure consistency, reduces misinterpretation, and improves collaboration between business and technical teams. For example, when multiple teams use the term “customer ID”, the dictionary clarifies exactly how it is defined, where it is stored, and how it should be used. Modern platforms like Decube extend the concept of a data dictionary by connecting it directly with lineage, quality checks, and governance—so it’s not just documentation, but an active part of ensuring data trust across the enterprise.
What is an MCP Server?
An MCP Server stands for Model Context Protocol Server—a lightweight service that securely exposes tools, data, or functionality to AI systems (MCP clients) via a standardized protocol. It enables LLMs and agents to access external resources (like files, tools, or APIs) without custom integration for each one. Think of it as the “USB-C port for AI integrations.”
How does MCP architecture work?
The MCP architecture operates under a client-server model: MCP Host: The AI application (e.g., Claude Desktop or VS Code). MCP Client: Connects the host to the MCP Server. MCP Server: Exposes context or tools (e.g., file browsing, database access). These components communicate over JSON‑RPC (via stdio or HTTP), facilitating discovery, execution, and contextual handoffs.
Why does the MCP Server matter in AI workflows?
MCP simplifies access to data and tools, enabling modular, interoperable, and scalable AI systems. It eliminates repetitive, brittle integrations and accelerates tool interoperability.
How is MCP different from Retrieval-Augmented Generation (RAG)?
Unlike RAG—which retrieves documents for LLM consumption—MCP enables live, interactive tool execution and context exchange between agents and external systems. It’s more dynamic, bidirectional, and context-aware.
What is a data dictionary?
A data dictionary is a centralized repository that provides detailed information about the data within an organization. It defines each data element—such as tables, columns, fields, metrics, and relationships—along with its meaning, format, source, and usage rules. Think of it as the “glossary” of your data landscape. By documenting metadata in a structured way, a data dictionary helps ensure consistency, reduces misinterpretation, and improves collaboration between business and technical teams. For example, when multiple teams use the term “customer ID”, the dictionary clarifies exactly how it is defined, where it is stored, and how it should be used. Modern platforms like Decube extend the concept of a data dictionary by connecting it directly with lineage, quality checks, and governance—so it’s not just documentation, but an active part of ensuring data trust across the enterprise.
What is the purpose of a data dictionary?
The primary purpose of a data dictionary is to help data teams understand and use data assets effectively. It provides a centralized repository of information about the data, including its meaning, origins, usage, and format, which helps in planning, controlling, and evaluating the collection, storage, and use of data.
What are some best practices for data dictionary management?
Best practices for data dictionary management include assigning ownership of the document, involving key stakeholders in defining and documenting terms and definitions, encouraging collaboration and communication among team members, and regularly reviewing and updating the data dictionary to reflect any changes in data elements or relationships.
How does a business glossary differ from a data dictionary?
A business glossary covers business terminology and concepts for an entire organization, ensuring consistency in business terms and definitions. It is a prerequisite for data governance and should be established before building a data dictionary. While a data dictionary focuses on technical metadata and data objects, a business glossary provides a common vocabulary for discussing data.
What is the difference between a data catalog and a data dictionary?
While a data catalog focuses on indexing, inventorying, and classifying data assets across multiple sources, a data dictionary provides specific details about data elements within those assets. Data catalogs often integrate data dictionaries to provide rich context and offer features like data lineage, data observability, and collaboration.
What challenges do organizations face in implementing data governance?
Common challenges include resistance from business teams, lack of clear ownership, siloed systems, and tool fragmentation. Many organizations also struggle to balance strict governance with data democratization. The right approach involves embedding governance into workflows and using platforms that unify governance, observability, and catalog capabilities.
How does data governance impact AI and machine learning projects?
AI and ML rely on high-quality, unbiased, and compliant data. Poorly governed data leads to unreliable predictions and regulatory risks. A governance framework ensures that data feeding AI models is trustworthy, well-documented, and traceable. This increases confidence in AI outputs and makes enterprises audit-ready when regulations apply.
What is data governance and why is it important?
Data governance is the framework of policies, ownership, and controls that ensure data is accurate, secure, and compliant. It assigns accountability to data owners, enforces standards, and ensures consistency across the organization. Strong governance not only reduces compliance risks but also builds trust in data for AI and analytics initiatives.
What is the difference between a data catalog and metadata management?
A data catalog is a user-facing tool that provides a searchable inventory of data assets, enriched with business context such as ownership, lineage, and quality. It’s designed to help users easily discover, understand, and trust data across the organization. Metadata management, on the other hand, is the broader discipline of collecting, storing, and maintaining metadata (technical, business, and operational). It involves defining standards, policies, and processes for metadata to ensure consistency and governance. In short, metadata management is the foundation—it structures and governs metadata—while a data catalog is the application layer that makes this metadata accessible and actionable for business and technical users.
What features should you look for in a modern data catalog?
A strong catalog includes metadata harvesting, search and discovery, lineage visualization, business glossary integration, access controls, and collaboration features like data ratings or comments. More advanced catalogs integrate with observability platforms, enabling teams to not only find data but also understand its quality and reliability.
Why do businesses need a data catalog?
Without a catalog, employees often struggle to find the right datasets or waste time duplicating efforts. A data catalog solves this by centralizing metadata, providing business context, and improving collaboration. It enhances productivity, accelerates analytics projects, reduces compliance risks, and enables data democratization across teams.
What is a data catalog and how does it work?
A data catalog is a centralized inventory that organizes metadata about data assets, making them searchable and easy to understand. It typically extracts metadata automatically from various sources like databases, warehouses, and BI tools. Users can then discover datasets, understand their lineage, and see how they’re used across the organization.
What are the key features of a data observability platform?
Modern platforms include anomaly detection, schema and freshness monitoring, end-to-end lineage visualization, and alerting systems. Some also integrate with business glossaries, support SLA monitoring, and automate root cause analysis. Together, these features provide a holistic view of both technical data pipelines and business data quality.
How is data observability different from data monitoring?
Monitoring typically tracks system metrics (like CPU usage or uptime), whereas observability provides deep visibility into how data behaves across systems. Observability answers not only “is something wrong?” but also “why did it go wrong?” and “how does it impact downstream consumers?” This makes it a foundational practice for building AI-ready, trustworthy data systems.
What are the key pillars of Data Observability?
The five common pillars include: Freshness, Volume, Schema, Lineage, and Quality. Together, they provide a 360° view of how data flows and where issues might occur.
What is Data Observability and why is it important?
Data observability is the practice of continuously monitoring, tracking, and understanding the health of your data systems. It goes beyond simple monitoring by giving visibility into data freshness, schema changes, anomalies, and lineage. This helps organizations quickly detect and resolve issues before they impact analytics or AI models. For enterprises, data observability builds trust in data pipelines, ensuring decisions are made with reliable and accurate information.

Table of Contents

Read other blog articles

Grow with our latest insights

Sneak peek from the data world.

Thank you! Your submission has been received!
Talk to a designer