How Organizations Can Move From Static Data Classification To Data Trust

1 hour ago 2

Saurabh Gupta - Technology, AI and Innovation Leader.

getty

Many data governance practices were built for a much slower and more predictable world, where data was easier to locate, classify and protect because it stayed closer to the systems where it was created.

Today, in the ​AI-first era, data moves across cloud platforms, applications, analytics tools, reports, logs, APIs and AI systems, where it is copied, enriched, combined and reused for machine learning and generative AI. ​

Many companies try to solve these challenges by adding more data to their systems. However, in my experience, the more important fix is ensuring that data is imbued with trusted context.

How Context Impacts Data Governance In The AI Era ​

Adding context can make your data more useful. An attribute like “address” or “postal_code” may look limited, but it can reveal identity or financial behavior when combined with customer name, email, account number, amount, merchant, time stamp, device or IP data. ​

Perhaps more importantly, without the context, systems don't have the necessary information to approve access, apply masking requirements, highlight retention risk or prevent sensitive data from being included in AI workflows.​

In working with enterprises, I have seen datasets labelled as "confidential" without any context about where the data is sourced from, who owns it, who can access it or whether it is approved for AI use. ​​​​​

This lack of context creates a governance gap that is one of the leading causes of AI-related data breaches, with IBM finding that 97% of organizations that suffered a data-related breach said they lacked proper access controls​.

This governance gap can impact nearly every stakeholder in your organization.

Privacy teams need to know where personal and regulated data exists. Compliance teams need proof that policies are applied consistently. Security teams need to know which data creates the highest exposure risk. AI teams need to know whether data is safe, accurate, permitted and explainable before it is used in models, prompts, retrieval systems or automated decisions.​

The Challenges Of ​Static Classification

While static classification remains essential, data changes constantly in modern environments: schemas evolve, reports are exported and replicated, access shifts and sensitive data can surface in logs, prompts, embeddings or generated responses.

With static classification, a table marked confidential may be properly controlled inside one governed source system, but the risk changes when the same data is copied into another system, exported into a report, joined with customer identifiers or indexed for search and AI retrieval.

The label may stay the same, but the exposure, users, controls and audit needs have changed.

How Classification Can Build Data Trust

​When updating your data classification, the key is not to abandon rules, but to give classification enough context to support governance decisions.

In other words, the goal is to build data trust, which is the measurable confidence that data can be safely, responsibly and compliantly used for a specific business or AI purpose, established through a holistic evaluation of governance context.​

AI-powered contextual classification can help with this process by connecting data to context, ownership, lineage, access, policy, risk and evidence to enable data trust.

AI-powered contextual classification combines metadata intelligence, contextual relationships, governance policies and AI reasoning to determine whether data can be trusted for a specific business purpose.

This matters because sensitivity is not always obvious from field names or values. Some data is inherently sensitive, such as payment details, health information or contact information. Other data becomes sensitive only when combined with identity, financial activity, customer behavior or regulated processes.

LLMs and GenAI can complement and strengthen deterministic checks and help interpret ambiguous metadata, summarize why data was flagged, compare intended use with policy language and prepare review notes.

That said, like with nearly every other application of AI, AI-powered classification needs both guardrails and human oversight.

Also, when implementing AI, organizations can run into issues with the data environment around the AI model that have nothing to do with the model itself, such as when metadata can be incomplete and field names are misleading. In these cases, models can make confident but wrong suggestions that can only be solved by cleaning up the data. ​

What's Required For Contextual Classification

​When building a classification system that provides the appropriate context for the AI era, four strategies can help ensure that the classification system is properly functioning:

1. Establish control decisions to build data trust. If personal information appears in a report, the response may be masking, restricted access or retention review. If the same data is exposed to a GenAI retrieval workflow, the response may require stronger approval, prompt safeguards, output monitoring or exclusion from retrieval.

2. Provide risk scoring to differentiate urgent from low sensitivity exposures. A practical control and risk scoring considers sensitivity, identifiability, access breadth, protection level, retention status, policy gaps, prompt exposure, retrieval exposure and model-use context.​

3. Monitor all decisions. Data, access and prompts change, embeddings are refreshed, and new attributes emerge. Organizations can leverage AI-assisted monitoring to track model drift, audit access permissions, flag over-retained data and quickly isolate sensitive data exposure across active AI workflows.

4. Mandate evidence. Data leaders should have a clear record of what was discovered, why it was classified, what risk was assigned, what controls were recommended or applied, how exceptions were reviewed and whether the data is safe for a specific use.​​​​

The Future Is Data Trust

AI and data governance depend on data foundations that are trusted, explainable, policy-aware and continuously governed. As new data sources, regulations, threats and AI use cases continue raising the standard for responsible governance, stronger AI models will need to detect sensitive data, understand context, apply controls, prioritize risk and prove responsible use.

The defining question will no longer be only what data is sensitive, but whether that data can be trusted for safe and responsible AI use.

Disclaimer: Opinions expressed here belong solely to the author and do not reflect the views of their employer.​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?


Read Entire Article