/Scalable Systems And Methods For Context-aware Sensitive Data Detection, Hierarchical Labeling, And Protection In Natural Language Processing Environments
Abstract

This disclosure relates to scalable systems and methods for detecting, labeling, and protecting sensitive data in natural language processing (NLP) environments. This includes NLP applications in artificial intelligence (AI) systems, such as language models (LMs) and generative AI (GenAI). More particularly, the present disclosure introduces a hierarchical, context-aware labeling mechanism that is optimized using an LM in conjunction with machine learning (ML) techniques to ensure the utility-preserving effective protection of sensitive data with, for example, minimal false positives and false negatives and/or optimal precision and recall. This disclosure also relates to data privacy, security, and context-sensitive access control, particularly for AI inference and data retrieval processes involving sensitive information in unstructured data. This disclosure also encompasses detection, labeling, tokenization, classification, provenance tracking, and selective disclosure enforcement, and extends to database-layer query authorization for sensitive unstructured data. The techniques disclosed herein may be applied to structured, semi-structured, and unstructured data formats.

Full Text

What is claimed is:

This disclosure relates to scalable systems and methods for detecting, labeling, and protecting sensitive data in natural language processing (NLP) environments. This includes NLP applications in artificial intelligence (AI) systems, such as language models (LMs) and generative AI (GenAI). More particularly, the present disclosure introduces a hierarchical, context-aware labeling mechanism that is optimized using an LM in conjunction with machine learning (ML) techniques to ensure the utility-preserving effective protection of sensitive data with, for example, minimal false positives and false negatives and/or optimal precision and recall. This disclosure also relates to data privacy, security, and context-sensitive access control, particularly for AI inference and data retrieval processes involving sensitive information in unstructured data. This disclosure also encompasses detection, labeling, tokenization, classification, provenance tracking, and selective disclosure enforcement, and extends to database-layer query authorization for sensitive unstructured data. The techniques disclosed herein may be applied to structured, semi-structured, and unstructured data formats.
Timeline
Filed
03/26/2026
Published
07/30/2026
Granted
Not Available
IPC Codes(1)
G06F 21/62:Protecting access to data via a platform, e.g. using keys or access control rules