Abstract
In some embodiments, techniques for self-labeling to extract a representative set of samples from a large-scale set of unlabeled documents (e.g., a set that represents a distribution of the large-scale set) are provided. The samples of the representative set may then be used to classify the documents of the large-scale set.
Full Text
What is claimed is:
In some embodiments, techniques for self-labeling to extract a representative set of samples from a large-scale set of unlabeled documents (e.g., a set that represents a distribution of the large-scale set) are provided. The samples of the representative set may then be used to classify the documents of the large-scale set.
Timeline
Filed
05/07/2026Published
09/10/2026Granted
Not AvailableIPC Codes(2)
G06F 16/906:Clustering; Classification
G06F 16/93:Document management systems