beta
/Representation Method, Search Method, Model Training Method, Device, And Storage Medium
Abstract

A method for generating a representation of a search object comprises: obtaining a textual description of a search object and an image of the search object; dividing the image of the search object into a plurality of image patches, performing feature extraction on the plurality of image patches respectively through an image feature extraction model, and obtaining a feature sequence of the plurality of image patches; extracting a pixel-level feature from the image of the search object through a text feature extraction model, and determining a representation vector of the search object based on the feature sequence, the pixel-level feature, and the image of the search object. Since the extracted feature sequence and pixel-level feature represent high-level and low-level features of the image, the determined features of the image of the search object are comprehensive and accurate, and the constructed representation vector of the search object is accurate.

Full Text

What is claimed is:

A method for generating a representation of a search object comprises: obtaining a textual description of a search object and an image of the search object; dividing the image of the search object into a plurality of image patches, performing feature extraction on the plurality of image patches respectively through an image feature extraction model, and obtaining a feature sequence of the plurality of image patches; extracting a pixel-level feature from the image of the search object through a text feature extraction model, and determining a representation vector of the search object based on the feature sequence, the pixel-level feature, and the image of the search object. Since the extracted feature sequence and pixel-level feature represent high-level and low-level features of the image, the determined features of the image of the search object are comprehensive and accurate, and the constructed representation vector of the search object is accurate.
Timeline
Filed
03/12/2026
Published
07/16/2026
Granted
Not Available
IPC Codes(5)
G06V 10/77:Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
G06N 3/088:Non-supervised learning, e.g. competitive learning
G06T 7/11:Region-based segmentation