A method for generating a representation of a search object comprises: obtaining a textual description of a search object and an image of the search object; dividing the image of the search object into a plurality of image patches, performing feature extraction on the plurality of image patches respectively through an image feature extraction model, and obtaining a feature sequence of the plurality of image patches; extracting a pixel-level feature from the image of the search object through a text feature extraction model, and determining a representation vector of the search object based on the feature sequence, the pixel-level feature, and the image of the search object. Since the extracted feature sequence and pixel-level feature represent high-level and low-level features of the image, the determined features of the image of the search object are comprehensive and accurate, and the constructed representation vector of the search object is accurate.
Full Text
What is claimed is: