This application provides a video grounding method. The method includes: obtaining a video data set including video data, where the video data includes a plurality of frames of images; separately obtaining a first video feature of the video data and a first text feature of the video data, where the first text feature includes a plurality of words describing each frame of image in the video data; then segmenting the first video feature, to obtain features of a plurality of video clips; mapping the plurality of words to the plurality of video clips, to obtain a text description corresponding to each video clip of the plurality of video clips; and training a video grounding model based on the text description corresponding to each video clip, to obtain a trained video grounding model.
Full Text
What is claimed is: