A method of generating a video clip from a text description is provided. The method may include receiving the text description of the video clip to be generated. The method may include obtaining, based on the received text description, a vector representation of the text description. The method may include obtaining a vector representation of a sequence of key points of the video clip to be generated, by a diffusion motion model based on the obtained vector representation of the text description. The method may include mapping the vector representation of the sequence of key points to a series of key point images of the video clip being generated, wherein each key point image from the series corresponding to a respective frame of the video clip being generated. The method may include generating a frame sequence of the video clip.
Full Text
What is claimed is: