In a method for video generation model training, at least two first sample pairs are obtained, each of the first sample pairs including a first sample speech segment and a first sample video segment, the first sample video segment of each of the first sample pairs including images of a respective virtual object, and a facial expression of the virtual object in the first sample video segment matching speech content of the first sample speech segment. In the method, a respective predicted video segment corresponding to the first sample speech segment is generated through a video generation model based on the first sample speech segment of each of the first sample pairs. In the method, the video generation model is trained based on the predicted video segment and the first sample video segment corresponding to the first sample speech segment. Apparatus and non-transitory computer-readable storage medium counterparts are also contemplated.
Full Text
What is claimed is: