Provided is a method for evaluating text-to-image generation, an electronic device and a storage medium, relating to the field of computer technology, and in particular to the fields of artificial intelligence, generative search, knowledge graph, text-to-image generation and others. The method includes: receiving an evaluation request instruction described in natural language and a target image to be evaluated, wherein the target image is an image correspondingly generated based on a text described in the natural language; when parsing the evaluation request instruction, decomposing the natural language description into multiple pieces of structured semantic information based on intent analysis; obtaining a set of visual question answering tasks for evaluation and verification based on the multiple pieces of structured semantic information in combination with visual analysis of the target image; and performing multi-dimensional evaluation according to the visual question answering tasks to obtain a quantitative evaluation result.
Full Text
What is claimed is: