UPDF AI

Multi-caption Text-to-Face Synthesis: Dataset and Algorithm

Jianxin Sun,Qi Li,2 Authors,Zhenan Sun

2021 · DOI: 10.1145/3474085.3475391
ACM Multimedia · 52 Citations

TLDR

A Semantic Embedding and Attention (SEA-T2F) network that allows multiple captions as input to generate highly semantically related face images and introduces an attribute loss to guide the network to generate sentence-related attributes.

Abstract

Text-to-Face synthesis with multiple captions is still an important yet less addressed problem because of the lack of effective algorithms and large-scale datasets. We accordingly propose a Semantic Embedding and Attention (SEA-T2F) network that allows multiple captions as input to generate highly semantically related face images. With a novel Sentence Features Injection Module, SEA-T2F can integrate any number of captions into the network. In addition, an attention mechanism named Attention for Multiple Captions is proposed to fuse multiple word features and synthesize fine-grained details. Considering text-to-face generation is an ill-posed problem, we also introduce an attribute loss to guide the network to generate sentence-related attributes. Existing datasets for text-to-face are either too small or roughly generated according to attribute labels, which is not enough to train deep learning based methods to synthesize natural face images. Therefore, we build a large-scale dataset named CelebAText-HQ, in which each image is manually annotated with 10 captions. Extensive experiments demonstrate the effectiveness of our algorithm.