Semantic-Aware Facial Image Synthesis from Natural Language Descriptions Using Joint Bi-LSTM and Generative Adversarial Networks
DOI:
https://doi.org/10.64751/ajaccm.2026.v6.n3.794Abstract
Recent advancements in deep learning have enabled the generation of realistic images directly from natural language descriptions. This paper presents a semantic-aware framework for text-to-face image synthesis using a joint Bidirectional Long Short-Term Memory (BiLSTM) network and Generative Adversarial Network (GAN). The proposed approach simultaneously trains the text encoder and image generator, allowing effective learning of semantic relationships between textual attributes and facial features. Initially, input descriptions are transformed into meaningful vector representations using Bi-LSTM, which are then utilized by the GAN to synthesize high-quality facial images. Unlike conventional methods that rely on separately trained text encoders, the proposed end-to-end architecture improves semantic consistency and visual realism. The model is trained on the CelebA dataset with corresponding facial descriptions and evaluated using similarity and image quality measures. Experimental results demonstrate improved face generation accuracy and better preservation of facial attributes, making the framework suitable for applications in forensic investigations, digital character creation, intelligent human-computer interaction, and public safety systems.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.







