Text-driven Video Manipulation with GPT4V
DOI:
https://doi.org/10.31224/3402Abstract
Generative Adversarial Networks (GANs) have revolutionized image synthesis including faces and scenes, with recent style-based generative models. Traditionally, researchers used the Gaussian distribution as the input latent code and then generate images. Recently, GAN inversion techniques are utilized to retrieval the latent code by the pretrained model from the real images. After that, the GAN generator can reconstruct the image using the estimated latent code, which provides challenges and opportunities for the image manipulation task. In our project, we work on semantic manipulation meaning changing the semantic attributes of an object.
Downloads
Downloads
Posted
License
Copyright (c) 2023 John Feng

This work is licensed under a Creative Commons Attribution 4.0 International License.