Preprint / Version 1

Text-driven Video Manipulation with GPT4V

##article.authors##

  • John Feng Virginia Tech

DOI:

https://doi.org/10.31224/3402

Abstract

Generative Adversarial Networks (GANs) have revolutionized image synthesis including faces and scenes, with recent style-based generative models. Traditionally, researchers used the Gaussian distribution as the input latent code and then generate images. Recently, GAN inversion techniques are utilized to retrieval the latent code by the pretrained model from the real images. After that, the GAN generator can reconstruct the image using the estimated latent code, which provides challenges and opportunities for the image manipulation task. In our project, we work on semantic manipulation meaning changing the semantic attributes of an object.

Downloads

Download data is not yet available.

Downloads

Posted

2023-12-29