Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Diptych Prompting (CVPR 2025)


arXiv

Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
Chaehun Shin, Jooyoung Choi, Heeseung Kim, Sungroh Yoon, Data Science and Artificial Intelligence Lab, Seoul National University


Introduction

We introduce Diptych Prompting, a novel zero-shot approach that reinterprets as an inpainting task with precise subject alignment by leveraging the emergent property of diptych generation in large-scale text-to-image models. Diptych Prompting arranges an incomplete diptych with the reference image in the left panel, and performs text-conditioned inpainting on the right panel. We further prevent unwanted content leakage by removing the background in the reference image and improve fine-grained details in the generated subject by enhancing attention weights between the panels during inpainting.

How to Use

Note

Memory consumption

Flux can be quite expensive to run on consumer hardware devices and as a result this implementation comes with high memory requirements exceeding 40GB of VRAM .

Setup (Optional)

  1. Environment setup
conda create -n diptychprompting python=3.10
conda activate diptychprompting
  1. Requirements installation
conda install pytorch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 pytorch-cuda=12.1 -c pytorch -c nvidia
pip install -r requirements.txt

Usage example

python diptych_prompting_inference.py --input_image_path /path/to/your-image --subject_name "name of your subject" --target_prompt "text prompt you want to generate"

For example,

python diptych_prompting_inference.py --input_image_path ./assets/bear_plushie.jpg --subject_name "bear plushie" --target_prompt "a bear plushie riding a skateboard"

Awesome Concurrent Work

You can explore the outstanding concurrent work, In-Context LoRA.

In-Context LoRA: IC-LoRA trains a LoRA model to generate image sets with intrinsic relationships, and conditions the image generation process on another image set using the SDEdit inpainting approach.

Additionally, you might enjoy exploring the following community extensions based on In-Context LoRA:

License

Some of the implementations are based on FLUX-Controlnet-Inpainting, and thus parts of the code would apply under the license of it.

Citation

@article{shin2024largescaletexttoimagemodelinpainting,
      title={Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator}, 
      author={Chaehun Shin and Jooyoung Choi and Heeseung Kim and Sungroh Yoon},
      journal={arXiv preprint arXiv:2411.15466},
      year={2024}
}

About

DipthchPrompting for subject-driven image editing

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages