While generative models have advanced significantly in synthesizing high-fidelity images, precise control over their outputs remains a fundamental challenge. Early breakthroughs like Generative Adversarial Networks (GANs) demonstrated great success in generating natural images, but manipulating specific semantic attributes (e.g., object placement, time of day, or camera pose) was highly constrained. The advent of diffusion models introduced intuitive control via natural language prompts. However, the underlying mechanics of this control remain opaque. It is not entirely clear how text prompts, once translated into embeddings, dictate the generation process, nor is it understood exactly what information the image model requires from these text encoders. In this research, we investigate these mechanisms to advance the interpretable and efficient control of generative models.
Nurit Spingarn is a PhD candidate under the supervision of Prof. Tomer Michaeli.