\text{Ph.D. Dissertation}
\textbf{Naresh Kumar Devulapally}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Dr. Vishnu Lokhande }\textit{(Ph.D. Supervisor)}
\text{Dr. Junsong Yuan}
\text{Dr. Siwei Lyu}
\textbf{\underline{Committee Members:}}
\textbf{Controlling Diffusion Models Across Modalities:}
\textit{Aligning Generation with Human Intent for:}
\textit{AI Provenance, Safety, Personalization, and Reliability}
\text{July 10, 2025}
\text{GenAI has seen immense progress lately}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

Google Veo 3

\text{July 10, 2025}
\text{Startups based on Diffusion Models!}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

Startups working on Diffusion Models

\text{July 10, 2025}
\text{Need to control GenAI models for human intent!}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

AI Provenance

NSF and DARPA

mention AI control as core GenAI challenge

\text{July 10, 2025}
\text{Controllability in Diffusion Models - Human Intents}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance}
\text{Controlling}
\text{Diffusion Models}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{Safety}
\textit{generation}
\textbf{Personalization}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{Reliability}
\textit{Mitigate factually}
\textit{incorrect generation}
\textit{Unlearnable data}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Controlling}
\text{Diffusion Models}
\textbf{Desirable characteristics of techniques to control Diffusion Models:}
\textbf{\textit{Training-free}}
\textbf{\textit{inference-time steering}}
\textit{Parameter-efficient finetuning}
\textbf{\textit{Generalization across}}
\textbf{\textit{diffusion architectures}}
\textbf{\textit{Internal/External signal guided}}
\textbf{\textit{intervention}}
\textbf{\textit{Aligning Generation with Human Intent for:}}
\textbf{\textit{AI Provenance, Safety, Personalization, and Reliability}}
\text{Preliminaries on Diffusion Models - Image Generation}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Preliminaries on Diffusion Models - dLLMs}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

Autoregressive Formulation in Traditional LLMs

p_\theta(x) = p_\theta(x^1) \prod_{i=2}^{L} p_\theta(x^i \mid x^1, \ldots, x^{i-1})

Each token prediction depends on previous tokens.

Sequential sampling

No parallelism

Error compounding

Sequential sampling

\text{Preliminaries on Diffusion Models - dLLMs}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

Large Language Diffusion Models - Loss

\mathcal{L}(\theta) \triangleq - \mathbb{E}_{t, x_0, x_t} \left[ \frac{1}{t} \sum_{i=1}^{L} \mathbf{1}[x_t^i = \text{M}] \log p_\theta(x_0^i \mid x_t) \right]

LLaDA’s core is a mask predictor: a model \( p_\theta(\cdot \mid x_t) \) that takes a masked sequence \( x_t \) and predicts all masked tokens (set \( M \)) simultaneously.

Pre-training

Supervised Finetuning

Sampling

\text{Control knobs in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Modifying Text Embeddings}
\text{Control knobs in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Modifying Initial Noise}
\text{Control knobs in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Modifying Trajectory}
\text{using reward signal}
\text{Control knobs in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Modifying Trajectory}
\text{using internal signal}
\text{July 10, 2025}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance}
\text{Controlling}
\text{Diffusion Models}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{Safety}
\textit{generation}
\textbf{Personalization}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{Reliability}
\textit{Mitigate factually}
\textit{incorrect generation}
\textit{Unlearnable data}
\text{July 10, 2025}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Controlling}
\text{Diffusion Models}
\textbf{AI Provenance}
\textit{In-generation}
\textit{object-level watermarking}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\text{Watermark Embedder } (E_{\theta})
\text{Message } (m)
\text{Predicted } (m')

We want to match

these by training \( E_\theta \)

Blind Watermarking: No additional metadata other than a watermarked image is required for detection.

Input Image \( I_0 \)

Watermarked Image \( I_w \)

\text{Watermark Detector } (D_{w})
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

Payload

Robustness

Invisibility

\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

Payload

Robustness

Invisibility

  • Watermarking within T2I generation.
  • Object Level watermark localization.
  • Towards training-free methods.
  • Natural control for watermarking.

Additional constraints while using

Generative models for watermarking.

\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
  • Watermarking within T2I generation.
  • Object Level watermark localization.
  • Towards training-free methods.
  • Natural control for watermarking.
\textcolor{blue}{\text{Integrate Watermarking into the generation pipeline}}
\textcolor{blue}{\text{Watermark localization during image generation}}
\textcolor{blue}{\text{Very low parameter requirement}}
\textcolor{blue}{\text{User facing components during generation}}
\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textcolor{green}{\text{Integrate Watermarking into the generation pipeline}}
\textcolor{green}{\text{Watermark localization during image generation}}
\textcolor{green}{\text{Very low parameter requirement}}
\textcolor{green}{\text{User facing components during generation}}
\textcolor{blue}{\text{Text controlled}}
\textcolor{blue}{\text{768 params}}
\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
L_w = \text{BCE}(D_w(\text{Dec}(z'_0, w))), m)
\mathcal{D}_w - \textcolor{blue}{\text{Message Detector}}
z'_0 - \textcolor{blue}{\text{Watermarked Latent}}
m - \textcolor{blue}{\text{Watermark Key}}
\text{(Binary string) i.e., } m \in \{0, 1 \}^k
\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
L_z = \min_{\mathcal{W}_*} \mathbb{E}_t \left[ \| z_t^* - z_t(\mathcal{W}_*) \|^2 \right]
z_t - \textcolor{blue}{\text{Latent at }} t
\mathcal{W}_* - \textcolor{blue}{\text{Watermark token}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

Datasets:

MS COCO, WikiArt

Image similarity:

PSNR, SSIM, FID

Robustness Metrics:

Basic and Adversarial Attacks

\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

 Multiple Object Watermarking

\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}

\( \text{Naresh Kumar Devulapally} \)

devulapa@buffalo.edu

\( \text{The State University of New York} \)

\( \text{at Buffalo} \)

\( \text{Shruti Agarwal} \)

shragarw@adobe.com

\( \text{Research Scientist} \)

\( \text{Adobe Research} \)

\( \text{Vishnu Suresh Lokhande} \)

vishnulo@buffalo.edu

\( \text{The State University of New York} \)

\( \text{at Buffalo} \)

\( \text{Mingzhen Huang} \)

mhuang33@buffalo.edu

\( \text{The State University of New York} \)

\( \text{at Buffalo} \)

\( \text{Adobe Research} \)

\( \text{Vishal Asnani} \)

vasnani@adobe.com

\( \text{Research Scientist} \)

\( \text{Adobe Research} \)

\( \text{Siwei Lyu} \)

siweilyu@buffalo.edu

\( \text{The State University of New York} \)

\( \text{at Buffalo} \)

\textbf{AI Provenance:}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{\textit{ICCV 2025}}
\text{July 10, 2025}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance}
\text{Controlling}
\text{Diffusion Models}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{Safety}
\textit{generation}
\textbf{Personalization}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{Reliability}
\textit{Mitigate factually}
\textit{incorrect generation}
\textit{Unlearnable data}
\text{July 10, 2025}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Controlling}
\text{Diffusion Models}
\textbf{Safety}
\textit{generation}
\textit{Unlearnable data}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textit{Unlearnable Data Generation}
\textbf{\textit{ACMMM 2025}}
  • Diffusion models (e.g., Stable Diffusion) can adapt to new subjects or styles with just a few images. Textual Inversion and DreamBooth achieve high-fidelity personalization rapidly.

Power of Personalization

Risks and Concerns

  • Privacy: Risk of misuse of private or identity-sensitive data.

  • Intellectual Property: Unauthorized style or content cloning.

  • Security: Uncontrolled replication of copyrighted material

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}
  • Pixel-space perturbations make samples unlearnable. However...
    • Visibly degrade quality (noise/artifacts).
    • Vulnerable to purification defenses (e.g., DiffPure).

Limitations of Existing Defenses

Proposed Method

  • Shift perturbations into latent space, altering the diffusion trajectory.
  • Ensure perturbations are imperceptible yet robust to personalization and purification.
  • Practical and scalable defense for safeguarding sensitive data.
\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}

Approach: Introduce adversarial noise perturbations to ensure that certain images become unlearnable.

Goal: Prevent a diffusion model from learning and reproducing specific images while maintaining its generalization capability.

\min_{\|\delta^{u}\| \leq \rho_u} D(p_{\theta}^{u}(x), q(x))
\text{where } x \sim q(x), \|\delta\| \leq \rho_u
\text{s.t.} \max_{\delta^*} \mathbb{E}_{t, x' \sim p_{\theta}^{u}(x), x \sim q^{c}(x)} [ \mid \mid \mathcal{L}_{DM}(x') - x \mid \mid_2^2]
\delta^* =\argmin_\delta D(p_{\theta}^{u}(x), q(x))
\text{where } p_{\theta}^{u}(x) \sim x + \delta_\theta
\textcolor{green}{\text{Unlearnable for LDM}}
\textcolor{green}{\text{Imperceptible}}
\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}

We consider Latent Diffusion Models for Image Generation

Textual Inversion and DreamBooth as Personalization models

\mathcal{L}_{\text{DM}} = \; \mathbb{E}_{\substack{x,y,\epsilon \sim \mathcal{N}(0,1),t}} \Bigl[ \|\epsilon - \epsilon_\theta\bigl(x_t,\, t,\, \mathcal{T}(y)\bigr)\|_2^2 \Bigr]

Diffusion Noise-Matching Loss

\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}

We consider Latent Diffusion Models for Image Generation

Textual Inversion and DreamBooth as Personalization models

\mathcal{L}^{\text{TI}}_{\text{personalize}} = \; \mathbb{E}_{\substack{x,\epsilon \sim \mathcal{N}(0,1),t}} \Bigl[ \|\epsilon - \epsilon_\theta\bigl(x_t,\, t,\, \mathcal{T}_{\theta}(S_*)\bigr)\|_2^2 \Bigr]

Textual Inversion Loss

\mathcal{L}^{\text{DB}}_{\text{personalize}} = \mathbb{E}_{x_p,\, x_{\text{cls}},\, \epsilon,\, t} \left[ \left\| \epsilon - \epsilon_\theta(x_p^t, t, c_p) \right\|^2 + \lambda \left\| \epsilon - \epsilon_\theta(x_{\text{cls}}^t, t, c_{\text{cls}}) \right\|^2 \right]

DreamBooth Loss

\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}

We consider Latent Diffusion Models for Image Generation

Textual Inversion and DreamBooth as Personalization models

\begin{aligned} &\text{maximize } \; \mathcal{L}_{\text{personalize}}(\bar{z}^{\text{ul}}_0) \\ &\text{subject to } \; \| \bar{x}^{\text{ul}}_0 - x_0 \| \leq \delta^u \end{aligned}
\mathcal{L}_{\text{personalize}}(\bar{z}^{\text{ul}}) = \mathbb{E}_{\bar{z}^{\text{ul}},\, t} \Bigl[ \| \epsilon - \epsilon_\theta(\bar{z}^{\text{ul}}_t,\, t,\, \mathcal{T}(y)) \|_2^2 \Bigr].

where the personalization loss at the latent level is defined as

\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}

We consider Latent Diffusion Models for Image Generation

Textual Inversion and DreamBooth as Personalization models

\mathcal{L}_{\text{personalize}}(\bar{z}^{\text{ul}}) = \mathbb{E}_{\bar{z}^{\text{ul}},\, t} \Bigl[ \| \epsilon - \epsilon_\theta(\bar{z}^{\text{ul}}_t,\, t,\, \mathcal{T}(y)) \|_2^2 \Bigr].
\textcolor{green}{\mathcal{L} = \lambda \cdot (|| (\bar{x}^{\text{ul}}_0) - (x_0) || - \delta^u) + (-\mathcal{L}_{\text{personalize}}( \bar{z}^{\text{ul}}_0))}

Overall Loss:

\( \lambda \) - hyper-parameter

\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}

Few-step Diffusion Models maintain data distribution integrity while allowing perturbations in reconstruction:

e_L = \Delta(\text{Curve Fit}, \text{Data})
e_R = \Delta(\Phi^{\text{denoise}}, \text{Data})
\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}

Quantitative results comparison to baseline methods: Metrics: (1) Imperceptibility (PSNR, SSIM, FID), (2) Face-specific metrics (ISM, BRIS, FDFR), (3) Non-Face metrics (SDS, ISM, BRIS, IQAC, LIQE). The metrics presented utilize  \( 10/255 \) as the maximum budget, with \( 15 \) steps of DiffPure purification after Unlearnable Sample Generation and TI training with the prompt ``A photo of an sks person/object'' for personalization. We generate \( 16 \) images per prompt post personalization to report personalization metrics. We notice from the above table that our method shows significant improvements in imperceptible perturbation at the image-level while maintaining enhanced counter personalization metrics.

\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}
\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}
\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}
  • All Current unlearnable sample methods = static: perturbations crafted once after personalization model is fixed.

  • But real-world personalization is continual → models adapt incrementally.

  • Need defenses that also adapt continually to evolving personalization systems.

This min-max formulation can be written as:

\min_{\rho} \max_{\tau_i} \; \mathcal{L}^{\text{TI}}(x_0, \tau_i; \rho) - \lambda \cdot \mathcal{L}_{\text{unlearn}}(\bar{x}_0^{\text{ul}}, \tau_i)

This setup better reflects real-world usage, where users may upload new training data, and personalization systems are incrementally updated.

\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Safety:}
\textbf{\textit{ACMMM 2025}}

\( \text{Naresh Kumar Devulapally} \)

devulapa@buffalo.edu

\( \text{The State University of New York} \)

\( \text{at Buffalo} \)

\( \text{Shruti Agarwal} \)

shragarw@adobe.com

\( \text{Research Scientist} \)

\( \text{Adobe Research} \)

\( \text{Tejas Gokhale} \)

gokhale@umbc.edu

\( \text{University of Maryland} \)

\( \text{Baltimore County} \)

\( \text{Vishnu Suresh Lokhande} \)

vishnulo@buffalo.edu

\( \text{The State University of New York} \)

\( \text{at Buffalo} \)

\textit{Unlearnable Data Generation}
\text{July 10, 2025}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance}
\text{Controlling}
\text{Diffusion Models}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{Safety}
\textit{generation}
\textbf{Personalization}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{Reliability}
\textit{Mitigate factually}
\textit{incorrect generation}
\textit{Unlearnable data}
\text{July 10, 2025}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Controlling}
\text{Diffusion Models}
\textbf{Personalization}
\textit{Condition-to-output}
\textit{faithfulness}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

What is Prompt Inversion?

Existing hard prompt inversion methods produce prompts that are either incoherent or too brittle under downstream token edits, so we invert a reference image into a coherent, edit-robust prompt that reliably reconstructs it via a T2I model, sidestepping the slow trial-and-error of manual prompt engineering.

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}
  • Prompt engineering for text-to-image (T2I) models is a slow trial-and-error process.
  • Hard prompt inversion aims to recover a discrete text prompt that reconstructs an input reference image. However, existing inversion techniques optimize only for one-shot reconstruction.
  • Prompts by existing methods are incoherent (PEZ) or brittle; a small token edit can collapse the whole image.
  • Goal: prompts that are interpretable, aligned, and stay stable under token swap & append edits.

Interpretable and Edit-Friendly Prompt Inversion

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Edit-friendly objective. Hard inversion that directly optimizes alignment under token swap/append edits.

CLIP-guided discrete diffusion. A dLLM decoder steered by CLIP for faithful, fully-readable prompts.

Fast & plug-and-play. ~10x faster than baselines; no fine-tuning of the dLLM or any T2I model.

Contributions

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Pipeline

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Method

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Method

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Results - Quantitative

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Results - Prompt Inversion comparison

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Results - Concept Swap

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Results - Concept Append

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

Results - TIFA and LLM as a Judge

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization:}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{\textit{CVPR 2026}}

\( \text{Naresh Kumar Devulapally} \)

devulapa@buffalo.edu

\( \text{The State University of New York} \)

\( \text{at Buffalo} \)

\( \text{Shruti Agarwal} \)

shragarw@adobe.com

\( \text{Research Scientist} \)

\( \text{Adobe Research} \)

\( \text{Vishnu Suresh Lokhande} \)

vishnulo@buffalo.edu

\( \text{The State University of New York} \)

\( \text{at Buffalo} \)

\( \text{Adobe Research} \)

\( \text{Vishal Asnani} \)

vasnani@adobe.com

\( \text{Research Scientist} \)

\( \text{Adobe Research} \)

\text{July 10, 2025}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{AI Provenance}
\text{Controlling}
\text{Diffusion Models}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{Safety}
\textit{generation}
\textbf{Personalization}
\textit{Condition-to-output}
\textit{faithfulness}
\textbf{Reliability}
\textit{Mitigate factually}
\textit{incorrect generation}
\textit{Unlearnable data}
\text{July 10, 2025}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Controlling}
\text{Diffusion Models}
\textbf{Personalization}
\textbf{Concept}
\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization}
\textbf{Concept}
\textit{Collaboration work}

Incremental concept addition without forgetting

\textbf{\textit{WACV, AAAI 2026}}

Arjun Ramesh Koushik et al. 2026

\text{July 10, 2025}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Personalization}
\textbf{Concept}
\textit{Collaboration work}
\textbf{\textit{WACV, AAAI 2026}}

Arjun Ramesh Koushik et al. 2026

\textbf{AI Provenance}
\text{Controlling}
\text{Diffusion Models}
\textit{In-generation}
\textit{object-level watermarking}
\textbf{Safety}
\textit{generation}
\textbf{Reliability}
\textit{Mitigate factually}
\textit{incorrect generation}
\textit{Unlearnable data}
\textbf{Personalization}
\textbf{Concept}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Controlling}
\text{Diffusion Models}
\textbf{Reliability}
\textit{Mitigate factually}
\textit{incorrect generation}
\text{Controllability in Diffusion Models}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Reliability:}
\textit{Mitigate factually}
\textit{incorrect generation}

Internal-signal guided Hallucination Detection and Mitigation

\textbf{\textit{CoLM 2026}}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Reliability:}
\textit{Mitigate factually}
\textit{incorrect generation}

Internal-signal guided Hallucination Detection and Mitigation

\textbf{\textit{CoLM 2026}}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Reliability:}
\textit{Mitigate factually}
\textit{incorrect generation}

Internal-signal guided Hallucination Detection and Mitigation

\textbf{\textit{CoLM 2026}}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Reliability:}
\textit{Mitigate factually}
\textit{incorrect generation}
\textbf{\textit{ArXIV}}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Reliability:}
\textit{Mitigate factually}
\textit{incorrect generation}
\textbf{\textit{ArXIV}}
\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Reliability:}
\textit{Mitigate factually}
\textit{incorrect generation}
\textbf{\textit{ArXIV}}

Validator-guided Hallucination Mitigation in Diffusion LLMs and VLMs

ongoing work

\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Reliability:}
\textit{Mitigate factually}
\textit{incorrect generation}
\textbf{\textit{ArXIV}}

ongoing work

\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Reliability:}
\textit{Mitigate factually}
\textit{incorrect generation}
\textbf{\textit{ArXIV}}

ongoing work

May 2024

Jan 2025

Feb 2025

May 2025

Nov. 2025

Feb. 2026

May 2026

July 2026

Aug. 2026

...

Chair's Fellowship

CSE 573 Instructor

Teaching Award 🌟

CSE 573

Instructor

...

\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{PhD Trajectory so far...}

Research

Intern 👨🏻‍💻

\text{Naresh Kumar Devulapally}
\text{Dissertation Proposal}
\text{August 3, 2026}
\textbf{Questions?}

The floor is open to questions :)

Dissertation Proposal - Naresh

By Naresh Kumar Devulapally

Dissertation Proposal - Naresh

My PhD Dissertation Proposal slides

  • 368