experimenting with regional prompting on the Hunyuan video model, giving some inception vibes
left side prompt: cyberpunk & pan left
right side prompt: steampunk & pan right
Just published a set of ComfyUI nodes to use Genmo's Mochi to edit videos.
github.com/logtd/ComfyUI-…
It uses rf-inversion, the gift that keeps on giving.
Been revisiting Reference-Only Control for Flux. It uses the diffusion model as a pseudo image encoder on a reference image to influence the generation.
Results are somewhere between style and content transfer.
RAVE and FLATTEN were two of the papers that originally got me into diffusion models. They take inverse noise and apply consistency to image models.
Now with RF-Inversion (thanks @litu_rout_ and @natanielruizg) I can try these on Flux.
Not production quality, but still fun.