Skip to content

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RefVFX: Tuning-free Visual Effect Transfer Across Videos

tuning_free.mov

Maxwell Jones1, Rameen Abdal2, Or Patashnik2, Ruslan Salakhutdinov1, Sergey Tulyakov2, Jun-Yan Zhu1, Kuan-Chieh "Jackson" Wang2

1 Carnegie Mellon University   ·   2 Snap Research

⚠️ Unofficial reimplementation. This repository is an independent reimplementation built at CMU from the arXiv preprint. Because the dataset, exact training scaffolding, and compute used here differ from the original setup, the results you reproduce may not exactly match those in the paper. Hopefully the results are still competitive though!!

📖 Introduction

TL;DR. RefVFX transfers a temporal visual effect (cakeify, dissolve, acid trip, stylization, …) from a reference effect video onto any target video or image with a single model. The reference is the spec: whatever happens in it is the effect that gets applied to your input, preserving the original motion. Under the hood it's a single LoRA on top of Wan2.1-FLF2V-14B-720P trained on the published refVFX_dataset (166 717 pairs, 2 898 effect types). The Video diffusion model takes in the reference video effect and input image/video, and outputs the a new video with the effect applied. See the Method section below for the conditioning units and CFG scheme.

This repo ships the full stack the paper describes: trainer, single-GPU and 4× USP-parallel inference, a Gradio demo, and the three independent dataset-generation pipelines.


Quickstart (inference setup)

git clone https://github.com/maxwelljones14/refVFX.git refVFX && cd refVFX

# 0) Install uv once (https://docs.astral.sh/uv/getting-started/installation/).
#    Skip if you already have it (`uv --version`).
curl -LsSf https://astral.sh/uv/install.sh | sh

# 1) One-time config — copy the template and point REFVFX_CACHE_DIR at a fast
#    disk with plenty of free space. Every other directory defaults under it.
#    `.env` only holds *directories*; per-run inputs (videos, prompts, LoRA path,
#    output path) are passed as CLI flags to the inference scripts.
cp .env.example .env
vim .env        
# At this stage just set the directory vars: REFVFX_CACHE_DIR (weights root) and
# REFVFX_MODEL_DIR. You'll fill in REFVFX_LORA_PATH / REFVFX_CAUSVID_LORA after
# downloading the weights below.

# 2) Install the trainer/inference env. `--seed` puts pip into the venv so
#    step 3's install_flash_attn.py (which shells out to pip) works.
uv venv --seed --python 3.10 .venv && source .venv/bin/activate
uv pip install -r refVFX_trainer/requirements_trainer.txt
uv pip install -e refVFX_trainer    # installs the vendored `diffsynth`

# 3) Optional (recommended for speed): install a prebuilt flash-attention wheel
#    matching your python/torch/CUDA. Without it, attention falls back to PyTorch
#    SDPA. The script detects your env and pulls the newest matching wheel from
#    https://github.com/mjun0812/flash-attention-prebuild-wheels/releases.
python refVFX_trainer/install_flash_attn.py

The block above also copies .env. REFVFX_CACHE_DIR is the single root for all model weights — every other path (REFVFX_QWEN_CACHE_DIR, REFVFX_LORAS_DIR, REFVFX_WAN_I2V_DIR, …) defaults under it. Make sure REFVFX_CACHE_DIR and REFVFX_MODEL_DIR are set before downloading the weights next.

Download the model weights

Now that the env is installed, huggingface-cli is available. Pull the two LoRA checkpoints from the Hugging Face Hub (the Wan2.1-FLF2V-14B-720P base model is fetched automatically by ensure_model.py on first run):

Weight Repo Required?
refVFX LoRA (our trained model) maxwelljones14/refVFX-LoRA required
CausVid few-step LoRA Kijai/WanVideo_comfy strongly recommended

Best results come from our refVFX LoRA stacked with the CausVid LoRA (6-step accelerated sampling), together with ref_effect_guidance — both are on by default in the launchers, so just keep them.

# Download into your weights directory — the same path you set as
# REFVFX_CACHE_DIR in .env. Public refVFX repo, so no login is needed.
export REFVFX_CACHE_DIR=/path/to/weights        # match your .env

huggingface-cli download maxwelljones14/refVFX-LoRA step-10000.safetensors \
    --local-dir "$REFVFX_CACHE_DIR"             # newer hub versions: `hf download ...`

huggingface-cli download Kijai/WanVideo_comfy \
    Wan21_CausVid_14B_T2V_lora_rank32.safetensors \
    --local-dir "$REFVFX_CACHE_DIR"

Set the LoRA paths and load .env

Point the two LoRA file variables in .env at the files you just downloaded (and set REFVFX_MODEL_DIR if you haven't already):

# in .env
REFVFX_LORA_PATH=$REFVFX_CACHE_DIR/step-10000.safetensors
REFVFX_CAUSVID_LORA=$REFVFX_CACHE_DIR/Wan21_CausVid_14B_T2V_lora_rank32.safetensors

.env is not auto-sourced for direct python invocations — load it once now, and again any time you change these variables. NOTE: best results come with the CausVid LoRA applied and ref_effect_guidance on (both defaulted, so keep them):

set -a; source .env; set +a

Running inference

The two bash launchers run the recommended default cakeify example — refVFX LoRA + CausVid few-step LoRA, 6 denoising steps, cfg_scale=6.0, cfg_scale_ref=2.0, at 480×832 / 33 frames. They take no CLI arguments: to customize a run, call the python entry point directly (a fully-annotated reference invocation with every flag is in the comment block at the bottom of each .sh file).

If you have a GPU with >= 45 Gb VRAM, simply run:

bash refVFX_inference/infer_refvfx.sh

For multi-GPU sequence-parallel inference (auto-detects visible GPUs):

bash refVFX_inference/infer_refvfx_usp.sh

For custom inputs (different reference video, prompt, resolution, seed, …), call the python files directly. After set -a; source .env; set +a:

python refVFX_inference/infer_refvfx.py [args]              # single-GPU

torchrun --nproc_per_node=[num_gpus] --master_port=29506 \  # multi-GPU
    refVFX_inference/infer_refvfx_usp.py [args]

See the comment block at the bottom of refVFX_inference/infer_refvfx.sh / infer_refvfx_usp.sh for the full list of supported flags with example values.

Running the Gradio demo

If you have a GPU with >= 80 Gb VRAM, run the demo single-GPU:

python refVFX_inference/inference_gradio.py --server_port 7860

IF you have a single A6000 only, launch with

python refVFX_inference/inference_gradio.py --vram_limit 20 --server_port 7860 --high_quality

For 4× <80 Gb GPUs (tested on 4×A6000 49 Gb), launch under torchrun — the script auto-detects WORLD_SIZE > 1 and patches the pipeline with Ulysses + Ring sequence parallelism (defaults to ulysses=2, ring=2):

PYTHONUNBUFFERED=1 torchrun --nproc_per_node=4 --master_port=29501 \
    refVFX_inference/inference_gradio.py --server_port 7860 --vram_limit 20 --high_quality

Same as the inference scripts, video size auto-drops from 480P (480×832) to 544×320 in the multi-GPU path to fit per-rank memory; only rank-0 hosts the web UI.


Repository tour

Directory (README link) What it does Where to look
refVFX_inference/ Single-GPU and 4×GPU USP-parallel inference + Gradio demo infer_refvfx.py, inference_gradio.py
refVFX_trainer/ LoRA fine-tuning of Wan2.1-FLF2V-14B-720P with the RefVFX conditioning units train_refvfx.py, refvfx_pipeline.py
dataset_generation/ Three pipelines that produce the three subsets of the published dataset per-subset READMEs below
└ I2V_pipeline_from_scratch/ LoRA-driven image-to-video subset (Wan-I2V-14B + community LoRAs) generate_lora_i2v_pipeline.py
└ V2V_pipeline_from_scratch/ Pose-hybrid video-to-video subset (VACE + first/last-frame conditioning) pose_first_last_hybrid.py, batch_v2v.py
└ scalable_code_based_video_generation/ CPU-only deterministic effects (21 spatial + 15 temporal) run_examples.py, generator.py

Method Details

RefVFX learns to transfer temporal visual effects from a reference video onto an arbitrary target video or image, without per-effect optimization.

Problem

Existing video-editing systems are either text-driven (which struggles to describe complex temporal effects like "the subject dissolves into pixels and re-forms") or keyframe-based (which requires manual setup and breaks for non-rigid effects). RefVFX uses a reference effect video as the specification: whatever happens in the reference is the effect to be applied.

Architecture

RefVFX builds on top of Wan2.1-FLF2V-14B-720P (a text + first/last-frame video diffusion transformer) and adds three conditioning units, all wired through a LoRA adapter on the DiT (refvfx_pipeline.py):

  1. Reference Video Encoder — VAE-encodes the reference effect video; its latents are concatenated width-wise onto the target latents inside the transformer. The transformer's attention then naturally cross-attends between the target frames and the reference effect.
  2. Control Video Embedder — VAE-encodes the input target video and applies per-frame masks (1.0 = copy as-is, 0.5 = transform, 0.0 = drop) to mix the original content into the y-conditioning channel.
  3. Image Embedder (CLIP) — embeds the explicit first and last frames into the CLIP-conditioning slot for FLF2V.

Classifier-free guidance is multi-axis: separate dropout probabilities and inference-time scales for the reference, the control video, and the text prompt, so users can rebalance "fidelity to the reference effect" vs. "fidelity to the input motion" at sample time.

Data

Three complementary subsets, all packaged in maxwelljones14/refVFX_dataset:

Subset Effects Pairs How it was built
code_based_edits 2 736 136 800 Procedural — 21 spatial × 15 temporal effect combos applied with masks
neural_v2v_data 114 22 922 Pose-hybrid V2V using VACE + Wan2.1-FLF2V on real motion
I2V_LoRA 48 6 995 Community Wan-I2V LoRAs run on Qwen-Image-generated portraits
Total 2 898 166 717

See dataset_generation/README.md for the full pipeline, including VLM-based filtering for the I2V and V2V subsets.

Training

Single-stage LoRA fine-tune on the combined dataset, weighted {code_based_edits: 0.33, neural_v2v_data: 0.51, I2V_LoRA: 0.16} by default (train_refvfx.sh). See refVFX_trainer/README.md for the full hparam table, multi-GPU recipe, and resume instructions.


Citation

@article{jones2026tuning,
  title={Tuning-free Visual Effect Transfer across Videos},
  author={Jones, Maxwell and Abdal, Rameen and Patashnik, Or and Salakhutdinov, Ruslan and Tulyakov, Sergey and Zhu, Jun-Yan and Wang, Kuan-Chieh Jackson},
  journal={arXiv preprint arXiv:2601.07833},
  year={2026}
}

Acknowledgments

  • DiffSynth-Studio (modelscope/DiffSynth-Studio, Apache 2.0) — vendored under refVFX_trainer/DiffSynth-Studio/. The upstream README lives there. We thank the ModelScope team for the Wan video pipeline and the training scaffolding.
  • VACE (ali-vilab/VACE, Apache 2.0) — vendored under dataset_generation/V2V_pipeline_from_scratch/VACE/ for pose extraction in the V2V data pipeline.
  • Wan2.1-FLF2V-14B-720P (Wan-AI) — base video diffusion model.
  • Qwen3-VL-8B-Instruct — VLM used by the dataset filtering pipelines.

License

Apache 2.0 (see refVFX_trainer/LICENSE). Vendored dependencies retain their original licenses.

About

Reference-guided VFX: tuning-free visual effect transfer across videos.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages