tuning_free.mov
[arXiv] · [Project Page] · [Dataset]
Maxwell Jones1, Rameen Abdal2, Or Patashnik2, Ruslan Salakhutdinov1, Sergey Tulyakov2, Jun-Yan Zhu1, Kuan-Chieh "Jackson" Wang2
1 Carnegie Mellon University · 2 Snap Research
⚠️ Unofficial reimplementation. This repository is an independent reimplementation built at CMU from the arXiv preprint. Because the dataset, exact training scaffolding, and compute used here differ from the original setup, the results you reproduce may not exactly match those in the paper. Hopefully the results are still competitive though!!
TL;DR. RefVFX transfers a temporal visual effect (cakeify, dissolve, acid
trip, stylization, …) from a reference effect video onto any target video
or image with a single model. The reference is
the spec: whatever happens in it is the effect that gets applied to your
input, preserving the original motion. Under the hood it's a single LoRA on
top of Wan2.1-FLF2V-14B-720P
trained on the published refVFX_dataset
(166 717 pairs, 2 898 effect types). The Video diffusion model takes in the reference video effect and input image/video, and outputs the a new video with the effect applied. See the Method
section below for the conditioning units and CFG scheme.
This repo ships the full stack the paper describes: trainer, single-GPU and 4× USP-parallel inference, a Gradio demo, and the three independent dataset-generation pipelines.
git clone https://github.com/maxwelljones14/refVFX.git refVFX && cd refVFX
# 0) Install uv once (https://docs.astral.sh/uv/getting-started/installation/).
# Skip if you already have it (`uv --version`).
curl -LsSf https://astral.sh/uv/install.sh | sh
# 1) One-time config — copy the template and point REFVFX_CACHE_DIR at a fast
# disk with plenty of free space. Every other directory defaults under it.
# `.env` only holds *directories*; per-run inputs (videos, prompts, LoRA path,
# output path) are passed as CLI flags to the inference scripts.
cp .env.example .env
vim .env
# At this stage just set the directory vars: REFVFX_CACHE_DIR (weights root) and
# REFVFX_MODEL_DIR. You'll fill in REFVFX_LORA_PATH / REFVFX_CAUSVID_LORA after
# downloading the weights below.
# 2) Install the trainer/inference env. `--seed` puts pip into the venv so
# step 3's install_flash_attn.py (which shells out to pip) works.
uv venv --seed --python 3.10 .venv && source .venv/bin/activate
uv pip install -r refVFX_trainer/requirements_trainer.txt
uv pip install -e refVFX_trainer # installs the vendored `diffsynth`
# 3) Optional (recommended for speed): install a prebuilt flash-attention wheel
# matching your python/torch/CUDA. Without it, attention falls back to PyTorch
# SDPA. The script detects your env and pulls the newest matching wheel from
# https://github.com/mjun0812/flash-attention-prebuild-wheels/releases.
python refVFX_trainer/install_flash_attn.pyThe block above also copies .env. REFVFX_CACHE_DIR is the single root for all
model weights — every other path (REFVFX_QWEN_CACHE_DIR, REFVFX_LORAS_DIR,
REFVFX_WAN_I2V_DIR, …) defaults under it. Make sure REFVFX_CACHE_DIR and
REFVFX_MODEL_DIR are set before downloading the weights next.
Now that the env is installed, huggingface-cli is available. Pull the two LoRA
checkpoints from the Hugging Face Hub (the Wan2.1-FLF2V-14B-720P base model is
fetched automatically by ensure_model.py on first run):
| Weight | Repo | Required? |
|---|---|---|
| refVFX LoRA (our trained model) | maxwelljones14/refVFX-LoRA |
required |
| CausVid few-step LoRA | Kijai/WanVideo_comfy |
strongly recommended |
Best results come from our refVFX LoRA stacked with the CausVid LoRA (6-step accelerated sampling), together with
ref_effect_guidance— both are on by default in the launchers, so just keep them.
# Download into your weights directory — the same path you set as
# REFVFX_CACHE_DIR in .env. Public refVFX repo, so no login is needed.
export REFVFX_CACHE_DIR=/path/to/weights # match your .env
huggingface-cli download maxwelljones14/refVFX-LoRA step-10000.safetensors \
--local-dir "$REFVFX_CACHE_DIR" # newer hub versions: `hf download ...`
huggingface-cli download Kijai/WanVideo_comfy \
Wan21_CausVid_14B_T2V_lora_rank32.safetensors \
--local-dir "$REFVFX_CACHE_DIR"Point the two LoRA file variables in .env at the files you just downloaded (and
set REFVFX_MODEL_DIR if you haven't already):
# in .env
REFVFX_LORA_PATH=$REFVFX_CACHE_DIR/step-10000.safetensors
REFVFX_CAUSVID_LORA=$REFVFX_CACHE_DIR/Wan21_CausVid_14B_T2V_lora_rank32.safetensors.env is not auto-sourced for direct python invocations — load it once now, and
again any time you change these variables. NOTE: best results come with the
CausVid LoRA applied and ref_effect_guidance on (both defaulted, so keep them):
set -a; source .env; set +aThe two bash launchers run the recommended default cakeify example —
refVFX LoRA + CausVid few-step LoRA, 6 denoising steps, cfg_scale=6.0,
cfg_scale_ref=2.0, at 480×832 / 33 frames. They take no CLI arguments:
to customize a run, call the python entry point directly (a fully-annotated
reference invocation with every flag is in the comment block at the bottom of
each .sh file).
If you have a GPU with >= 45 Gb VRAM, simply run:
bash refVFX_inference/infer_refvfx.sh
For multi-GPU sequence-parallel inference (auto-detects visible GPUs):
bash refVFX_inference/infer_refvfx_usp.sh
For custom inputs (different reference video, prompt, resolution, seed, …),
call the python files directly. After set -a; source .env; set +a:
python refVFX_inference/infer_refvfx.py [args] # single-GPU
torchrun --nproc_per_node=[num_gpus] --master_port=29506 \ # multi-GPU
refVFX_inference/infer_refvfx_usp.py [args]See the comment block at the bottom of refVFX_inference/infer_refvfx.sh /
infer_refvfx_usp.sh for the full list of supported flags with example values.
If you have a GPU with >= 80 Gb VRAM, run the demo single-GPU:
python refVFX_inference/inference_gradio.py --server_port 7860IF you have a single A6000 only, launch with
python refVFX_inference/inference_gradio.py --vram_limit 20 --server_port 7860 --high_qualityFor 4× <80 Gb GPUs (tested on 4×A6000 49 Gb), launch under torchrun —
the script auto-detects WORLD_SIZE > 1 and patches the pipeline with
Ulysses + Ring sequence parallelism (defaults to ulysses=2, ring=2):
PYTHONUNBUFFERED=1 torchrun --nproc_per_node=4 --master_port=29501 \
refVFX_inference/inference_gradio.py --server_port 7860 --vram_limit 20 --high_qualitySame as the inference scripts, video size auto-drops from 480P (480×832) to 544×320 in the multi-GPU path to fit per-rank memory; only rank-0 hosts the web UI.
| Directory (README link) | What it does | Where to look |
|---|---|---|
refVFX_inference/ |
Single-GPU and 4×GPU USP-parallel inference + Gradio demo | infer_refvfx.py, inference_gradio.py |
refVFX_trainer/ |
LoRA fine-tuning of Wan2.1-FLF2V-14B-720P with the RefVFX conditioning units | train_refvfx.py, refvfx_pipeline.py |
dataset_generation/ |
Three pipelines that produce the three subsets of the published dataset | per-subset READMEs below |
└ I2V_pipeline_from_scratch/ |
LoRA-driven image-to-video subset (Wan-I2V-14B + community LoRAs) | generate_lora_i2v_pipeline.py |
└ V2V_pipeline_from_scratch/ |
Pose-hybrid video-to-video subset (VACE + first/last-frame conditioning) | pose_first_last_hybrid.py, batch_v2v.py |
└ scalable_code_based_video_generation/ |
CPU-only deterministic effects (21 spatial + 15 temporal) | run_examples.py, generator.py |
RefVFX learns to transfer temporal visual effects from a reference video onto an arbitrary target video or image, without per-effect optimization.
Existing video-editing systems are either text-driven (which struggles to describe complex temporal effects like "the subject dissolves into pixels and re-forms") or keyframe-based (which requires manual setup and breaks for non-rigid effects). RefVFX uses a reference effect video as the specification: whatever happens in the reference is the effect to be applied.
RefVFX builds on top of Wan2.1-FLF2V-14B-720P (a text + first/last-frame
video diffusion transformer) and adds three conditioning units, all wired
through a LoRA adapter on the DiT (refvfx_pipeline.py):
- Reference Video Encoder — VAE-encodes the reference effect video; its latents are concatenated width-wise onto the target latents inside the transformer. The transformer's attention then naturally cross-attends between the target frames and the reference effect.
- Control Video Embedder — VAE-encodes the input target video and applies
per-frame masks (
1.0= copy as-is,0.5= transform,0.0= drop) to mix the original content into the y-conditioning channel. - Image Embedder (CLIP) — embeds the explicit first and last frames into the CLIP-conditioning slot for FLF2V.
Classifier-free guidance is multi-axis: separate dropout probabilities and inference-time scales for the reference, the control video, and the text prompt, so users can rebalance "fidelity to the reference effect" vs. "fidelity to the input motion" at sample time.
Three complementary subsets, all packaged in
maxwelljones14/refVFX_dataset:
| Subset | Effects | Pairs | How it was built |
|---|---|---|---|
code_based_edits |
2 736 | 136 800 | Procedural — 21 spatial × 15 temporal effect combos applied with masks |
neural_v2v_data |
114 | 22 922 | Pose-hybrid V2V using VACE + Wan2.1-FLF2V on real motion |
I2V_LoRA |
48 | 6 995 | Community Wan-I2V LoRAs run on Qwen-Image-generated portraits |
| Total | 2 898 | 166 717 |
See dataset_generation/README.md for the full
pipeline, including VLM-based filtering for the I2V and V2V subsets.
Single-stage LoRA fine-tune on the combined dataset, weighted
{code_based_edits: 0.33, neural_v2v_data: 0.51, I2V_LoRA: 0.16} by default
(train_refvfx.sh). See refVFX_trainer/README.md
for the full hparam table, multi-GPU recipe, and resume instructions.
@article{jones2026tuning,
title={Tuning-free Visual Effect Transfer across Videos},
author={Jones, Maxwell and Abdal, Rameen and Patashnik, Or and Salakhutdinov, Ruslan and Tulyakov, Sergey and Zhu, Jun-Yan and Wang, Kuan-Chieh Jackson},
journal={arXiv preprint arXiv:2601.07833},
year={2026}
}- DiffSynth-Studio (modelscope/DiffSynth-Studio, Apache 2.0) —
vendored under
refVFX_trainer/DiffSynth-Studio/. The upstream README lives there. We thank the ModelScope team for the Wan video pipeline and the training scaffolding. - VACE (ali-vilab/VACE, Apache 2.0) — vendored under
dataset_generation/V2V_pipeline_from_scratch/VACE/for pose extraction in the V2V data pipeline. - Wan2.1-FLF2V-14B-720P (Wan-AI) — base video diffusion model.
- Qwen3-VL-8B-Instruct — VLM used by the dataset filtering pipelines.
Apache 2.0 (see refVFX_trainer/LICENSE). Vendored
dependencies retain their original licenses.