Follow-up to #14824, which was closed as "noise replay" (same seed/stream/shape for generation and edit). I see a degradation at output_resolution=1024 that does not depend on the seed, so replay does not explain it.
Setup: diffusers 0121a91 and current main 0377f0c (same result), torch 2.14.0, transformers 5.17.0, bf16, MPS (M5 Max, macOS 26.6), 40 steps. Script: one edit per process, below.
| input |
seed |
result |
RGB std (input 0.259) |
PSNR vs input |
| image generated by Qwen-Image-2.1 (via mflux, MLX RNG, seed 42), red circle drawn in |
42 |
instruction followed (bicycle appears) but whole image grey and embossed |
0.090 |
12.1 dB |
| same |
12345 |
same degradation |
0.099 |
11.9 dB |
| real photo (Kodak kodim01, 768×512) |
42 |
clean, correct edit |
0.176 (input 0.162) |
17.9 dB |
| image generated by the diffusers pipeline (seed 700467539, 1376×768) |
12345 |
edit applied, image darker, edge energy ×2.0 |
0.129 (input 0.167) |
16.5 dB |
The generation seed/stream differs from the edit's in rows 1, 2 and 4 (row 1 is even a different framework's RNG), so the edit's initial noise is not the draw that produced the image. At 512 and 768 all of these edits are clean. This may be what @peterc described in #14824 as a separate issue.
Script (one edit per process)
"""One edit at output_resolution=1024, run in its own process."""
import sys, time, json, numpy as np, torch
from PIL import Image
from diffusers import QwenImage21Pipeline
src, prompt, seed, out = sys.argv[1], sys.argv[2], int(sys.argv[3]), sys.argv[4]
img = Image.open(src).convert("RGB")
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", dtype=torch.bfloat16).to("mps")
pipe.set_progress_bar_config(disable=True)
t = time.time()
arr = pipe(image=img, prompt=prompt, num_inference_steps=40, generator=torch.Generator("mps").manual_seed(seed),
output_resolution=1024, output_type="np").images[0]
rgb = (np.clip(arr[..., :3], 0, 1) * 255).round().astype("uint8")
res = Image.fromarray(rgb); res.save(out)
# RGB std as a contrast measure (washed-out ~0.09, clean ~0.26 on this image) and edge energy relative to the input
def grad(a): a = a.mean(-1); return np.abs(np.diff(a, axis=0)).mean() + np.abs(np.diff(a, axis=1)).mean()
inp = np.asarray(img.resize(res.size)).astype(float) / 255; o = rgb.astype(float) / 255
print(json.dumps({"out": out, "s": round(time.time() - t), "std": round(float(o.std()), 3), "std_in": round(float(inp.std()), 3),
"edge_ratio": round(float(grad(o) / grad(inp)), 2), "psnr": round(float(10 * np.log10(1 / ((o - inp) ** 2).mean())), 1)}))
Usage: python edit1024.py <image> <prompt> <seed> <out.png>
Follow-up to #14824, which was closed as "noise replay" (same seed/stream/shape for generation and edit). I see a degradation at
output_resolution=1024that does not depend on the seed, so replay does not explain it.Setup: diffusers 0121a91 and current main 0377f0c (same result), torch 2.14.0, transformers 5.17.0, bf16, MPS (M5 Max, macOS 26.6), 40 steps. Script: one edit per process, below.
The generation seed/stream differs from the edit's in rows 1, 2 and 4 (row 1 is even a different framework's RNG), so the edit's initial noise is not the draw that produced the image. At 512 and 768 all of these edits are clean. This may be what @peterc described in #14824 as a separate issue.
Script (one edit per process)
Usage:
python edit1024.py <image> <prompt> <seed> <out.png>