InstantFusion✳

A shared latent interface for
heterogeneous image generators.

Xuanlang Dai1,2,*Jiazi Bu2,3,*Yujie Zhou3Yuhong Liu2,4 Beichen Zhang2,4Penghui Yang2,3Ziyu Liu2,3Yuhang Zang2,†
1 Fudan University2 Shanghai AI Laboratory3 Shanghai Jiao Tong University4 The Chinese University of Hong Kong

* Equal contribution.† Corresponding author.

SCROLL TO CONNECT
INSTANTFUSION

Shared latent
interface.

A learned connection between independently trained image generators.

SD3 + QWEN-IMAGE CROSS-MODEL COMPOSITION
Four SD3 examples of snowy mountains and flowers beneath a blue sky.
SD3Base model
ENCODELAESHARED SPACEDECODE
Four Qwen-Image examples of snowy mountains and flower-covered valleys.
Qwen-ImageBase model
COMPOSED GENERATION ↘
Four InstantFusion examples combining SD3 and Qwen-Image, depicting snowy mountains and flowers.
InstantFusionSD3 + Qwen-Image
GENERATION EXAMPLES / INDIVIDUAL MODELS AND THEIR COMPOSITION

Each frozen generator connects to a common representation through a model-specific Latent AutoEncoder (LAE). Paired states from the same image and noise level are aligned while preserving the information needed for reconstruction.

Alignment also constrains local denoising dynamics, allowing an intermediate state to be translated into another model’s latent space and generation to continue.

01.A / LATENT ALIGNMENT

Lalign State alignment

Aligns encoded states from the same image at matched noise levels, establishing correspondence in the shared space.

01.B / RECONSTRUCTION

Lrec Latent reconstruction

Reconstructs each model’s input latent from its shared representation, preserving the information needed for subsequent denoising.

01.C / DYNAMICS ALIGNMENT

Ldyn Denoising dynamics

Aligns local denoising updates projected into the shared space, encouraging compatible evolution across models.

Training and inference pipelines of InstantFusion.

Preserve states.
Continue denoising.

We first evaluate what the shared interface preserves, then examine generation when denoising moves between models.

02.A / RECONSTRUCTION QUALITY

Repeated latent round trips

Reconstruction fidelity is measured after 1, 5, and 10 round trips. A single round trip retains high fidelity; repeated translations accumulate reconstruction error.

SD3–Qwen
TripsPSNR ↑SSIM ↑LPIPS ↓
131.3750.94170.04028
522.9770.84310.12630
1020.0430.76700.18892
FLUX–Qwen
TripsPSNR ↑SSIM ↑LPIPS ↓
132.5790.94960.03663
522.6630.84500.13064
1019.1060.76210.21014
SD3–FLUX–Qwen
TripsPSNR ↑SSIM ↑LPIPS ↓
132.3970.95110.03257
523.4450.85800.11006
1020.1190.78080.17591
Visual reconstruction after 1, 5, and 10 latent round trips.
02.B / CROSS-MODEL DENOISING

Generation across model boundaries

Translate an intermediate state through the shared space and continue denoising with another generator. The examples compare model combinations; the table reports their generation quality.

Cross-model generation with SD3, FLUX.1, and Qwen-Image.

Faster
generation.

Combine efficient generators with Qwen-Image to reduce inference time. Adaptive switching uses the relative change in predicted-clean latents to decide when to hand off.

Multiple preferences.
One generation.

Combine existing preference experts during denoising without training a new joint model for every preference combination.

Qualitative examples of multi-preference composition using existing experts.
TRAINING CHECKPOINTS / NOT SAMPLING STEPS
SD3 BASELINEBEFORE DISTILLATION
SD3 baseline: a fox carrying a green backpack in a snowy forest.
WITH CROSS-MODEL DISTILLATIONSTEP 2,400
Cross-model on-policy distillation at training step 2400: a fox carrying a green backpack in a snowy forest.
3006009001,2001,5001,8002,1002,400

SUBJECT / A fox with a green backpack in a snowy pine forest.

Inside the
interface.

A / OBSERVATION

Cross-model compatibility

Related spatial patterns and preserved coarse content suggest partial compatibility between independently trained generators.

B / DYNAMICS ALIGNMENT

The role of dynamics alignment

Adding the dynamics objective preserves state alignment and improves velocity alignment between models.

C / TRAINING DYNAMICS

LAE training dynamics

Reconstruction, state alignment, and dynamics alignment are optimized together while the generators stay frozen.