Mixture Gradient Flow

Unifying Distributional Training
for One-Step Visual Generation

A common gradient-flow view of distributional training.
Gaussian mixtures capture multimodal feature distributions with a compact, adjustable representation.

Anonymous project page

One forward pass. FLUX.2 [klein] 4B post-trained with MGFlow · Selected samples

1 NFE

One-step sampling

1.45 ↓

FDr6 · pMF-H · ImageNet

0.900 ↑

GenEval · FLUX.2 [klein] 4B

21.98 ↑

PickScore · FLUX.2 [klein] 4B

01 / Sample explorer

From a single step.

Browse selected generated images.
Choose an image to view it at full resolution.

FLUX.2 [klein] 4B + MGFlow · 512 × 512 · 1 NFE

Loading samples…

This is a gallery of pre-generated, selected qualitative examples from the paper, not a live inference service or an unbiased sample set.

02 / The idea

One framework.
Different distribution models.

Distributional training matches generated and reference features. Its update depends on how we represent the distributions and how we compare them.

𝒮

Batch sampling

ℰ

Frozen encoders

ℳ

Distribution model

𝒟

Matching objective

A unified distributional training framework: FD-Loss models a single Gaussian, MGFlow uses Gaussian mixtures, and Drifting uses kernel density estimates. Distribution matching supplies a Wasserstein gradient-flow update.
MGFlow uses Gaussian mixtures to capture multimodal structure between a single-Gaussian approximation and sample-based representations.
Distribution models and matching objectives
ObjectiveGaussianGaussian mixtureSample-based
OT-basedFD-LossYang et al., 2026 ↗MGFlow-W2W-FlowHan et al., 2026 ↗
Score-based (KL)Gaussian KLMGFlow-KLGaussian-kernel DriftingDeng et al., 2026 ↗

W-Flow uses empirical measures and Sinkhorn divergence; Gaussian-kernel Drifting uses kernel-smoothed densities. MGFlow-W2 uses paired Gaussian transport costs, and MGFlow-KL uses paired component score fields.

Match the components,
not just their local scores.

A more expressive density model alone does not ensure stable matching. Soft assignments can lose components, while global mixture scores can miss incorrect mass allocation between separated modes.

MGFlow constrains each component's batch mass with a linear program, updates its statistics, and matches corresponding reference and generated components. This pairing is used for both Gaussian transport and KL-based updates.

Toy experiment comparing posterior-assigned global scores, LP-assigned global scores, and LP-paired updates over time. The LP-paired update aligns component positions and reference mass allocation.
Component assignment and the choice of score field both matter. The toy experiment compares soft-assigned global, LP-assigned global, and LP-paired updates.

03 / Results

Image quality. Prompt alignment.
One-step generation.

Selected comparisons from the paper.
↓ Lower is better. ↑ Higher is better.

ImageNet 256 × 256

JiT-H & pMF-H

100 epochs · SigLIP, Inception & MAE encoders

MethodNFEFDr6 ↓FDr3 ↓FID ↓
JiT-HLi & He, 2026 ↗
Pretrained50×2×2†7.669.071.97
FD-LossYang et al., 2026 ↗12.654.440.75
AdvFDGao et al., 2026 ↗11.802.930.72
AMFD-CLiu et al., 2026 ↗12.153.100.85
AMFD-ULiu et al., 2026 ↗11.792.780.83
MGFlow-W212.554.420.99
MGFlow-KL11.642.570.94
pMF-HLu et al., 2026 ↗
Pretrained16.876.092.29
FD-LossYang et al., 2026 ↗11.892.690.77
AdvFDGao et al., 2026 ↗11.742.500.74
AMFD-CLiu et al., 2026 ↗11.932.430.86
AMFD-ULiu et al., 2026 ↗11.752.300.85
MGFlow-W211.502.051.09
MGFlow-KL11.451.881.07

FDr6 averages normalized Fréchet distances across six encoders; FDr3 uses three held-out encoders. † Full-CFG NFE upper bound for interval CFG. Bold marks the best value within each backbone among the methods shown.

Text-to-image 512 × 512

FLUX.2 [klein] 4B

MGFlow: 1,000 training steps · K = 1 + 4

MethodNFEGenEval ↑PickScore ↑
FLUX.2 [klein]Black Forest Labs, 2026 ↗40.79421.85
One-step initialization10.53520.10
DMD2Yin et al., 2024 ↗10.804—
FD-SIMYang et al., 2026 ↗10.80221.62
iRDMFeng et al., 2026 ↗10.82621.82
AMFD-U-SIMLiu et al., 2026 ↗10.81321.77
AMFD-C-SIMLiu et al., 2026 ↗10.84621.85
AMFD-C-10 enc.Liu et al., 2026 ↗10.84521.82
MGFlow · image-only10.85621.86
MGFlow · joint10.90021.98

The one-step initialization is evaluated before MGFlow post-training. Joint matching uses image and SigLIP2 text features. PickScore uses 499 Pick-a-Pic prompts. Published baseline values follow AMFD, Table 3; — denotes an unreported value.

04 / Models & data

Build on MGFlow.

Six ImageNet backbones, two objectives.
Two text-to-image variants, training data, and references.

ImageNet

Post-trained JiT and pMF models at B, L, and H scales. MGFlow-W2 uses K = 1 + 4; MGFlow-KL uses K = 1 + 4 + 16.

Model checkpoints

Selected post-trained checkpointpMF-H · MGFlow-KL weights · 1 step Selected base checkpointpMF-H · pretrained weights

Text-to-image

FLUX.2 [klein] 4B post-trained with MGFlow-KL. Image-only and joint image–text distribution matching use the same reference construction.

Model checkpoints contain trained weights and the training step. See the model repository for the current release and loading information.