Equivalent Flows, Unequal Learning: Clean-Latent Prediction in Transformers

Paper Project Page Code

Figure 1: Matched Base-scale qualitative comparison on class-conditional ImageNet 256×256. Generations from matched B/1 models: clean prediction (x̂) preserves global coherence, boundary crispness, and fine details, whereas direct velocity prediction (v̂) frequently distorts, blurs, or introduces semantic artifacts under identical sampling.

Authors

Funing Fu1* · Tenghui Wang2* · Guanyu Zhou3 · Junyong Cen1 · Qichao Zhu4
1 Independent Researcher · 2 Wuhan University of Technology · 3 Technology Innovation Institute · 4 Hangzhou Jiyi AI (*Equal contribution)

Overview

JLT (Just Latent Transformer) addresses a fundamental question in flow-based generative modeling: which quantity should the neural network directly predict?

Given corrupted state zt and time t, predicting clean data x and predicting velocity v = x - ε are algebraically equivalent under the sample-wise linear path zt = t·x + (1 - t)·ε, via the exact affine readout: v^=x^−zt1−t\hat{v} = \frac{\hat{x} - z_t}{1 - t}

While an ideal oracle can express either form, a finite-capacity Transformer behaves very differently. Clean prediction delegates the known linear residual and time gain (1 - t)-1 to the analytic readout, sparing the network from internalizing this known response. Moreover, in anisotropic latent spaces (such as the FLUX.2 VAE), clean prediction naturally attenuates low-variance directions, whereas velocity prediction adds an isotropic unit floor Σv = Σx + I that spreads variance across weakly supported directions.

Across identical architectures, training setups, and frozen FLUX.2 VAE representations:

  • Base (B/1): FID-50K improves from 6.56 to 2.84 (IS 204.83) under Lv, and to 2.38 (IS 256.88) with unweighted clean MSE Lx.
  • Large (L/1): FID-50K improves from 2.12 to 1.83 (IS 301.07).
  • Huge (H/1, 951.3M): Reaches FID-50K 1.19 (IS 271.96).

Results

Matched Prediction Target (Class-Conditional ImageNet 256×256)
Within each scale, the VAE, architecture, velocity loss, schedule, and sampler are strictly matched:

Scale Model Target Loss FID-50K ↓ IS ↑
Base Baseline velocity v Lv 6.56 132.12
Base JLT-B/1 clean latent x Lv 2.84 204.83
Base JLT-B/1 clean latent x Lx 2.38 256.88
Large Baseline velocity v Lv 2.12 236.21
Large JLT-L/1 clean latent x Lv 1.83 301.07
Huge Baseline velocity v Lv 1.60 327.41
Huge JLT-H/1 clean latent x Lv 1.19 271.96

Model Family

Model Depth Width Heads Params Tokenizer
JLT-B/1 12 768 12 130.5M FLUX.2 VAE (frozen)
JLT-L/1 24 1024 16 458.1M FLUX.2 VAE (frozen)
JLT-H/1 32 1280 16 951.3M FLUX.2 VAE (frozen)

Usage

Download Checkpoints

# Default Base checkpoint
huggingface-cli download dawn-neo/JLT checkpoint-last.pth

# Or download specific model checkpoints: JLT-L-1.pth, JLT-H-1.pth

Evaluation

Run evaluation using pre-encoded ImageNet latents with main_jlt.py:

python main_jlt.py \
    --model JLT-B/1 \
    --vae_type flux2 \
    --img_size 256 \
    --data_path /path/to/imagenet_latents_256 \
    --use_latent_cache \
    --online_eval \
    --eval_freq 1 \
    --gen_bsz 128 \
    --num_images 50000 \
    --cfg 2.9 \
    --num_sampling_steps 50 \
    --resume /path/to/checkpoint-last.pth \
    --output_dir ./eval_output

For full training recipes and data preparation, see the GitHub repository.

Citation

@article{fu2026jlt,
  title={Equivalent Flows, Unequal Learning: Clean-Latent Prediction in Transformers},
  author={Fu, Funing and Wang, Tenghui and Zhou, Guanyu and Cen, Junyong and Zhu, Qichao},
  journal={arXiv preprint arXiv:2605.27102},
  year={2026}
}

Acknowledgements

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including dawn-neo/JLT

Papers for dawn-neo/JLT