Equivalent Flows, Unequal Learning: Clean-Latent Prediction in Transformers
Figure 1: Matched Base-scale qualitative comparison on class-conditional ImageNet 256×256. Generations from matched B/1 models: clean prediction (x̂) preserves global coherence, boundary crispness, and fine details, whereas direct velocity prediction (v̂) frequently distorts, blurs, or introduces semantic artifacts under identical sampling.
Authors
Funing Fu1* · Tenghui Wang2* · Guanyu Zhou3 · Junyong Cen1 · Qichao Zhu4
1 Independent Researcher · 2 Wuhan University of Technology · 3 Technology Innovation Institute · 4 Hangzhou Jiyi AI (*Equal contribution)
Overview
JLT (Just Latent Transformer) addresses a fundamental question in flow-based generative modeling: which quantity should the neural network directly predict?
Given corrupted state zt and time t, predicting clean data x and predicting velocity v = x - ε are algebraically equivalent under the sample-wise linear path zt = t·x + (1 - t)·ε, via the exact affine readout:
While an ideal oracle can express either form, a finite-capacity Transformer behaves very differently. Clean prediction delegates the known linear residual and time gain (1 - t)-1 to the analytic readout, sparing the network from internalizing this known response. Moreover, in anisotropic latent spaces (such as the FLUX.2 VAE), clean prediction naturally attenuates low-variance directions, whereas velocity prediction adds an isotropic unit floor Σv = Σx + I that spreads variance across weakly supported directions.
Across identical architectures, training setups, and frozen FLUX.2 VAE representations:
- Base (B/1): FID-50K improves from 6.56 to 2.84 (IS 204.83) under Lv, and to 2.38 (IS 256.88) with unweighted clean MSE Lx.
- Large (L/1): FID-50K improves from 2.12 to 1.83 (IS 301.07).
- Huge (H/1, 951.3M): Reaches FID-50K 1.19 (IS 271.96).
Results
Matched Prediction Target (Class-Conditional ImageNet 256×256)
Within each scale, the VAE, architecture, velocity loss, schedule, and sampler are strictly matched:
| Scale | Model | Target | Loss | FID-50K ↓ | IS ↑ |
|---|---|---|---|---|---|
| Base | Baseline | velocity v | Lv | 6.56 | 132.12 |
| Base | JLT-B/1 | clean latent x | Lv | 2.84 | 204.83 |
| Base | JLT-B/1 | clean latent x | Lx | 2.38 | 256.88 |
| Large | Baseline | velocity v | Lv | 2.12 | 236.21 |
| Large | JLT-L/1 | clean latent x | Lv | 1.83 | 301.07 |
| Huge | Baseline | velocity v | Lv | 1.60 | 327.41 |
| Huge | JLT-H/1 | clean latent x | Lv | 1.19 | 271.96 |
Model Family
| Model | Depth | Width | Heads | Params | Tokenizer |
|---|---|---|---|---|---|
| JLT-B/1 | 12 | 768 | 12 | 130.5M | FLUX.2 VAE (frozen) |
| JLT-L/1 | 24 | 1024 | 16 | 458.1M | FLUX.2 VAE (frozen) |
| JLT-H/1 | 32 | 1280 | 16 | 951.3M | FLUX.2 VAE (frozen) |
Usage
Download Checkpoints
# Default Base checkpoint
huggingface-cli download dawn-neo/JLT checkpoint-last.pth
# Or download specific model checkpoints: JLT-L-1.pth, JLT-H-1.pth
Evaluation
Run evaluation using pre-encoded ImageNet latents with main_jlt.py:
python main_jlt.py \
--model JLT-B/1 \
--vae_type flux2 \
--img_size 256 \
--data_path /path/to/imagenet_latents_256 \
--use_latent_cache \
--online_eval \
--eval_freq 1 \
--gen_bsz 128 \
--num_images 50000 \
--cfg 2.9 \
--num_sampling_steps 50 \
--resume /path/to/checkpoint-last.pth \
--output_dir ./eval_output
For full training recipes and data preparation, see the GitHub repository.
Citation
@article{fu2026jlt,
title={Equivalent Flows, Unequal Learning: Clean-Latent Prediction in Transformers},
author={Fu, Funing and Wang, Tenghui and Zhou, Guanyu and Cen, Junyong and Zhu, Qichao},
journal={arXiv preprint arXiv:2605.27102},
year={2026}
}
Acknowledgements
- JiT — Base architecture
- FLUX.2 VAE — Latent space
- Li & He. "Back to Basics" — Clean prediction insight