diffusion-vlv-cosmos-checkpoints
Control row of the captioning experiment in Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): the same image-conditioned latent-diffusion captioner as arkanathp/llmae-captioning, trained on the latents of the public COSMOS text autoencoder instead of LLMAE latents. Everything else is held fixed (VLV-6M captions, frozen SigLIP so400m-patch14-384 conditioning through cross-attention, 12-layer score network, 226K steps at batch 64 x 8 GPUs, LR 2e-4), so the autoencoder is the only variable between the rows.
| File | Autoencoder | Latent | Score network |
|---|---|---|---|
checkpoint_step_226000.pt |
COSMOS (Meshchaninov et al.), 512-token checkpoint | 512 x 768 | hidden 768, 12 layers (159M) |
Training config: diffusion/config_captioning_cosmos.py in the code release; the COSMOS autoencoder is loaded through the same load_autoencoder entry point from a local clone of the COSMOS repository (COSMOS_REPO). Decoding follows the paper protocol: raw (non-EMA) weights in model_state_dict, CFG 3.0, 250 sampling steps, at most 256 decoded tokens.
The score-network weights are released under the MIT license of the code; the COSMOS autoencoder they depend on is distributed by its authors under their own terms.