diffusion-vlv-cosmos-checkpoints

Control row of the captioning experiment in Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): the same image-conditioned latent-diffusion captioner as arkanathp/llmae-captioning, trained on the latents of the public COSMOS text autoencoder instead of LLMAE latents. Everything else is held fixed (VLV-6M captions, frozen SigLIP so400m-patch14-384 conditioning through cross-attention, 12-layer score network, 226K steps at batch 64 x 8 GPUs, LR 2e-4), so the autoencoder is the only variable between the rows.

File Autoencoder Latent Score network
checkpoint_step_226000.pt COSMOS (Meshchaninov et al.), 512-token checkpoint 512 x 768 hidden 768, 12 layers (159M)

Training config: diffusion/config_captioning_cosmos.py in the code release; the COSMOS autoencoder is loaded through the same load_autoencoder entry point from a local clone of the COSMOS repository (COSMOS_REPO). Decoding follows the paper protocol: raw (non-EMA) weights in model_state_dict, CFG 3.0, 250 sampling steps, at most 256 decoded tokens.

The score-network weights are released under the MIT license of the code; the COSMOS autoencoder they depend on is distributed by its authors under their own terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support