--- license: mit tags: - diffusion - image-captioning - llmae --- # llmae-captioning Image-conditioned latent-diffusion captioning models from the paper *Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders* ([code](https://github.com/arkanath/LLMAE)): a 12-layer transformer score network that denoises LLMAE latents conditioned on frozen SigLIP so400m-patch14-384 patch features through cross-attention. The autoencoder and the image encoder stay frozen; only the score network is trained (VLV-6M captions, 226K steps, batch 64 x 8 GPUs, LR 2e-4). | File | Autoencoder | Latent | Score network | |---|---|---|---| | `checkpoint_step_226000.pt` | `arkanathp/llmae-gemma-270m-s2` | 256 x 640 | hidden 640, 12 layers, 10 heads (111M) | | `qwen_s2/checkpoint_step_226000.pt` | `arkanathp/llmae-qwen-0.5b-s2` | 256 x 896 | hidden 896, 12 layers, 14 heads (183M) | The paper decodes with the raw (non-EMA) score-network weights in `model_state_dict`, CFG 3.0, 250 sampling steps and at most 256 decoded tokens: ```bash python -m diffusion.eval_captioning --images_dir \ --candidate "diffusion:checkpoint_step_226000.pt,autoencoder=arkanathp/llmae-gemma-270m-s2,config_module=diffusion.config_captioning,encoder=siglip,feature_mode=post_ln,cfg=3.0,steps=250,max_tokens=256|name=llmae-gemma" ``` Training configs: `diffusion/config_captioning.py` (Gemma) and `diffusion/config_captioning_qwen.py` (Qwen) in the code release; the COSMOS control row is `arkanathp/diffusion-vlv-cosmos-checkpoints`.