llmae-captioning / README.md
arkanathp's picture
Add model card
bcf9cd8 verified
|
Raw History Blame Contribute Delete
1.53 kB
metadata
license: mit
tags:
  - diffusion
  - image-captioning
  - llmae

llmae-captioning

Image-conditioned latent-diffusion captioning models from the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): a 12-layer transformer score network that denoises LLMAE latents conditioned on frozen SigLIP so400m-patch14-384 patch features through cross-attention. The autoencoder and the image encoder stay frozen; only the score network is trained (VLV-6M captions, 226K steps, batch 64 x 8 GPUs, LR 2e-4).

File Autoencoder Latent Score network
checkpoint_step_226000.pt arkanathp/llmae-gemma-270m-s2 256 x 640 hidden 640, 12 layers, 10 heads (111M)
qwen_s2/checkpoint_step_226000.pt arkanathp/llmae-qwen-0.5b-s2 256 x 896 hidden 896, 12 layers, 14 heads (183M)

The paper decodes with the raw (non-EMA) score-network weights in model_state_dict, CFG 3.0, 250 sampling steps and at most 256 decoded tokens:

python -m diffusion.eval_captioning --images_dir <images> \
  --candidate "diffusion:checkpoint_step_226000.pt,autoencoder=arkanathp/llmae-gemma-270m-s2,config_module=diffusion.config_captioning,encoder=siglip,feature_mode=post_ln,cfg=3.0,steps=250,max_tokens=256|name=llmae-gemma"

Training configs: diffusion/config_captioning.py (Gemma) and diffusion/config_captioning_qwen.py (Qwen) in the code release; the COSMOS control row is arkanathp/diffusion-vlv-cosmos-checkpoints.