Download README.md from arkanathp/llmae-captioning: direct link, hf CLI and curl.
- Browser
- Download file 1.53 kB
-
https://huggingface.co/arkanathp/llmae-captioning/resolve/main/README.md
- Command line
-
hf download hf://arkanathp/llmae-captioning/README.md
-
curl -L -o README.md https://huggingface.co/arkanathp/llmae-captioning/resolve/main/README.md
license: mit
tags:
- diffusion
- image-captioning
- llmae
llmae-captioning
Image-conditioned latent-diffusion captioning models from the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): a 12-layer transformer score network that denoises LLMAE latents conditioned on frozen SigLIP so400m-patch14-384 patch features through cross-attention. The autoencoder and the image encoder stay frozen; only the score network is trained (VLV-6M captions, 226K steps, batch 64 x 8 GPUs, LR 2e-4).
| File | Autoencoder | Latent | Score network |
|---|---|---|---|
checkpoint_step_226000.pt |
arkanathp/llmae-gemma-270m-s2 |
256 x 640 | hidden 640, 12 layers, 10 heads (111M) |
qwen_s2/checkpoint_step_226000.pt |
arkanathp/llmae-qwen-0.5b-s2 |
256 x 896 | hidden 896, 12 layers, 14 heads (183M) |
The paper decodes with the raw (non-EMA) score-network weights in model_state_dict, CFG 3.0, 250 sampling steps and at most 256 decoded tokens:
python -m diffusion.eval_captioning --images_dir <images> \
--candidate "diffusion:checkpoint_step_226000.pt,autoencoder=arkanathp/llmae-gemma-270m-s2,config_module=diffusion.config_captioning,encoder=siglip,feature_mode=post_ln,cfg=3.0,steps=250,max_tokens=256|name=llmae-gemma"
Training configs: diffusion/config_captioning.py (Gemma) and diffusion/config_captioning_qwen.py (Qwen) in the code release; the COSMOS control row is arkanathp/diffusion-vlv-cosmos-checkpoints.