Add pipeline tag and paper link

#1
by nielsr HF Staff - opened
Files changed (1) hide show
  1. README.md +3 -2
README.md CHANGED
@@ -4,11 +4,12 @@ tags:
4
  - diffusion
5
  - image-captioning
6
  - llmae
 
7
  ---
8
 
9
  # llmae-captioning
10
 
11
- Image-conditioned latent-diffusion captioning models from the paper *Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders* ([code](https://github.com/arkanath/LLMAE)): a 12-layer transformer score network that denoises LLMAE latents conditioned on frozen SigLIP so400m-patch14-384 patch features through cross-attention. The autoencoder and the image encoder stay frozen; only the score network is trained (VLV-6M captions, 226K steps, batch 64 x 8 GPUs, LR 2e-4).
12
 
13
  | File | Autoencoder | Latent | Score network |
14
  |---|---|---|---|
@@ -22,4 +23,4 @@ python -m diffusion.eval_captioning --images_dir <images> \
22
  --candidate "diffusion:checkpoint_step_226000.pt,autoencoder=arkanathp/llmae-gemma-270m-s2,config_module=diffusion.config_captioning,encoder=siglip,feature_mode=post_ln,cfg=3.0,steps=250,max_tokens=256|name=llmae-gemma"
23
  ```
24
 
25
- Training configs: `diffusion/config_captioning.py` (Gemma) and `diffusion/config_captioning_qwen.py` (Qwen) in the code release; the COSMOS control row is `arkanathp/diffusion-vlv-cosmos-checkpoints`.
 
4
  - diffusion
5
  - image-captioning
6
  - llmae
7
+ pipeline_tag: image-to-text
8
  ---
9
 
10
  # llmae-captioning
11
 
12
+ Image-conditioned latent-diffusion captioning models from the paper *Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders* ([paper](https://huggingface.co/papers/2609.27248), [code](https://github.com/arkanath/LLMAE)): a 12-layer transformer score network that denoises LLMAE latents conditioned on frozen SigLIP so400m-patch14-384 patch features through cross-attention. The autoencoder and the image encoder stay frozen; only the score network is trained (VLV-6M captions, 226K steps, batch 64 x 8 GPUs, LR 2e-4).
13
 
14
  | File | Autoencoder | Latent | Score network |
15
  |---|---|---|---|
 
23
  --candidate "diffusion:checkpoint_step_226000.pt,autoencoder=arkanathp/llmae-gemma-270m-s2,config_module=diffusion.config_captioning,encoder=siglip,feature_mode=post_ln,cfg=3.0,steps=250,max_tokens=256|name=llmae-gemma"
24
  ```
25
 
26
+ Training configs: `diffusion/config_captioning.py` (Gemma) and `diffusion/config_captioning_qwen.py` (Qwen) in the code release; the COSMOS control row is `arkanathp/diffusion-vlv-cosmos-checkpoints`.