Add pipeline tag and paper link
#1
by nielsr HF Staff - opened
README.md
CHANGED
|
@@ -4,11 +4,12 @@ tags:
|
|
| 4 |
- diffusion
|
| 5 |
- image-captioning
|
| 6 |
- llmae
|
|
|
|
| 7 |
---
|
| 8 |
|
| 9 |
# llmae-captioning
|
| 10 |
|
| 11 |
-
Image-conditioned latent-diffusion captioning models from the paper *Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders* ([code](https://github.com/arkanath/LLMAE)): a 12-layer transformer score network that denoises LLMAE latents conditioned on frozen SigLIP so400m-patch14-384 patch features through cross-attention. The autoencoder and the image encoder stay frozen; only the score network is trained (VLV-6M captions, 226K steps, batch 64 x 8 GPUs, LR 2e-4).
|
| 12 |
|
| 13 |
| File | Autoencoder | Latent | Score network |
|
| 14 |
|---|---|---|---|
|
|
@@ -22,4 +23,4 @@ python -m diffusion.eval_captioning --images_dir <images> \
|
|
| 22 |
--candidate "diffusion:checkpoint_step_226000.pt,autoencoder=arkanathp/llmae-gemma-270m-s2,config_module=diffusion.config_captioning,encoder=siglip,feature_mode=post_ln,cfg=3.0,steps=250,max_tokens=256|name=llmae-gemma"
|
| 23 |
```
|
| 24 |
|
| 25 |
-
Training configs: `diffusion/config_captioning.py` (Gemma) and `diffusion/config_captioning_qwen.py` (Qwen) in the code release; the COSMOS control row is `arkanathp/diffusion-vlv-cosmos-checkpoints`.
|
|
|
|
| 4 |
- diffusion
|
| 5 |
- image-captioning
|
| 6 |
- llmae
|
| 7 |
+
pipeline_tag: image-to-text
|
| 8 |
---
|
| 9 |
|
| 10 |
# llmae-captioning
|
| 11 |
|
| 12 |
+
Image-conditioned latent-diffusion captioning models from the paper *Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders* ([paper](https://huggingface.co/papers/2609.27248), [code](https://github.com/arkanath/LLMAE)): a 12-layer transformer score network that denoises LLMAE latents conditioned on frozen SigLIP so400m-patch14-384 patch features through cross-attention. The autoencoder and the image encoder stay frozen; only the score network is trained (VLV-6M captions, 226K steps, batch 64 x 8 GPUs, LR 2e-4).
|
| 13 |
|
| 14 |
| File | Autoencoder | Latent | Score network |
|
| 15 |
|---|---|---|---|
|
|
|
|
| 23 |
--candidate "diffusion:checkpoint_step_226000.pt,autoencoder=arkanathp/llmae-gemma-270m-s2,config_module=diffusion.config_captioning,encoder=siglip,feature_mode=post_ln,cfg=3.0,steps=250,max_tokens=256|name=llmae-gemma"
|
| 24 |
```
|
| 25 |
|
| 26 |
+
Training configs: `diffusion/config_captioning.py` (Gemma) and `diffusion/config_captioning_qwen.py` (Qwen) in the code release; the COSMOS control row is `arkanathp/diffusion-vlv-cosmos-checkpoints`.
|