BootsofLagrangian/danbooru-multitier-captions-202606
Viewer • Updated • 11.5M • 710 • 4
Natural language compression into booru tokens
The refiner model takes a text within 512 tokens, and slices it up, each sentence fragment has a danbooru token in it.
The model reuses the architecture of the blocks and the embeddings of the linked model.
It identifies the variable number of split points within the source sentence, such as:
[1, 11, 21, 29, 43, 47, 73, 78, 85, 96, 104, 108, 122, 133, 139, 143, 161, 165, 175]
The following text was divided to these fragments:
The leading whitespace is due to the tokenizer.
Ideally, this could be developed further so that the diffusion model only takes these compressed tokens and not the entire page of text.
Base model
RikkaBotan/NexteraBERT-Mezzoforte-220M-en