Safetensors
English
compression

Natural language compression into booru tokens

The refiner model takes a text within 512 tokens, and slices it up, each sentence fragment has a danbooru token in it.

The model reuses the architecture of the blocks and the embeddings of the linked model.

It identifies the variable number of split points within the source sentence, such as:

[1, 11, 21, 29, 43, 47, 73, 78, 85, 96, 104, 108, 122, 133, 139, 143, 161, 165, 175]

The following text was divided to these fragments:

  • 'A young woman with short, choppy white hair',
  • ' and striking teal eyes stands against a wall covered',
  • ' in large, jagged black graffiti',
  • ' letters. She has a bold, playful expression, winking one eye',
  • ' and sticking her tongue out',
  • ' while raising her hand in a peace sign. She wears a black cross-shaped hair clip, multiple ear piercings,',
  • ' a black choker,',
  • ' and layered gold necklaces.',
  • ' Her outfit is a mix of edgy streetwear:',
  • ' an oversized, bright cyan sweatshirt',
  • ' with a black skull',
  • ' and dragon graphic, worn off one shoulder to reveal a red strap.',
  • ' She wears a short, pleated purple plaid skirt',
  • ' with a black garter strap',
  • ' on her thigh,',
  • ' and distressed black thigh-high stockings with large rips. Her chunky black boots',
  • ' have red laces',
  • ' and accents. A small red and black bag'...

The leading whitespace is due to the tokenizer.

Ideally, this could be developed further so that the diffusion model only takes these compressed tokens and not the entire page of text.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
70.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nebulette/nextera-token-compression

Finetuned
(1)
this model

Dataset used to train nebulette/nextera-token-compression