Safetensors
botanic1
custom_code

Is the released Botanic1-S the 8k or 128k checkpoint?

#1
by liujh0223 - opened

Hi Living Models team,

Thank you for releasing Botanic1-S and the associated code/resources.

I have a question about the relationship between the currently released living-models/Botanic1-S checkpoint and the long-context Botanic1-S model described in the BOTANIC-1 preprint.

From the model card, my understanding is that the currently released Botanic1-S was pretrained with an 8,192-token sequence length. However, the paper also describes a context-extension curriculum for a Botanic1-S-sized model:

8,192 → 16,384 → 32,768 → 65,536 → 131,072 bp

and refers to a final long-context model/checkpoint, which I understand as the Botanic1-S-128k model used in the long-context experiments.

Could you please clarify the following?

  1. Is the currently released living-models/Botanic1-S checkpoint the 8,192-bp pretrained model, or does it already include the context-extension training up to 131,072 bp?

  2. If the currently released checkpoint is the 8k version, is the final 131,072-bp context-extended Botanic1-S checkpoint publicly available somewhere?

  3. If it is not yet public, are there plans to release the 128k checkpoint?

  4. For sequence lengths substantially longer than 8,192 bp, such as ~78 kb, would you recommend using only the context-extended checkpoint rather than directly running the current 8k checkpoint at a longer sequence length?

My use case is nucleotide-resolution functional track prediction from ~77,824-bp plant genomic sequences, so the distinction between an 8k-pretrained checkpoint and a genuinely context-extended 128k checkpoint is very important for us.

Thanks very much for the clarification and for making these resources available.

Best regards,
Jianhong Liu

Living Models org

Hi Jianhong!

Thanks for your interest in BOTANIC-1!

The currently released Botanic1-S checkpoint is the 8,192-bp model. The 128k context-extended checkpoint described in the paper has not been released yet.
We do plan to make the long-context model available, but we haven’t finalized the details/timeline of the public release yet.

For sequences substantially longer than 8k, such as ~78 kb, we would recommend using the context-extended checkpoint rather than simply running the current 8k checkpoint at a longer sequence length.

Thanks again for the detailed question and for your interest in the model!

Thank you very much for the clarification!

I’m really looking forward to the release of the 128k context-extended checkpoint, as it should be highly relevant to our long-context nucleotide-resolution track prediction work!

Sign up or log in to comment