IDiom-300M

IDiom is an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million intrinsically disordered protein regions (IDRs) curated from the AlphaFold Database. IDiom generates IDRs with or without flanking protein context using a fill-in-the-middle (FIM) format.

Preprint   |   IDiom GitHub   |   IDiom Models   |   IDiom-DB dataset

Model family

Model Parameters Layers Width Heads
IDiom-300M 302M 24 1,024 16
IDiom-85M 85M 12 768 12
IDiom-20M 18.9M 6 512 8

All three models use RMSNorm, SwiGLU, QK normalization, RoPE, and tied input/output embeddings. They share a context length of 1,024 input tokens, including special tokens, and a vocabulary of 20 amino acids, three FIM markers, and four control tokens.

Training data

The models were trained on IDiom-DB-v1, a dataset of 54M IDRs curated from the AlphaFold Database.

Usage

Please see the IDiom Github repository for usage information.

Sparse autoencoder

Only IDiom-300M is accompanied by IDiomSAE, a Top-K sparse autoencoder trained on the residual stream of layer-18.

License

Model weights and code: MIT. The IDiom-DB corpus is CC BY 4.0.

Downloads last month
327
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train jxliu2/idiom-300M

Space using jxliu2/idiom-300M 1

Collection including jxliu2/idiom-300M

Paper for jxliu2/idiom-300M