IDiom-300M
IDiom is an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million intrinsically disordered protein regions (IDRs) curated from the AlphaFold Database. IDiom generates IDRs with or without flanking protein context using a fill-in-the-middle (FIM) format.
Preprint | IDiom GitHub | IDiom Models | IDiom-DB dataset
Model family
| Model | Parameters | Layers | Width | Heads |
|---|---|---|---|---|
| IDiom-300M | 302M | 24 | 1,024 | 16 |
| IDiom-85M | 85M | 12 | 768 | 12 |
| IDiom-20M | 18.9M | 6 | 512 | 8 |
All three models use RMSNorm, SwiGLU, QK normalization, RoPE, and tied input/output embeddings. They share a context length of 1,024 input tokens, including special tokens, and a vocabulary of 20 amino acids, three FIM markers, and four control tokens.
Training data
The models were trained on IDiom-DB-v1, a dataset of 54M IDRs curated from the AlphaFold Database.
Usage
Please see the IDiom Github repository for usage information.
Sparse autoencoder
Only IDiom-300M is accompanied by IDiomSAE, a Top-K sparse autoencoder trained on the residual stream of layer-18.
License
Model weights and code: MIT. The IDiom-DB corpus is CC BY 4.0.
- Downloads last month
- 327