mRNA-GPT-eukaryote
Pretrained on eukaryotic coding sequences.
mRNA-GPT is a decoder-only transformer language model over codons. The tokenizer operates on codons rather than nucleotides, so one token is exactly one residue position. That makes protein-constrained decoding exact: at step t the logits are masked to the synonymous codon set of residue t, so the generated coding sequence translates back to the requested protein by construction, not with high probability.
| Architecture | decoder-only transformer, 24 layers, 1024-dim, 16 heads |
| Parameters | 302,109,696 |
| Vocabulary | 68 tokens (4 special + 64 codons) |
| Context | 2048 codons |
| Positional encoding | rope |
| Precision | bfloat16 |
| Code | https://github.com/ZHymLumine/mRNA-GPT |
Installation
git clone https://github.com/ZHymLumine/mRNA-GPT && cd mRNA-GPT
pip install -r requirements.txt
pip install huggingface_hub safetensors
huggingface-cli download ZYMScott/mRNA-GPT-eukaryote --local-dir mRNA-GPT-eukaryote
Design a coding sequence for your protein
This is the main use. Put your target protein in a FASTA file:
>my_target
MKAIFVLKGSLDRDLEHHHHHHGSMSTAVLENPGLGRKLSDFGQETSYIEDNSNQ
python -m mrnagpt.generate \
--ckpt mRNA-GPT-eukaryote \
--proteins my_target.fasta \
--temperature 0.8 --top-p 0.95 \
--out designs.fasta --report designs.md
The output FASTA keeps your sequence names, and the report states what fraction of designs start with ATG, end with a single in-frame stop, contain no internal stop, and translate to exactly the requested protein:
| | starts with ATG | ends with a stop codon | has an internal stop codon | all three satisfied |
|---|---:|---:|---:|---:|
| constrained | 100.00% | 100.00% | 0.00% | 100.00% |
- target protein exact match: 100.00% (1/1)
Sampling several designs per target and ranking them is the usual workflow: repeat the call, or pass a FASTA with the target repeated.
From Python
from mrnagpt.generate import load_model, constrained_sample
from mrnagpt.vocab import translate
model = load_model("mRNA-GPT-eukaryote", device="cuda") # or "cpu"
target = "MKAIFVLKGSLDRDLEHHHHHHGSMSTAVLENPGLGRKLSDFGQETSYIEDNSNQ"
designs = constrained_sample(
model, [target] * 8, # eight independent designs
temperature=0.8, top_p=0.95, device="cuda",
)
for codons in designs:
cds = "".join(codons)
assert translate(codons, stop_at_first_stop=True) == target
print(cds)
load_model accepts this directory, the model.safetensors inside it, or a
.pt checkpoint written during training.
Unconstrained generation
Sampling coding sequences without a target protein, for characterising what the model has learned about the domain:
python -m mrnagpt.generate --ckpt mRNA-GPT-eukaryote --n 100 --out sampled.fasta
Files
| file | contents |
|---|---|
model.safetensors |
weights, bfloat16 |
config.json |
architecture, vocabulary size, token ids |
vocab.txt |
the 68-token codon vocabulary, in id order |
provenance.json |
source checkpoint, training step, SHA-256 of the weights |
The tied lm_head.weight is not stored separately, because safetensors will not
serialise two names backed by the same storage. config.json records
tie_weights, and the key is restored when the model is built.
Related models
ZYMScott/mRNA-GPT-bacteriaโ pretrained on bacterial coding sequences.ZYMScott/mRNA-GPT-archaeaโ pretrained on archaeal coding sequences.ZYMScott/mRNA-GPT-bacterial-expressionโ fine-tuned from mRNA-GPT-bacteria on the high-expression arm of a bacterial protein-expression library.ZYMScott/mRNA-GPT-fungal-expressionโ fine-tuned on the high-expression arm of a fungal expression dataset.ZYMScott/mRNA-GPT-stabilityโ fine-tuned on the high-stability arm of an mRNA stability dataset.
Limitations
- The model generates coding sequences only โ no UTRs, no poly(A), no cap-proximal structure. Expression depends on those too.
- Sampling is per-sequence and has no notion of a host's codon supply beyond what it learned from the pretraining domain. For a host far from that domain, fine-tune rather than relying on the pretrained model.
- A property fine-tune shifts the codon distribution toward the high-property arm of its training set. That is a statistical shift, not a guarantee about any individual design: validate experimentally.
- Protein-constrained decoding guarantees the translated protein, not that the resulting mRNA folds or expresses well.
License
Apache-2.0.
- Downloads last month
- -