mRNA-GPT-eukaryote

Pretrained on eukaryotic coding sequences.

mRNA-GPT is a decoder-only transformer language model over codons. The tokenizer operates on codons rather than nucleotides, so one token is exactly one residue position. That makes protein-constrained decoding exact: at step t the logits are masked to the synonymous codon set of residue t, so the generated coding sequence translates back to the requested protein by construction, not with high probability.

Architecture decoder-only transformer, 24 layers, 1024-dim, 16 heads
Parameters 302,109,696
Vocabulary 68 tokens (4 special + 64 codons)
Context 2048 codons
Positional encoding rope
Precision bfloat16
Code https://github.com/ZHymLumine/mRNA-GPT

Installation

git clone https://github.com/ZHymLumine/mRNA-GPT && cd mRNA-GPT
pip install -r requirements.txt
pip install huggingface_hub safetensors
huggingface-cli download ZYMScott/mRNA-GPT-eukaryote --local-dir mRNA-GPT-eukaryote

Design a coding sequence for your protein

This is the main use. Put your target protein in a FASTA file:

>my_target
MKAIFVLKGSLDRDLEHHHHHHGSMSTAVLENPGLGRKLSDFGQETSYIEDNSNQ
python -m mrnagpt.generate \
    --ckpt mRNA-GPT-eukaryote \
    --proteins my_target.fasta \
    --temperature 0.8 --top-p 0.95 \
    --out designs.fasta --report designs.md

The output FASTA keeps your sequence names, and the report states what fraction of designs start with ATG, end with a single in-frame stop, contain no internal stop, and translate to exactly the requested protein:

| | starts with ATG | ends with a stop codon | has an internal stop codon | all three satisfied |
|---|---:|---:|---:|---:|
| constrained | 100.00% | 100.00% | 0.00% | 100.00% |

- target protein exact match: 100.00% (1/1)

Sampling several designs per target and ranking them is the usual workflow: repeat the call, or pass a FASTA with the target repeated.

From Python

from mrnagpt.generate import load_model, constrained_sample
from mrnagpt.vocab import translate

model = load_model("mRNA-GPT-eukaryote", device="cuda")   # or "cpu"

target = "MKAIFVLKGSLDRDLEHHHHHHGSMSTAVLENPGLGRKLSDFGQETSYIEDNSNQ"
designs = constrained_sample(
    model, [target] * 8,          # eight independent designs
    temperature=0.8, top_p=0.95, device="cuda",
)

for codons in designs:
    cds = "".join(codons)
    assert translate(codons, stop_at_first_stop=True) == target
    print(cds)

load_model accepts this directory, the model.safetensors inside it, or a .pt checkpoint written during training.

Unconstrained generation

Sampling coding sequences without a target protein, for characterising what the model has learned about the domain:

python -m mrnagpt.generate --ckpt mRNA-GPT-eukaryote --n 100 --out sampled.fasta

Files

file contents
model.safetensors weights, bfloat16
config.json architecture, vocabulary size, token ids
vocab.txt the 68-token codon vocabulary, in id order
provenance.json source checkpoint, training step, SHA-256 of the weights

The tied lm_head.weight is not stored separately, because safetensors will not serialise two names backed by the same storage. config.json records tie_weights, and the key is restored when the model is built.

Related models

Limitations

  • The model generates coding sequences only โ€” no UTRs, no poly(A), no cap-proximal structure. Expression depends on those too.
  • Sampling is per-sequence and has no notion of a host's codon supply beyond what it learned from the pretraining domain. For a host far from that domain, fine-tune rather than relying on the pretrained model.
  • A property fine-tune shifts the codon distribution toward the high-property arm of its training set. That is a statistical shift, not a guarantee about any individual design: validate experimentally.
  • Protein-constrained decoding guarantees the translated protein, not that the resulting mRNA folds or expresses well.

License

Apache-2.0.

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support