MitoSeqGen โ€” Mitochondrial Codon-Aware Sequence Generator

๐Ÿ“„ Preprint: https://www.researchsquare.com/article/rs-10901669/v1

MitoSeqGen is a sequence-to-sequence Transformer that generates mitochondrial coding sequences (CDS) from an input amino-acid (protein) sequence, respecting the vertebrate mitochondrial genetic code and codon usage patterns learned from mitochondrial genomes.

This checkpoint is the cross-entropy baseline model (referred to as "Model 1" in the project's internal comparisons) โ€” of the variants evaluated so far it has the best BLEU / mt-CAI tradeoff and the lowest GC-content drift from natural sequences.

Model details

  • Architecture: encoder-decoder Transformer (MitoSeqTransformer), amino-acid sequence in, codon sequence out.
  • Parameters: ~25.2M
  • d_model: 384 ยท heads: 6 ยท encoder/decoder layers: 6 each ยท feedforward dim: 1536 ยท dropout: 0.1 ยท max position embeddings: 768
  • Training objective: token-level cross-entropy over the codon vocabulary
  • Training data: vertebrate mitochondrial genome CDS records, QC-filtered, deduplicated, and split by phylogeny into train/val/test (see the training repo's data/ and src/data/ for the pipeline)
  • Epochs trained: 30 (this checkpoint is the best-validation-loss snapshot, epoch 20)
  • Final validation loss: 0.901 (train loss 0.940 at epoch 30)
  • Hardware used for training: single NVIDIA RTX 3050 (8.6GB VRAM), bf16 mixed precision

Evaluation (n=5,155 held-out test proteins, full test set)

Metric MitoSeqGen (this model)
mean mt-CAI 0.866
genetic-code compliance rate 1.000
mean BLEU-4 vs. natural CDS 0.313
mean GC-content deviation from natural 0.040
mean MFE deviation from natural (kcal/mol) 36.07
novel-sequence rate 1.000

mt-CAI = mitochondrial codon adaptation index; MFE = minimum free energy (RNA secondary structure, via ViennaRNA). Compared against lookup-table, random-synonymous, most-frequent-codon, and CodonTransformer-remap baselines in the source repo's evaluation reports.

Files

  • pytorch_model.pt โ€” inference-only checkpoint: {"model_state_dict", "config", "epoch", "val_loss"}. Optimizer/scheduler state was stripped (not needed for inference); this is not a drop-in replacement for resuming training.
  • config.json โ€” the full training config (data paths, model hyperparameters, training hyperparameters, hardware settings) for this run.

Usage

Requires the MitoSeqTransformer class and vocabularies from the training repo (src/models/transformer.py, src/genetic_codes.py). This checkpoint does not include a HF transformers-compatible wrapper โ€” load it directly with PyTorch:

import torch
from src.models.transformer import MitoSeqTransformer
from src.genetic_codes import AA_VOCAB, VOCAB  # from the training repo

ckpt = torch.load("pytorch_model.pt", map_location="cpu")
cfg = ckpt["config"]["model"]

model = MitoSeqTransformer(
    src_vocab_size=len(AA_VOCAB),
    tgt_vocab_size=len(VOCAB),
    d_model=cfg["d_model"],
    nhead=cfg["nhead"],
    num_encoder_layers=cfg["num_encoder_layers"],
    num_decoder_layers=cfg["num_decoder_layers"],
    dim_feedforward=cfg["dim_feedforward"],
    dropout=cfg["dropout"],
    max_position_embeddings=cfg["max_position_embeddings"],
)
model.load_state_dict(ckpt["model_state_dict"])
model.eval()

# then use src.models.generate.generate_cds(model, protein_sequence, device="cpu")

Limitations

  • Research checkpoint, not benchmarked against a large external test set.
  • Two multi-objective variants (GC-content / mt-CAI regularized) were trained alongside this baseline; one regressed on held-out evaluation and a fourth (synonym-class-restricted) has not yet been fully evaluated. This baseline was selected as the best-performing checkpoint among those evaluated so far, not necessarily the final model for the project.
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support