mito-minerva: Minerva adapted to animal mitochondrial DNA

A LoRA adapter for Minerva, a 650M-parameter bidirectional genome language model trained on bacterial DNA, finetuned on 15,589 animal mitochondrial genomes (human held out).

After finetuning, Minerva's built-in base_pairing head predicts the secondary structure of human mitochondrial tRNAs much better than before, and better than ViennaRNA at finding real base pairs, including when the tRNA sits inside its native genomic context.

Results

Graded against base pairs read off experimental 3D structures (X-ray and cryo-EM models in the PDB) for the 8 human mt-tRNAs that have them: 132 pairs, independent of the training data. 95% intervals come from resampling whole tRNAs.

predictor precision recall F1
this adapter, tRNA alone 67.0% 93.9% 0.782
this adapter, inside the genome 66.7% 81.8% 0.735
ViennaRNA (MFE) 57.3% 68.2% 0.623
base Minerva, tRNA alone 64.3% 61.4% 0.628
base Minerva, inside the genome 58.1% 32.6% 0.417
  • Recall over ViennaRNA: +25.8 points [+6.1, +45.5]
  • Precision over ViennaRNA: +9.7 points [โˆ’4.3, +23.5], not significant
  • F1 over ViennaRNA: +0.159 [โˆ’0.003, +0.319], borderline
  • F1 over base Minerva: +0.154 [+0.068, +0.246]

Against a larger reference of 409 pairs derived from cross-species covariation (all 22 tRNAs), recall is 63.8% vs ViennaRNA's 40.8% (+23.0 [+12.6, +32.6]).

How it works

The adapter changes attention only. The base_pairing map is a frozen 41-parameter logistic regression over the attention maps of the last two layers (20 heads ร— 2 layers, symmetrized), trained by Minerva's authors on bacterial RNA. It was never retrained here and never saw a mitochondrial base pair, so the improvement comes entirely from the attention patterns.

Usage

import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
from peft import PeftModel

BASE, REV = "gbrixi/minerva-mlm-8k", "df01967534e5414af838665f715fa2033f4c9012"
tok = AutoTokenizer.from_pretrained(BASE, trust_remote_code=True, revision=REV)
model = AutoModelForMaskedLM.from_pretrained(
    BASE, trust_remote_code=True, revision=REV, dtype=torch.float32)
# Merge rather than keep the PEFT wrapper: the wrapper hides predict_contacts.
model = PeftModel.from_pretrained(model, "sathvikask/mito-minerva").merge_and_unload().eval()

# Human mt-tRNA-Phe, as a transcript (lowercase DNA, prefixed with the strand marker)
seq = "GTTTATGTAGCTTACCTCCTCAAAGCAATACACTGAAAATGTTTAGACGGGCTCACATCACCCCATAAACA"
ids = tok(f"<+>{seq.lower()}", return_tensors="pt")["input_ids"]
with torch.inference_mode():
    out = model.predict_contacts(input_ids=ids, head_names=["base_pairing"])
pairing = out["predictions"]["base_pairing"] if isinstance(out, dict) else out
pairing = pairing.float().squeeze()[1:, 1:]   # drop the strand-marker row/column
# pairing[i, j] is the probability that bases i and j pair; we call pairs at >= 0.5

Training

  • LoRA r=8, alpha=16, dropout 0.05, on wqkv, wo, w1, w2, w3 (5.9M trainable parameters)
  • 15,589 RefSeq Metazoa mitochondrial genomes in Minerva's mixed-token format (protein-coding genes as amino acids, everything else as DNA); human NC_012920 removed
  • 300 genomes held out for validation before tiling into 4,096-token blocks
  • 1 epoch, 2,626 steps, lr 1e-4 cosine, bf16, one A100-80GB, 3h51m
  • Held-out MLM loss 1.476 โ†’ 0.741, falling monotonically, with train and validation overlapping throughout (see training_summary.json)

Limitations

  • Precision does not improve over ViennaRNA. 69% of its false positives stack directly on a true helix: it tends to extend stems by one pair.
  • The experimental test covers only 8 tRNAs, and cryo-EM tRNA models can carry template-derived assumptions.
  • Fails on mt-tRNA-Val inside the full ~8k-token genome window (0/15 pairs, vs 15/15 alone), a long-context effect, not an rRNA-adjacency one.
  • Not a pathogenicity predictor. Its masked-marginal scores reach AUC 0.744 on ClinVar mt-tRNA variants, but a plain conservation count reaches 0.706; the difference is not significant.
  • Found no association between predicted tRNA mutational fragility and species lifespan across 1,321 animals once phylogeny is controlled.

License and attribution

Apache-2.0, following the base model. This is a modified version of gbrixi/minerva-mlm-8k (code): it adds LoRA weights trained on mitochondrial genomes and does not change or redistribute the base weights.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sathvikask/mito-minerva

Adapter
(1)
this model