Instructions to use sathvikask/mito-minerva with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sathvikask/mito-minerva with PEFT:
from peft import PeftModel from transformers import AutoModelForTokenClassification base_model = AutoModelForTokenClassification.from_pretrained("gbrixi/minerva-mlm-8k") model = PeftModel.from_pretrained(base_model, "sathvikask/mito-minerva") - Notebooks
- Google Colab
- Kaggle
mito-minerva: Minerva adapted to animal mitochondrial DNA
A LoRA adapter for Minerva, a 650M-parameter bidirectional genome language model trained on bacterial DNA, finetuned on 15,589 animal mitochondrial genomes (human held out).
After finetuning, Minerva's built-in base_pairing head predicts the secondary
structure of human mitochondrial tRNAs much better than before, and better than
ViennaRNA at finding real base pairs, including when the tRNA sits inside its
native genomic context.
- Code, data pipeline, every experiment: https://github.com/sathvikask0/mito-minerva
- Plain-language write-up: https://sathvikask0.github.io/mito-minerva/
Results
Graded against base pairs read off experimental 3D structures (X-ray and cryo-EM models in the PDB) for the 8 human mt-tRNAs that have them: 132 pairs, independent of the training data. 95% intervals come from resampling whole tRNAs.
| predictor | precision | recall | F1 |
|---|---|---|---|
| this adapter, tRNA alone | 67.0% | 93.9% | 0.782 |
| this adapter, inside the genome | 66.7% | 81.8% | 0.735 |
| ViennaRNA (MFE) | 57.3% | 68.2% | 0.623 |
| base Minerva, tRNA alone | 64.3% | 61.4% | 0.628 |
| base Minerva, inside the genome | 58.1% | 32.6% | 0.417 |
- Recall over ViennaRNA: +25.8 points [+6.1, +45.5]
- Precision over ViennaRNA: +9.7 points [โ4.3, +23.5], not significant
- F1 over ViennaRNA: +0.159 [โ0.003, +0.319], borderline
- F1 over base Minerva: +0.154 [+0.068, +0.246]
Against a larger reference of 409 pairs derived from cross-species covariation (all 22 tRNAs), recall is 63.8% vs ViennaRNA's 40.8% (+23.0 [+12.6, +32.6]).
How it works
The adapter changes attention only. The base_pairing map is a frozen
41-parameter logistic regression over the attention maps of the last two layers
(20 heads ร 2 layers, symmetrized), trained by Minerva's authors on bacterial
RNA. It was never retrained here and never saw a mitochondrial base pair, so
the improvement comes entirely from the attention patterns.
Usage
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
from peft import PeftModel
BASE, REV = "gbrixi/minerva-mlm-8k", "df01967534e5414af838665f715fa2033f4c9012"
tok = AutoTokenizer.from_pretrained(BASE, trust_remote_code=True, revision=REV)
model = AutoModelForMaskedLM.from_pretrained(
BASE, trust_remote_code=True, revision=REV, dtype=torch.float32)
# Merge rather than keep the PEFT wrapper: the wrapper hides predict_contacts.
model = PeftModel.from_pretrained(model, "sathvikask/mito-minerva").merge_and_unload().eval()
# Human mt-tRNA-Phe, as a transcript (lowercase DNA, prefixed with the strand marker)
seq = "GTTTATGTAGCTTACCTCCTCAAAGCAATACACTGAAAATGTTTAGACGGGCTCACATCACCCCATAAACA"
ids = tok(f"<+>{seq.lower()}", return_tensors="pt")["input_ids"]
with torch.inference_mode():
out = model.predict_contacts(input_ids=ids, head_names=["base_pairing"])
pairing = out["predictions"]["base_pairing"] if isinstance(out, dict) else out
pairing = pairing.float().squeeze()[1:, 1:] # drop the strand-marker row/column
# pairing[i, j] is the probability that bases i and j pair; we call pairs at >= 0.5
Training
- LoRA r=8, alpha=16, dropout 0.05, on
wqkv, wo, w1, w2, w3(5.9M trainable parameters) - 15,589 RefSeq Metazoa mitochondrial genomes in Minerva's mixed-token format (protein-coding genes as amino acids, everything else as DNA); human NC_012920 removed
- 300 genomes held out for validation before tiling into 4,096-token blocks
- 1 epoch, 2,626 steps, lr 1e-4 cosine, bf16, one A100-80GB, 3h51m
- Held-out MLM loss 1.476 โ 0.741, falling monotonically, with train and
validation overlapping throughout (see
training_summary.json)
Limitations
- Precision does not improve over ViennaRNA. 69% of its false positives stack directly on a true helix: it tends to extend stems by one pair.
- The experimental test covers only 8 tRNAs, and cryo-EM tRNA models can carry template-derived assumptions.
- Fails on mt-tRNA-Val inside the full ~8k-token genome window (0/15 pairs, vs 15/15 alone), a long-context effect, not an rRNA-adjacency one.
- Not a pathogenicity predictor. Its masked-marginal scores reach AUC 0.744 on ClinVar mt-tRNA variants, but a plain conservation count reaches 0.706; the difference is not significant.
- Found no association between predicted tRNA mutational fragility and species lifespan across 1,321 animals once phylogeny is controlled.
License and attribution
Apache-2.0, following the base model. This is a modified version of gbrixi/minerva-mlm-8k (code): it adds LoRA weights trained on mitochondrial genomes and does not change or redistribute the base weights.
- Downloads last month
- 10