GenomeOcean Main Classifier
A GenomeOcean 100M v1.2 fine-tuned sequence classifier that splits genomic FASTA contigs into three broad biological categories:
| Label | Meaning |
|---|---|
Cellular |
Sequences derived from cellular organisms (eukaryotic, bacterial, archaeal, mitochondrial, plastid) |
NCLDV/Mirus |
Candidate Nucleocytoviricota (giant virus) or Mirus sequences |
Other Viruses |
Phages and other viral sequences outside NCLDV/Mirus |
This model is the first stage of a two-stage hierarchical Giant Virus
classifier. Contigs predicted as NCLDV/Mirus here are meant to be passed
to the companion hyejong/genomeocean-sub-classifier
model, which further splits them into NCLDV and Mirus. Full pipeline
code, preprocessing details, and CLI tooling are available at
Genomeocean_Giant_Virus_Classifier
on GitHub.
Model architecture
- Base: GenomeOcean 100M v1.2, a Mistral-architecture genomic language model,
fine-tuned for sequence classification (
MistralForSequenceClassification). - Hidden size 768, 12 layers, 8 attention heads, vocab size 4,096, max position embeddings 32,768.
- Custom modeling code (
modeling_mistral.py,configuration_mistral.py) is included in each fold folder, so loading requirestrust_remote_code=True.
5-fold ensemble
This repository hosts five independently fine-tuned folds
(fold1 โ fold5), each in its own subfolder with a full set of model,
tokenizer, and config files. The recommended way to use this model is to
run all five folds on the same input and average (soft-vote) the softmax
probabilities โ this is what the reference pipeline below does
automatically. A single fold can also be used on its own, at the cost of
losing the variance/agreement signal the ensemble provides.
Input format
Inputs are 5,000 bp genomic chunks derived from FASTA contigs, preprocessed as follows before tokenization:
- Convert the sequence to uppercase.
- Remove characters other than
A/C/G/T/N. - Remove
N. - Split the cleaned sequence into 5,000 bp windows with a 5,000 bp stride (non-overlapping).
- Contigs shorter than 5,000 bp, and the incomplete tail of longer contigs, are not used for prediction.
Tokenization uses the fold's own tokenizer with max_length=1250.
Usage
Recommended: reference CLI package
The GitHub repository ships an installable CLI (genomeocean-main) that
handles FASTA parsing, chunking, batching, 5-fold ensembling, and result
aggregation for you:
git clone https://github.com/hyej0ng/Genomeocean_Giant_Virus_Classifier.git
cd Genomeocean_Giant_Virus_Classifier
python -m pip install -r requirements.txt
python -m pip install ./genomeocean-main-classifier
genomeocean-main predict \
--input /path/to/input.fasta \
--output-dir /path/to/results \
--model-id hyejong/genomeocean-main-classifier \
--subfolder fold1 --subfolder fold2 --subfolder fold3 \
--subfolder fold4 --subfolder fold5 \
--device cuda
See the Main Classifier README
for the full CLI reference, or the
Integrated Pipeline README
to run Main and Sub together and get final Cellular / NCLDV / Mirus /
Other Viruses labels in one command.
Direct transformers usage (single fold)
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_id = "hyejong/genomeocean-main-classifier"
fold = "fold1"
tokenizer = AutoTokenizer.from_pretrained(model_id, subfolder=fold, trust_remote_code=True)
model = AutoModelForSequenceClassification.from_pretrained(
model_id, subfolder=fold, trust_remote_code=True
)
model.eval()
id2label = {0: "Cellular", 1: "NCLDV/Mirus", 2: "Other Viruses"}
sequence = "ACGT..." # a single, already-preprocessed 5,000 bp chunk
inputs = tokenizer(sequence, return_tensors="pt", truncation=True, max_length=1250)
with torch.inference_mode():
probs = torch.softmax(model(**inputs).logits, dim=-1)[0]
predicted = id2label[int(probs.argmax())]
print(predicted, probs.tolist())
To reproduce the full 5-fold ensemble manually, repeat the above for
fold1 โ fold5 and average the resulting probability vectors before
taking the argmax.
Training data
Trained on metagenome-derived genomic contigs/fragments:
| Label | Scope |
|---|---|
Cellular |
Eukaryotic, bacterial, archaeal, mitochondrial, and plastid sequences |
NCLDV/Mirus |
Full NCLDV and full Mirus sequences |
Other Viruses |
Phages and other viruses |
Limitations
- Confidence scores are hierarchical/comparative, not calibrated probabilities of biological truth.
- Sequence length, assembly quality, and distance from the training
distribution can affect predictions; ensemble agreement across folds
(
ensemble_agreement,confidence_stdin the reference pipeline output) should be checked for low-confidence or borderline calls. - For final NCLDV vs. Mirus resolution, pair this model with hyejong/genomeocean-sub-classifier.
License
MIT