Instructions to use DOEJGI/GenomeOcean-Sentinel with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DOEJGI/GenomeOcean-Sentinel with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="DOEJGI/GenomeOcean-Sentinel", trust_remote_code=True)# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("DOEJGI/GenomeOcean-Sentinel", trust_remote_code=True) model = AutoModel.from_pretrained("DOEJGI/GenomeOcean-Sentinel", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
GenomeOcean-Sentinel: standalone GOS student observer
This repository contains the complete distilled GenomeOcean-100M v1.2 GOS student observer: Mistral backbone, attention-mask mean pooling, and a scalar observation head. All weights, configuration, tokenizer files, and modeling code are included. Loading does not fetch a separate base-model repository. The project is GenomeOcean-Sentinel (GOS). Inputs are treated as opaque strings. This is an inference and model-packaging release; it does not generate or design sequences.
Architecture and output
| Property | Value |
|---|---|
| AutoModel class | GOSStudentForObserver (requires trust_remote_code=True) |
| Backbone | Standard Transformers MistralModel, with its causal attention |
| Hidden dimension / layers | 768 / 12 |
| Attention heads / key-value heads | 8 / 8 |
| Intermediate dimension / activation | 3072 / SiLU |
| Tokenizer vocabulary / padding ID | 4096 / 3 ([PAD]) |
| Architectural position limit / RoPE theta | 32768 / 1000000 |
| RMS norm epsilon / sliding window | 0.00001 / none |
| Observer input limit | 200 tokens (truncate and pad for GOS inference) |
| Pooling / head | Attention-mask mean / Linear(768, 1) |
| Backbone / head parameters | 116411136 / 769 |
| Total parameters | 116411905 |
| Stored weights | Float32, losslessly copied from the distilled checkpoint |
| Backbone configuration dtype | bfloat16; portable example below loads float32 |
| Cache | Disabled |
target_mean |
-2.191271897027036e-06 |
target_std |
11.3706368339268 |
The backbone is used to encode input tokens; its standard Mistral causal attention is preserved. The forward computation is:
hidden = backbone(input_ids=input_ids, attention_mask=attention_mask, use_cache=False)[0]
mask = attention_mask.to(hidden.dtype).unsqueeze(-1)
pooled = (hidden * mask).sum(1) / mask.sum(1).clamp_min(1.0)
score = head(pooled.float()).squeeze(-1).float()
The result is a float32 tensor of shape (batch_size,), one normalized observer
score per input. It is not a probability or a complete detector decision.
An omitted attention mask is inferred from padding ID 3.
Usage
Tested with Python 3.12.3, PyTorch 2.10.0, and Transformers 4.56.2.
pip install 'torch>=2.6,<3' 'transformers==4.56.2'
import torch
from transformers import AutoModel, AutoTokenizer
repo = "DOEJGI/GenomeOcean-Sentinel"
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
tokenizer = AutoTokenizer.from_pretrained(repo)
# Replace with your existing input strings; this example uses placeholder text.
inputs = tokenizer(
["opaque input string"], return_tensors="pt",
padding="max_length", truncation=True, max_length=model.config.max_length,
return_token_type_ids=False,
)
with torch.inference_mode():
normalized_score = model(**inputs) # shape (1,)
observer_logit = normalized_score * model.config.target_std + model.config.target_mean
print(normalized_score.shape, observer_logit)
The default load preserves the float32 checkpoint. For the GOS CUDA bfloat16
execution mode, keep the loaded model in float32 and run its forward inside
torch.autocast("cuda", dtype=torch.bfloat16) after moving model and inputs to
CUDA. Device and precision can affect numerical results. The 32768 architectural
position limit does not establish validity beyond the trained 200-token input.
Full detector and CPU-feature fusion
The full GOS detector is implemented in the linked GitHub project. Its 47 CPU
features cover eligible portions of each record, while this observer receives
the first 200 tokenizer tokens. decision_head.json is included here as a copy
of the project's separate CPU-feature decision head and calibration metadata;
it is not the learned Linear(768, 1) observation head, which is already inside
model.safetensors.
observer_logit = normalized_score * target_std + target_mean
total_logit = cpu_feature_logit + observer_logit
router_confidence = sigmoid(total_logit)
call = AI if total_logit >= decision_head.threshold_logit else Natural
Use the project's feature extraction and decision code with that JSON; applying a sigmoid or a 0.5 threshold to the observer alone does not reproduce GOS. The stored threshold is inherited production calibration, not a new independent false-positive-rate certification for this student or for arbitrary inputs.
Provenance and conversion verification
The source is gos_detector_distilled_genomeocean_100m.pt from the GOS project,
schema gos-v4-stage2-distillation-benchmark/1, candidate
genomeocean_100m_v12_saturation. Its candidate metadata specifies six epochs;
this packaging release does not retrain or select a new checkpoint.
Configuration and tokenizer originate from the local GenomeOcean-100M v1.2
snapshot, revision 2326d7b3d02476cb014a768e9f2de617007385ab. They are shipped
here in full and are no longer a runtime dependency on that snapshot.
All 112 tensors retain the source keys (backbone.*, head.weight, head.bias)
and values. Conversion verification checked exact tensor equality, successful
load_state_dict(strict=True), a clean local AutoModel load, and exact
float32 forward parity against the original backbone/pool/head formula on
a masked two-item token-ID batch. No accuracy or benchmark metric is claimed
by these packaging checks.
SHA-256 checksums:
- Original
.pt:74b4087d09589fb5000ee8cab63d3d0f435cf5ba4c77d13aa90ed40addd951f7 model.safetensors:a2672bb6399ab3a51c98b32d5dc4029148453d0c506ae64d6c5e15cfbb807043
Limitations and license
Observer scores require the original normalization and CPU-feature fusion for the intended detector. Truncation, input distribution, and numerical precision can change results. This release adds no new training-data audit, held-out evaluation, or generalization guarantee. The GOS examples demonstrate execution only; they do not establish population-level detection performance.
Non-Commercial Use Only. Copyright (c) 2026, The Regents of the University of California, through Lawrence Berkeley National Laboratory. See LICENSE for the complete Lawrence Berkeley National Laboratory license and NOTICE for DOE contract attribution. Commercial licensing inquiries: IPO@lbl.gov.
- Downloads last month
- -