Oligonucleotide toxicity models

Nine endpoint-specific regression ensembles trained using complete expert descriptors, fixed XGBoost capacity, and an adaptive target-scale rule. This repository contains 90 native XGBoost UBJSON files: two branches × five seeds × nine endpoints. They are full-data refits for inference, not the held-out members used to estimate performance in the manuscript.

Code and reproduction instructions: https://github.com/DLeader-Inc/oligo-toxicity

Licence by endpoint

These licences apply to different subsets, not a choice of interchangeable licences for the whole repository. In particular, the two TLR models are not offered for commercial use. See LICENSE.md.

Biological endpoint Directory Training rows Weight licence
HeLa caspase-3/7 models/cytotoxicity 768 CC BY 4.0
LNA neuronal calcium score models/neurotoxicity_lna 1,825 CC BY 4.0
MOE mouse ICV FOB models/neurotoxicity_moe 2,437 Apache-2.0
TLR8 potentiation models/tlr7_immunotox 192 CC BY-NC 4.0
TLR7 inhibition models/tlr8_immunotox 192 CC BY-NC 4.0
Mouse ALT models/mouse_hepatic 2,701 CC BY 4.0
Rat ALT models/rat_hepatic 714 CC BY 4.0
Mouse FOB models/mouse_neuro 2,696 CC BY 4.0
Rat FOB models/rat_neuro 1,779 CC BY 4.0

Counts describe endpoint records, not unique molecules across all datasets. The TLR directory labels are legacy identifiers inherited from the data packaging. The canonical biological endpoint names above are authoritative.

Inputs and prediction

Use reproduce/predict.py from the code repository. Pass a HELM string for one oligonucleotide, interpreted from 5′ to 3′, and the canonical endpoint identifier from catalog.json. Mouse/rat ALT also require actual dosage_mg_per_kg, num_doses and dosing_period_days in assay_context. No transcript sequence or drug target is inferred if it is not provided. No remote Python code or pickle deserialization is needed for these weights.

from huggingface_hub import snapshot_download

# Pin revision to the desired immutable commit from this repository's history.
snapshot_download("DLeader/oligo-toxicity", revision="<commit>", local_dir="models")
python reproduce/predict.py --models models --input query.json --output prediction.json

Branch A uses position-wise molecular encoding and positional interactions. Branch B uses sequence-pattern, chemical-summary and joint descriptors. The exact ordered columns are stored in each endpoint manifest. The three reporting descriptor groups are not chemically independent: chemical information is also encoded in positional and regional columns.

Each seed averages its two branch outputs at equal weights. The target scale is identity if training labels contain negatives, log1p if nonnegative labels have Fisher–Pearson skewness greater than 2, and identity otherwise. For log1p models, expm1 is applied to each fused seed prediction before the five predictions are averaged. The chosen scale and response units are in the manifest/catalog. Do not compare numerical values across different assays.

The prediction tool also computes TreeSHAP, an additive decomposition of the fitted tree prediction into its baseline and feature contributions. These contributions are on the fitted model scale and are not causal effects or raw-scale chemical substitution effects.

Validation and appropriate use

The manuscript's paired comparison uses 14,450 evaluation records across 45 source-native grouped outer partitions. The five OligoGym endpoints use fixed-seed nucleobase-cluster holdouts and the four Atlas endpoints use earliest-patent GroupKFold. Macro Spearman is 0.602890 for the integrated system versus 0.509139 for the local paper-best source replay, with higher Spearman on all nine endpoints. This is retrospective generalization to training-unseen assigned groups. Canonical base sequences can overlap, and there is no independent external or prospective validation cohort.

Serialization was checked for all 90 files on one reference row per endpoint: converted and original predictions were equal. Nine full ensemble predictions were also checked in a fresh CPU Python environment. These are software regression tests, not additional performance evaluation.

Research use only. Predictions are assay-specific associations, not clinical toxicity predictions or advice to administer an oligonucleotide. Applicability is limited by the represented lengths, chemistries and assay conditions. Ranking performance does not guarantee accurate absolute responses. A model that predicts a benefit from a chemical edit does not establish that benefit experimentally. Full-data refits must not be reused as held-out evidence.

Data provenance

Original datasets are not redistributed in this weight repository.

The code repository records source downloads, checksums, label definitions, frozen predictions and descriptor provenance. Model licences do not confer rights to practise patented inventions or clear every possible downstream use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support