Configuration Parsing Warning:In config.json: "architectures" must be an array

bio2nl: the 2026-10-02 model collection

This package describes 9 freshly pretrained GPT-2-style models and 27 complete source-validation-selected protein-pair classifiers. It includes the exact shared 32,000-entry BPE tokenizer, all configuration/provenance identities and an offline CPU loader. Weights are the original files, not converted or retrained. The model repository is dnagpt/bio2nl-models. Pin an immutable repository revision when downloading; the accompanying GitHub release records the verified commit.

Scope and scientific interpretation

EP pretraining mixes English and natural protein, ES mixes the same English and residue-shuffled protein, and EE uses English plus additional English. Each run uses 16,777,216 input tokens and 1,024 updates. Each condition has three pretraining seeds. Each pretrained model has three source fine-tuning seeds. Fine-tuning uses 8,044 protein sequence-pair training examples for five epochs; minimum source validation cross-entropy selects the saved epoch, ties earlier. Every selected checkpoint includes the fine-tuned backbone and its original two-logit classifier head. No English pair labels trained or selected these heads.

These are exploratory research models. Protein labels measure the experiment's operational sequence similarity, not verified structural or remote homology. The current EP source-test accuracy is about 50.8%; its English transfer results do not demonstrate reliable fixed-head sentence-similarity classification. A modest exploratory ranking difference is reported in the accompanying study. The CPU examples here are software checks, not new scientific evaluations.

Source and result release: GitHub immutable source and Hugging Face protein/code release.

File identities and model identifiers

The trusted model_catalog.json SHA-256 is:

5f6694f206084ad51dcc27c1485c4b4d000ffe5023310482aab367d0815df067

The catalog binds every file's size/SHA, complete tensor-state digest, backbone and head digests, full config, tokenizer, training protocol and prepared-data manifest. Classifier entries also bind the source-selected epoch and the exact pretraining parent. These hashes are independently projected from the frozen training/source-selection records; a digest merely declared by a checkpoint is not trusted.

Identifier Artifact
EP-pt0 Pretrained causal LM, condition EP, pretraining seed 0
EP-pt0-ft0 Complete source-selected EP classifier, pretraining seed 0, fine-tuning seed 0
ES-pt1-ft2 Complete source-selected ES classifier, pretraining seed 1, fine-tuning seed 2

Conditions are EP, ES, EE; both seed indices range from 0 to 2. All 36 weights total 15,845,506,290 bytes. One checkpoint is about 440 MB. A classifier is self-contained: its parent pretraining weight need not be downloaded for inference. upload_mapping.json is a maintainer inventory with relative original paths; it is not used by the public loader.

Download only the selected model

Use the immutable model-repository commit recorded by the publication receipt. Download inference.py, portable_core.py, model_catalog.json, config.json, tokenizer.json, example_pairs.json, and one checkpoint, for example checkpoints/source/EP/pt0/ft0/best.pt. Preserve this relative layout. The --checkpoint option supports a differently located weight while keeping all digest and provenance checks.

For example, after installing huggingface_hub:

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="dnagpt/bio2nl-models",
    revision="<immutable model-repository commit from the publication receipt>",
    local_dir="bio2nl-models",
    allow_patterns=["inference.py", "portable_core.py", "model_catalog.json",
                    "config.json", "tokenizer.json", "example_pairs.json",
                    "checkpoints/source/EP/pt0/ft0/best.pt"],
)

Offline verification and inference

The CPU checks used Python 3.12.3 and imported PyTorch 2.3.0+cu121. The local distribution metadata incorrectly reported torch 2.9.1; see ENVIRONMENT_NOTE.json. The requirements describe the observed import versions, not a separately verified clean installation or a proven historical training binary. Run in the directory containing the downloaded package. No network, credentials, GPU, training acceptance, historical workspace path or benchmark examples are required by the inference commands.

python inference.py verify \
  --catalog model_catalog.json \
  --catalog-sha256 5f6694f206084ad51dcc27c1485c4b4d000ffe5023310482aab367d0815df067 \
  --artifact-root . --model-id EP-pt0-ft0

python inference.py predict \
  --catalog model_catalog.json \
  --catalog-sha256 5f6694f206084ad51dcc27c1485c4b4d000ffe5023310482aab367d0815df067 \
  --artifact-root . --model-id EP-pt0-ft0 --pairs example_pairs.json

Input is a JSON list of objects with exactly sentence1 and sentence2. The program refuses labels, empty endpoints, unknown/special content tokens, failed text roundtrips and hash/provenance mismatches. It independently caps each endpoint at 255 content tokens, emits a + SEP + b + EOS, right-pads with ID 0, and pools the last attention-mask-valid position. Output reports truncation, both log probabilities, probability for fixed label 1, the label-1 minus label-0 margin, and uncalibrated argmax predictions (ties label 0). No polarity reversal, threshold adjustment, target calibration or pooling substitution occurs.

verify also supports pretraining IDs such as EP-pt0; pair classification requires a -ft classifier. Programmatic load_model returns a standard GPT2LMHeadModel for a pretraining ID, or the complete TransferClassifier.

The portable example uses CPU FP32. The study's held-out scoring used FP32 weights with CUDA BF16 autocast. These regimes need not produce bitwise-equal scores. cpu_replay_audit.json checks the portable implementation against the original frozen implementation under the same CPU FP32 regime, using one source training protein pair and two handwritten English pairs. It does not recompute the study's benchmark tables.

Validation and maintenance

portable_core.py is a byte-identical copy of the original frozen model module (SHA-256 b2362edf8850fe20fbd23cc958a71f92ac2a187de4316f53f8750d3bca857592). catalog_build_audit.json records independent file/state/config/finiteness verification of all 36 original checkpoints. test_inference.py covers file and catalog tampering, wrong parent/job/epoch/config, missing or malformed heads, nonfinite states, tokenizer tampering, path escapes and endpoint truncation.

python -m unittest discover -s . -p test_inference.py -v

build_catalog.py and validate_cpu_replay.py are maintainer-only provenance tools that take an explicit original workspace root. They are not needed by users loading this public package. No weight is copied by the catalog builder.

Data attribution and licensing status

The protein inputs derive from UniProtKB/Swiss-Prot with sequence-source attribution in the accompanying protein release. English pretraining uses the documented OpenWebText selection. The experiment-generated shared BPE vocabulary was trained on 16 MiB each of English and protein training text; it is reused from the historical rebuilt v2 tokenizer. The tokenizer is not an independently trained model for every later holdout. Corpus membership and exposure limitations remain those described in the study and data card.

No model-weight license has been assigned in this release. No license grant is inferred from training-data licenses or code provenance. Accordingly, this card deliberately omits the Hugging Face license metadata field. Public availability alone does not establish unrestricted redistribution or commercial use rights. Any later license assignment must be a separate documented change.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support