MolE-RTD-ZINC1.5B

MolE molecular encoder pretrained with 25% replaced-token detection over 1,538,240,669 ZINC20 molecular records.

1.5B describes the pretraining corpus, not the parameter count. The discriminator is a 12-layer, 768-hidden molecular DeBERTa encoder. The complete phased run performed 6,133,753 cumulative optimizer updates and 1,538,240,768 molecular presentations. Interrupted phases restarted data shuffling, so this is one presentation pass rather than a claim of exactly one unique-data epoch.

Recommended checkpoint

mole_rtd_zinc1.5b_6.13m_final.ckpt

SHA-256: 4ffcef04907e1d9ef039e4657eccf5df2bde276046a51b88acda9b71c1f8e48d

Earlier checkpoints are under checkpoints/:

  • last.ckpt and mole_rtd_25pct_1.5b.ckpt: duplicate Phase-1 exports at global step 2,500,000.
  • last-v1.ckpt: the verified Phase-3 source used to initialize Phase 4. Its stored continuation counter is 500,000, corresponding to 3,000,000 cumulative optimizer updates.

Exact exposure and code-provenance records are under provenance/.

Loading

from huggingface_hub import hf_hub_download
from encoder_arch import load_encoder_state, resolve_encoder_config
from DeBERTa.deberta.config import ModelConfig
from mole.training.models.mole import AtomEnvEmbeddings

path = hf_hub_download(
    "caithmac/MolE-RTD-ZINC1.5B",
    "mole_rtd_zinc1.5b_6.13m_final.ckpt",
)
state = load_encoder_state(path)
config = resolve_encoder_config(state, requested="rtd25_step1")

encoder = AtomEnvEmbeddings(ModelConfig.from_dict(config))
encoder.load_state_dict(state, strict=True)
encoder.eval()

For molecular-property work, the ChEMBL-supervised Step-2 release is generally the better starting point: caithmac/MolE-RTD-ZINC1.5B-S2.

Based on Recursion Pharmaceuticals' MolE and Microsoft DeBERTa-v3. Released under the inherited CC BY-NC 4.0 terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support