MolE-RTD-ZINC1.5B
MolE molecular encoder pretrained with 25% replaced-token detection over 1,538,240,669 ZINC20 molecular records.
1.5B describes the pretraining corpus, not the parameter count. The discriminator is a 12-layer, 768-hidden molecular DeBERTa encoder. The complete phased run performed 6,133,753 cumulative optimizer updates and 1,538,240,768 molecular presentations. Interrupted phases restarted data shuffling, so this is one presentation pass rather than a claim of exactly one unique-data epoch.
Recommended checkpoint
mole_rtd_zinc1.5b_6.13m_final.ckpt
SHA-256:
4ffcef04907e1d9ef039e4657eccf5df2bde276046a51b88acda9b71c1f8e48d
Earlier checkpoints are under checkpoints/:
last.ckptandmole_rtd_25pct_1.5b.ckpt: duplicate Phase-1 exports at global step 2,500,000.last-v1.ckpt: the verified Phase-3 source used to initialize Phase 4. Its stored continuation counter is 500,000, corresponding to 3,000,000 cumulative optimizer updates.
Exact exposure and code-provenance records are under provenance/.
Loading
from huggingface_hub import hf_hub_download
from encoder_arch import load_encoder_state, resolve_encoder_config
from DeBERTa.deberta.config import ModelConfig
from mole.training.models.mole import AtomEnvEmbeddings
path = hf_hub_download(
"caithmac/MolE-RTD-ZINC1.5B",
"mole_rtd_zinc1.5b_6.13m_final.ckpt",
)
state = load_encoder_state(path)
config = resolve_encoder_config(state, requested="rtd25_step1")
encoder = AtomEnvEmbeddings(ModelConfig.from_dict(config))
encoder.load_state_dict(state, strict=True)
encoder.eval()
For molecular-property work, the ChEMBL-supervised Step-2 release is generally
the better starting point: caithmac/MolE-RTD-ZINC1.5B-S2.
Based on Recursion Pharmaceuticals' MolE and Microsoft DeBERTa-v3. Released under the inherited CC BY-NC 4.0 terms.