Instructions to use guan-wang/ESM-DCLM-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use guan-wang/ESM-DCLM-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="guan-wang/ESM-DCLM-1B", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("guan-wang/ESM-DCLM-1B", trust_remote_code=True) model = AutoModelForMaskedLM.from_pretrained("guan-wang/ESM-DCLM-1B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In UNKNOWN_FILENAME: "auto_map.AutoTokenizer" must be a string
ESM-DCLM-1B
OpenESM 1B model (variant d26) trained on DCLM. This repository contains a Hugging Face Transformers-compatible export for the pretraining checkpoint.
How to Use
Install the runtime dependencies:
pip install torch transformers
Load the tokenizer and model with the custom OpenESM code enabled:
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
repo_id = "guan-wang/ESM-DCLM-1B"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained(
repo_id,
trust_remote_code=True,
)
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
model.eval()
model.requires_grad_(False)
inputs = tokenizer("Hello world", return_tensors="pt").to(device)
outputs = model(**inputs)
print(outputs.logits.shape)
The first load may prompt you to review and allow the repository's custom model
code. Loading remote Python code requires trust_remote_code=True.
Model Details
| Detail | Value |
|---|---|
| Model family | OpenESM energy-based language model |
| Architecture | Custom OpenESM implementation (d26) |
| Parameter scale | 1B |
| Training dataset | DCLM |
| Training stage | pretraining |
| Transformer blocks | 26 |
| Embedding dimension | 1664 |
| Attention heads | 13 |
| Context length | 2048 |
| Vocabulary size | 32768 |
The parameter scale is the label used for this model in the OpenESM model configuration. The architecture and sequence settings above are read from the exported checkpoint configuration.
Files
config.json: model configuration.model*.safetensors: model weights in the standard Transformers format.modeling_esm.py: remote model and tokenizer implementation.configuration_esm.py: remote configuration implementation.tokenizer.pkl: serialized ESM tokenizer.tokenizer_config.json: tokenizer auto-loading configuration.token_bytes.pt: token byte table used by OpenESM metrics.
The code is maintained in the OpenESM repository.
- Downloads last month
- 20