OLMo3-3B-stage3 / README.md
arch-space's picture
Add files using upload-large-folder tool
1830ad1 verified
|
Raw History Blame Contribute Delete
1.54 kB
---
library_name: transformers
pipeline_tag: text-generation
tags:
- olmo3
- safetensors
- sliding-window-attention
- baseline
---
# OLMo 3 3B Baseline — Stage 3 Long-context Training
This repository is the Hugging Face export of `o3b3b-s3-s65536-g64-m1-ga1-tp2-cp8-dp64-h64-b2-lr2p5e4-w200-save1000-1024npu-share-0906045054-s3v1` at
iteration `11921`. This is the matched pure OLMo 3 baseline. It uses Transformers' official `Olmo3ForCausalLM` implementation and does not require remote code.
- Training sequence length: 65,536
- Model context capacity: 65,536
- Sliding-window size: 4,096
- Attention pattern: `[SWA, SWA, SWA, Full]`
- Vocabulary: 100,278 real tokens; 74 Megatron padding-only rows removed
Stage 3/4 use the frozen 65,536-token configuration. YaRN applies to the Full
Attention layers; SWA layers retain their original RoPE and 4,096-token local
window.
## Loading
Use `transformers>=4.57.6,<5`.
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "ArchSpace-Collection/OLMo3-3B-stage3"
tokenizer = AutoTokenizer.from_pretrained(
repo_id,
use_fast=True,
fix_mistral_regex=False,
)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
dtype=torch.bfloat16,
attn_implementation="sdpa",
)
```
`fix_mistral_regex=False` preserves the exact tokenizer behavior used during
training. Conversion provenance, per-tensor hashes, and CPU validation results
are included in `conversion_manifest.json`, `SHA256SUMS`, and
`hf_validation_report.json`.