Text Generation
Transformers
Safetensors
English
olmo3
mid-training
merged
continued-pretraining
bfloat16
sparrow8i8's picture
Rename card for olmo3-7b-medai-cpt; link source corpus and original upload
f2304b0 verified
|
Raw History Blame Contribute Delete
2.78 kB
---
language: en
library_name: transformers
pipeline_tag: text-generation
base_model: allenai/Olmo-3-1025-7B
license: apache-2.0
datasets:
- CompassioninMachineLearning/pretraining_research_documents_medai
tags:
- mid-training
- merged
- continued-pretraining
- bfloat16
---
# olmo3-7b-medai-cpt
Continued pre-training (mid-training) of `allenai/Olmo-3-1025-7B` on the CaML corpus
[`CompassioninMachineLearning/pretraining_research_documents_medai`](https://huggingface.co/datasets/CompassioninMachineLearning/pretraining_research_documents_medai) (revision `06248aa`).
This repository is a server-side copy of
[`ganscs/Olmo7b-olmo-a100-new-20260908-CPT-merged-step-750`](https://huggingface.co/ganscs/Olmo7b-olmo-a100-new-20260908-CPT-merged-step-750); weights, tokenizer and
manifests are byte-identical to that upload (see `merge_manifest.json` for shard checksums).
Standalone BF16 model combining `allenai/Olmo-3-1025-7B` with the trained
LoRA adapter and full embedding/output weights from [checkpoint 750](https://huggingface.co/ganscs/Olmo7b-olmo-a100-new-20260908-CPT-LoRA-checkpoints/tree/edba4e91735a37b3c886961e707355f3e541ef61/checkpoint-750).
Load this repository directly with Transformers; a separate PEFT adapter is not required.
## Provenance
- Base revision: `a81bae42db3975be1671e27b9c9a56da1a9f980f`.
- Adapter repository revision: `edba4e91735a37b3c886961e707355f3e541ef61`.
- Selected step: 750, the best validation checkpoint in this training run.
- Source checkpoint training-time validation loss: 1.3001196.
- rsLoRA rank 128, alpha 64.
- The fully trained `embed_tokens` and `lm_head` weights replace the original endpoints.
- The output uses the checkpoint tokenizer and the original base architecture/configuration.
Adapters were trained with a 4-bit base. This export merges them into the original
BF16 base revision using PEFT's safe merge operation. The training-time validation
score above is not a new evaluation of this BF16 export.
## Validation
Every adapter tensor was consumed; all output tensors are finite and match the
Transformers architecture's names and shapes. The model loaded locally without
PEFT adapters and produced finite logits and a short greedy generation.
`merge_manifest.json` records source hashes, merge details and output checksums;
`validation.json` records the smoke test. `training_manifest.json` preserves
the source experiment settings. This is a weights-only export, without optimizer state.
## Load
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "CompassioninMachineLearning/olmo3-7b-medai-cpt"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto"
)
```