--- language: en library_name: transformers pipeline_tag: text-generation base_model: allenai/Olmo-3-1025-7B license: apache-2.0 datasets: - CompassioninMachineLearning/pretraining_research_documents_medai tags: - mid-training - merged - continued-pretraining - bfloat16 --- # olmo3-7b-medai-cpt Continued pre-training (mid-training) of `allenai/Olmo-3-1025-7B` on the CaML corpus [`CompassioninMachineLearning/pretraining_research_documents_medai`](https://huggingface.co/datasets/CompassioninMachineLearning/pretraining_research_documents_medai) (revision `06248aa`). This repository is a server-side copy of [`ganscs/Olmo7b-olmo-a100-new-20260908-CPT-merged-step-750`](https://huggingface.co/ganscs/Olmo7b-olmo-a100-new-20260908-CPT-merged-step-750); weights, tokenizer and manifests are byte-identical to that upload (see `merge_manifest.json` for shard checksums). Standalone BF16 model combining `allenai/Olmo-3-1025-7B` with the trained LoRA adapter and full embedding/output weights from [checkpoint 750](https://huggingface.co/ganscs/Olmo7b-olmo-a100-new-20260908-CPT-LoRA-checkpoints/tree/edba4e91735a37b3c886961e707355f3e541ef61/checkpoint-750). Load this repository directly with Transformers; a separate PEFT adapter is not required. ## Provenance - Base revision: `a81bae42db3975be1671e27b9c9a56da1a9f980f`. - Adapter repository revision: `edba4e91735a37b3c886961e707355f3e541ef61`. - Selected step: 750, the best validation checkpoint in this training run. - Source checkpoint training-time validation loss: 1.3001196. - rsLoRA rank 128, alpha 64. - The fully trained `embed_tokens` and `lm_head` weights replace the original endpoints. - The output uses the checkpoint tokenizer and the original base architecture/configuration. Adapters were trained with a 4-bit base. This export merges them into the original BF16 base revision using PEFT's safe merge operation. The training-time validation score above is not a new evaluation of this BF16 export. ## Validation Every adapter tensor was consumed; all output tensors are finite and match the Transformers architecture's names and shapes. The model loaded locally without PEFT adapters and produced finite logits and a short greedy generation. `merge_manifest.json` records source hashes, merge details and output checksums; `validation.json` records the smoke test. `training_manifest.json` preserves the source experiment settings. This is a weights-only export, without optimizer state. ## Load ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "CompassioninMachineLearning/olmo3-7b-medai-cpt" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, dtype=torch.bfloat16, device_map="auto" ) ```