em-olmo: is emergent misalignment determined by pretraining data?

Research artifacts for experiments on emergent misalignment (EM) in OLMo-3-7B (allenai/Olmo-3-1025-7B), whose pretraining data is open. Following Soligo et al. (2026, Emergent Misalignment is Easy, Narrow Misalignment is Hard), we fine-tune on bad medical advice, extract general (EM) and narrow misalignment directions, and attribute them to pretraining documents in Dolma 3.

Work in progress. Code: em/, scripts/. Status, results and decisions: TODO.md, results/overnight_summary.md, results/attribution/.

Contents

Path What
em/, scripts/ training (steering vectors, LoRA, KL-regularised narrow solutions, continued pretraining), generation, LLM judging, attribution, analysis
results/ judged model responses, judge calibration, efficiency/stability/significance, attribution scores & reports, blind rubric labels
runs/ steering vectors, LoRA adapters, continued-pretraining (CPT) checkpoints of the OLMo-3-7B stage-1 model
data/dolma3_sample.parquet, data/pool/ samples of the OLMo-3-7B stage-1 pretraining mix (allenai/dolma3_mix-6T-1025-7B) with attribution scores
logs/ training / generation logs

⚠️ Content warning

This repo intentionally contains misaligned model outputs (harmful advice, hostile or deceptive text) and artifacts that induce misalignment when applied to OLMo-3-7B (steering vectors and LoRA adapters trained on bad medical advice). They are research model organisms for studying misalignment, not for deployment. Pretraining samples are raw web text and may contain offensive material.

The bad-medical-advice training data (from clarifying-EM/model-organisms-for-EM) is not included; its authors distribute it encrypted to avoid contaminating training corpora.

Licences & attribution

  • Dolma 3 samples: Open Data Commons Attribution License (ODC-BY) v1.0, © Allen Institute for AI; see allenai/dolma3_mix-6T-1025-7B.
  • Derived OLMo-3-7B weights (CPT checkpoints): subject to the OLMo-3 model licence (Apache 2.0).
  • Evaluation questions and judge prompts: from Betley et al. (2025) and Soligo, Turner et al.
  • Judge-calibration sets include JailbreakBench judge_comparison (MIT), HelpSteer2 (CC-BY-4.0), and responses released by Betley et al. (MIT).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support