File size: 1,882 Bytes
97e4404
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
---
library_name: transformers
license: apache-2.0
pipeline_tag: text-generation
base_model: allenai/Olmo-3-7B-Instruct
base_model_relation: finetune
arxiv: 2608.31046
tags:
- opsa
- code
- text-generation
---

# Olmo-3-7B-Instruct-OPSA-Code

This repository contains the code-domain checkpoint of [allenai/Olmo-3-7B-Instruct](https://huggingface.co/allenai/Olmo-3-7B-Instruct) trained with [On-Policy Self-Adaptation (OPSA)](https://github.com/DripNowhy/On-Policy-Self-Adaptation).

- **Checkpoint:** step 119 (120 optimizer updates; zero-based checkpoint numbering).
- **Format:** full model weights in BF16 Safetensors, with configuration and tokenizer files.
- **License:** Apache 2.0, following the base model.

[Paper](https://arxiv.org/abs/2608.31046) · [Code](https://github.com/DripNowhy/On-Policy-Self-Adaptation) · [Collection](https://huggingface.co/collections/Tuwhy/on-policy-self-adaptation-6a62d0f36f1e42afa27c7215)

## Usage

Install `torch`, `accelerate`, and `transformers>=4.57.1`.

Use the included original OLMo Instruct chat template.

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Tuwhy/Olmo-3-7B-Instruct-OPSA-Code"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
messages = [{"role": "user", "content": "Write a Python function that checks whether a string is a palindrome."}]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)
outputs = model.generate(
    **inputs,
    max_new_tokens=2048,
    do_sample=True,
    temperature=0.7,
    top_p=0.8,
    top_k=20,
)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```