dusersad12 commited on
Commit
aef7584
·
verified ·
1 Parent(s): 317aa90

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +69 -0
README.md ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: transformers
4
+ tags:
5
+ - reward-model
6
+ - rlhf
7
+ - alignment
8
+ ---
9
+
10
+ # BestRewardModel
11
+
12
+ <div align="center">
13
+ <img src="figures/training_curve.png" width="70%" alt="Training Curve" />
14
+ </div>
15
+
16
+ ## Model Description
17
+
18
+ This is a reward model trained for RLHF alignment, selected from multiple experimental runs based on validation accuracy and reward alignment quality.
19
+
20
+ ## Selection Criteria
21
+
22
+ The best checkpoint was chosen according to:
23
+ - **Highest `val_accuracy`** among all final checkpoints
24
+ - **Minimum `reward_alignment_score` threshold of 0.80**
25
+
26
+ Only checkpoints satisfying **both** conditions were eligible.
27
+
28
+ ## Training Runs Comparison
29
+
30
+ | Run | Base Model | Learning Rate | Final Step | Val Accuracy | Reward Alignment | Train Loss |
31
+ |-----|-----------|--------------|------------|-------------|-----------------|------------|
32
+ | run_gpt2_base_lr1e4 | GPT-2 Base | 1e-4 | 1000 | 0.907 | 0.876 | 0.115 |
33
+ | run_gpt2_base_lr5e5 | GPT-2 Base | 5e-5 | 1000 | 0.870 | 0.839 | 0.207 |
34
+ | run_gpt2_large_lr1e4 | GPT-2 Large | 1e-4 | 1000 | 0.958 | 0.928 | 0.061 |
35
+ | run_gpt2_large_lr5e5 | GPT-2 Large | 5e-5 | 1000 | 0.901 | 0.854 | 0.159 |
36
+ | run_deberta_lr1e4 | DeBERTa-v2 | 1e-4 | 1000 | 0.837 | 0.827 | 0.301 |
37
+
38
+ ## Best Run Metrics
39
+
40
+ | Metric | Value |
41
+ |--------|-------|
42
+ | Run Name | run_gpt2_large_lr1e4 |
43
+ | Val Accuracy | 0.958 |
44
+ | Reward Alignment Score | 0.928 |
45
+ | Final Train Loss | 0.061 |
46
+
47
+ ## Intended Uses
48
+
49
+ This model is intended for use as a reward model in RLHF pipelines to score and rank model outputs based on human preference alignment.
50
+
51
+ ## How to Use
52
+
53
+ ```python
54
+ from transformers import AutoModelForSequenceClassification, AutoTokenizer
55
+
56
+ model = AutoModelForSequenceClassification.from_pretrained("BestRewardModel-TestRepo")
57
+ tokenizer = AutoTokenizer.from_pretrained("BestRewardModel-TestRepo")
58
+
59
+ inputs = tokenizer("prompt", "response", return_tensors="pt")
60
+ score = model(**inputs).logits[0].item()
61
+ ```
62
+
63
+ <div align="center">
64
+ <img src="figures/reward_dist.png" width="60%" alt="Reward Distribution" />
65
+ </div>
66
+
67
+ ## License
68
+
69
+ Apache-2.0