Add link to paper

#1
by nielsr HF Staff - opened
Files changed (1) hide show
  1. README.md +4 -2
README.md CHANGED
@@ -1,9 +1,9 @@
1
  ---
2
- license: apache-2.0
3
  base_model: Qwen/Qwen3-4B-Instruct-2507
4
  datasets:
5
  - XingYing-stack/TIPS-Training-Data
6
  library_name: transformers
 
7
  pipeline_tag: text-generation
8
  tags:
9
  - reward-model
@@ -16,8 +16,10 @@ tags:
16
 
17
  This TIPS math checkpoint is initialized from `Qwen/Qwen3-4B-Instruct-2507` and trained with outcome-only GRPO. It is a generative reward model that reasons over a mathematical solution before producing step-level and outcome labels.
18
 
 
 
19
  Training data and code are available at https://huggingface.co/datasets/XingYing-stack/TIPS-Training-Data and https://github.com/RUCBM/TIPS.
20
 
21
  Use the prompt templates and evaluation scripts in the TIPS repository. This checkpoint is intended for reward modeling and process verification rather than general-purpose chat.
22
 
23
- Built upon [verl](https://github.com/volcengine/verl) and released under Apache-2.0.
 
1
  ---
 
2
  base_model: Qwen/Qwen3-4B-Instruct-2507
3
  datasets:
4
  - XingYing-stack/TIPS-Training-Data
5
  library_name: transformers
6
+ license: apache-2.0
7
  pipeline_tag: text-generation
8
  tags:
9
  - reward-model
 
16
 
17
  This TIPS math checkpoint is initialized from `Qwen/Qwen3-4B-Instruct-2507` and trained with outcome-only GRPO. It is a generative reward model that reasons over a mathematical solution before producing step-level and outcome labels.
18
 
19
+ This model is presented in the paper [Inducing Process Supervision from Outcome-Only Reinforcement Learning](https://huggingface.co/papers/2609.36641).
20
+
21
  Training data and code are available at https://huggingface.co/datasets/XingYing-stack/TIPS-Training-Data and https://github.com/RUCBM/TIPS.
22
 
23
  Use the prompt templates and evaluation scripts in the TIPS repository. This checkpoint is intended for reward modeling and process verification rather than general-purpose chat.
24
 
25
+ Built upon [verl](https://github.com/volcengine/verl) and released under Apache-2.0.