dipta007 commited on
Commit
2856479
·
verified ·
1 Parent(s): c7f9f02

Fix ablation weighted averages, stored label values, reasoning-model drops, byte counts; document non-commercial training data

Browse files
Files changed (1) hide show
  1. README.md +9 -0
README.md CHANGED
@@ -184,6 +184,15 @@ For 4B models, always use SFT initialization before GRPO:
184
  | [dagger-4B_SFT](https://huggingface.co/dipta007/dagger-4B_SFT) | SFT | 25.1 |
185
  | [dagger-4B_SFT_GRPO](https://huggingface.co/dipta007/dagger-4B_SFT_GRPO) | SFT → GRPO | **31.4** |
186
 
 
 
 
 
 
 
 
 
 
187
  ## Citation
188
 
189
  ```bibtex
 
184
  | [dagger-4B_SFT](https://huggingface.co/dipta007/dagger-4B_SFT) | SFT | 25.1 |
185
  | [dagger-4B_SFT_GRPO](https://huggingface.co/dipta007/dagger-4B_SFT_GRPO) | SFT → GRPO | **31.4** |
186
 
187
+ ## License and Data Provenance
188
+
189
+ Model weights are released under the [Gemma Terms of Use](https://ai.google.dev/gemma/terms).
190
+
191
+ **Training data is not fully permissive.** Part of the SFT data and all GRPO prompts come
192
+ from `numina-math-cot-bn`, which is **CC BY-NC-SA 4.0 (NonCommercial, ShareAlike)**. For
193
+ commercial use, re-derive that portion from the Apache-2.0 upstream
194
+ [AI-MO/NuminaMath-CoT](https://huggingface.co/datasets/AI-MO/NuminaMath-CoT).
195
+
196
  ## Citation
197
 
198
  ```bibtex