GRPO Tax: Evaluation Data
Complete dense-checkpoint evaluation data from the paper:
The GRPO Tax is Smaller Than You Think: A Longitudinal Study of Capability Preservation During Reasoning Training Muhammad Usama. Transactions on Machine Learning Research, 2026. OpenReview | Code
Contents
432 JSON evaluation files holding 5,616 individual benchmark scores, from dense-checkpoint evaluation of 5 models during GRPO training and 2 models during DPO training.
qwen-1.5b/ # 76 files (base + 75 checkpoints)
qwen-3b/ # 76 files
phi-3.8b/ # 76 files
gemma-2b/ # 76 files
llama-3b/ # 76 files
qwen-1.5b-dpo/ # 26 files (base + 25 checkpoints)
qwen-3b-dpo/ # 26 files
Download
This is a model-type repository, so no --repo-type flag is needed:
pip install huggingface_hub
hf download usama10/grpo-tax-eval-data --local-dir results/
To regenerate every figure in the paper from this data:
git clone https://github.com/Usama1002/GRPO-Capability-Tax.git
cd GRPO-Capability-Tax
hf download usama10/grpo-tax-eval-data --local-dir results/
python analysis/generate_all_figures.py
File Format
Each file holds scores on 13 benchmarks for a single checkpoint:
{
"base_model": "Qwen/Qwen2.5-1.5B-Instruct",
"checkpoint": "results/qwen-1.5b/training/checkpoint-3736",
"results": {
"math_reasoning": {"benchmark": "math_reasoning", "metric": "accuracy", "score": 0.355, "num_examples": 200}
}
}
Benchmarks
| Benchmark | Metric | Examples |
|---|---|---|
| GSM8K (target) | Accuracy | 200 |
| MMLU | Accuracy | 200 |
| HellaSwag | Accuracy | 200 |
| ARC-Challenge | Accuracy | 200 |
| TruthfulQA | Accuracy | 200 |
| Winogrande | Accuracy | 200 |
| IFEval | Constraint satisfaction | 30 |
| XSum | ROUGE-L | 66 |
| WMT (en->de) | BLEU | 97 |
| Coding | Syntax + structure | 25 |
| Safety | Refusal rate | 30 |
| Creative Writing | Vocabulary diversity | 25 |
| Conversation | Helpfulness | 25 |
The five heuristic benchmarks (IFEval, Coding, Safety, Creative Writing, Conversation) use rule-based scoring and are approximate; the paper reports expanded-sample-size audits for them.
Related Resources
| Resource | Link |
|---|---|
| Paper | The GRPO Tax is Smaller Than You Think: A Longitudinal Study of Capability Preservation During Reasoning Training (TMLR, 2026) |
| Source code | github.com/Usama1002/GRPO-Capability-Tax |
| Evaluation data | usama10/grpo-tax-eval-data |
| GRPO adapters | qwen-1.5b, qwen-3b, phi-3.8b, gemma-2b, llama-3b |
| DPO adapters | qwen-1.5b-dpo, qwen-3b-dpo |
Citation
@article{usama2026grpotax,
title = {The {GRPO} Tax is Smaller Than You Think: A Longitudinal Study of Capability Preservation During Reasoning Training},
author = {Muhammad Usama},
journal = {Transactions on Machine Learning Research},
issn = {2835-8856},
year = {2026},
url = {https://openreview.net/forum?id=e0UVcimXdK}
}