Download all_scripts/_card_data.md from SeanWang0027/rose-opd-implementation: direct link, hf CLI and curl.
- Browser
- Download file 2.1 kB
-
https://huggingface.co/SeanWang0027/rose-opd-implementation/resolve/main/all_scripts/_card_data.md
- Command line
-
hf download hf://SeanWang0027/rose-opd-implementation/all_scripts/_card_data.md
-
curl -L -o _card_data.md https://huggingface.co/SeanWang0027/rose-opd-implementation/resolve/main/all_scripts/_card_data.md
license: apache-2.0
tags:
- evaluation
- rollouts
- rlve
- aime
- hmmt
- sciworld
data-dq — 评测 rollout 与训练数据归档
evalout/ — 逐条评测记录
281 个 8 月的目录, 每条样本的完整生成文本都在。这是做 per-problem 配对分析的唯一来源 —— Olmo-7B ROSE 那条的诊断(变好 8 题 / 变差 7 题 / 题7 从 16/16 掉到 11/16 / pass@16 涨的那 1 题只靠 1/16 个样本撑着)全靠它。
格式(见 aime_eval.py):
{"id": "0", "problem": "...", "gt": 70, "hits": 16,
"samples": [{"pred": 70, "ok": true, "finish": "stop", "len": 3832, "text": "..."}]}
⚠ 部分早期 AIME 文件的逐条记录被 summary 覆盖了: --out 不带 .jsonl 后缀时,
aime_eval.py:211 的 a.out.replace(".jsonl","_summary.json") 是空操作, 两者写同一路径。
受影响的是 4B 那两条(aime_rose_dapo_4b32b_s30、aime_qwen3_4b_base_think);
Olmo / Qwen3-32B / 全部 HMMT 的都完好(各约 25MB)。
sciworld_sft/ — 32B 教师采的 SciWorld 轨迹
messages + loss_mask 格式(只有 assistant turn 是 1), 按 score 分档:
| 文件 | 轨迹 |
|---|---|
sciworld_sft_teacher32b_all.jsonl |
3592 |
..._s050.jsonl (score>=0.5) |
2050 |
..._s080.jsonl (score>=0.8) |
913 |
..._s100.jsonl (score=1.0) |
786 |
另有一批只成功 158/3592 的残缺采集(48 并发下 fd 上限 1024 耗尽), 见 sciworld-teacher32b-sft-partial。
datasets/
AIME25(30 题)、HMMT25(MathArena/hmmt_feb_2025 转成同格式, 30 题)、
DAPO-Math-17k english、RLVE train/test。
HMMT25 的答案 30 题里只有 14 题是纯整数, 其余是 LaTeX 表达式
(\frac{1}{576}、\frac{9\sqrt{23}}{23}、1-\frac{2}{\pi}), 按 int() 比会把
16 题全判错、上限只有 14/30 —— 所以 hmmt_eval.py 改用 math_verify 做符号等价判定
(自检: 标准答案 30/30 判对, 错答案 0/30 误判, \frac{2}{4}≡\frac{1}{2} 认)。