Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| assets | 2 items | ||
| data | 20 items | ||
| .gitattributes | 3.53 kB xet | 9ba5c62a | |
| LICENSE | 11.3 kB xet | e2469559 | |
| README.md | 14.7 kB xet | 3092ade6 | |
| README_ZH.md | 13.1 kB xet | a4efecde |
UltraData-RL-2609
π¦ UltraData Collection | π UltraData | π€ MiniCPM5 Series
English | δΈζ
π Introduction
UltraData-RL-2609 is the L3 refined data for reinforcement learning within UltraData's L0-L4 tiered data management framework. Built for the RL stage of MiniCPM5-2B post-training, it complements UltraData-SFT-2605 with verifiable-reward tasks. It is also the training corpus used by JustRL II (Scaling Small LLMs to 128K Reasoning with a Critic), which extends the minimalist JustRL recipe with a critic and 128K-scale reasoning RL.
The release contains more than 85,000 training samples spanning mathematical reasoning, scientific / knowledge reasoning, long-context understanding, and code generation. Each sample is a verifiable task, normalized into JSONL for RL training, and constructed around three goals β verifiable outcomes, trustworthy rewards, and calibrated difficulty β so that the policy receives a stable, traceable training signal.
π’ What's New
- [2026.09.07] The UltraData-RL-2609 dataset is released! Verifiable-reward RL data for the post-training of MiniCPM5-2B, about 86K samples across Math, Knowledge (STEM), Long-Context, and Code. πππ
- [2026.09.07] MiniCPM5-2B is released!, the second model in the MiniCPM5 series after MiniCPM5-1B. It is a dense 2B Transformer that scales up the same training recipe, built for on-device, local deployment, and resource-constrained scenarios. It reaches 2B-class open-source SOTA, remains competitive with 4B-class models, and shows particular advantages in coding, mathematics, long-context understanding, tool use, and agentic tasks. UltraData-RL-2609 serves as the core RL dataset for MiniCPM5-2B. πππ
- [2026.02.08] The UltraData platform is now live, introducing the L0-L4 tiered data management framework. πππ
π― Dataset Statistics and Capability Coverage
The release contains 85,995 samples across four RL directions. Knowledge corresponds to the STEM / science-reasoning slice in the construction write-up.
| Direction | Samples | Share | Task | How outcomes are verified |
|---|---|---|---|---|
| Math | 32,412 | 37.7% | Competition- and textbook-style problems with a single extractable answer | Answer match against ground_truth |
| Code | 23,665 | 27.5% | Program synthesis from a natural-language specification | Execute the submission against test cases in ground_truth |
| Long-Context | 18,046 | 21.0% | Long-document multi-hop QA (context is included in query) |
Answer match against ground_truth; context is guaranteed to support the answer |
| Knowledge | 11,872 | 13.8% | Short-answer science / knowledge reasoning | Answer match against ground_truth |
| Total | 85,995 | 100% |
π Dataset Characteristics
- Complete domain coverage. Math, Knowledge (STEM), Long-Context, and Code target mathematical reasoning, scientific reasoning, long-context understanding, and code generation, giving a complementary distribution of RL signals.
- Verifiable outcomes. Every domain keeps an explicit reference target and a defined way to judge the result: Math and Knowledge use extractable, checkable answers; Long-Context guarantees that the context supports the answer; Code judges outputs by program execution against test cases. Multiple-choice, true/false, proof, multi-part, and image-dependent items were removed from the construction pipeline.
- Reliable reward signal. Reference answers, gold labels, and execution verdicts passed consistency checks (LLM judge, multi-model consensus, test-case cross-validation). Items whose label could not be confirmed were dropped, not guessed.
- Difficulty matched to RL training. Difficulty is controlled around the RL initialization model: items that are already fully mastered (pass rate 1) are removed; learnable items are kept; hard-but-valid items (pass rate 0 with a confirmed label) are retained and scheduled by online dynamic sampling. Difficulty filtering never modifies a label.
π§ͺ Usage Example: JustRL II
JustRL II uses UltraData-RL-2609 as its verifiable-reward training set for scaling small LLMs to 128K reasoning with a critic. Building on the minimalist JustRL recipe, it further adopts length-adaptive advantage estimation and a critic for more stable long-horizon RL.
In the JustRL II experimental setting with this dataset, AIME 2025 rises from 61 to 81 within about 300 RL steps. The final MiniCPM5-2B checkpoint that consumes UltraData-RL-2609 in post-training further reaches 86 on AIME 2025. See the JustRL II technical blog for the full ablations, and additional benchmarks.
ποΈ Data Construction Pipeline
Stages 1β2 prepare the data; stages 3β5 enforce verifiable β trustworthy β calibrated; stage 6 packages the release. All four domains share the same six stages with domain-specific verifiers and filters.
- Data collection & integration. Math: union of DAPO, DeepScaler, and DeepMath. Knowledge / STEM: OpenScienceReasoning-2. Long-Context: HotpotQA, Qasper, and MuSiQue extended with longer contexts, plus in-house synthetic data. Code: OpenCodeReasoning, OpenCodeReasoning-2, and HardTests with synthesized test cases.
- Task standardization & rewriting. Unify task format, answer field, and metadata for the trainer and verifiers. Math / Knowledge are rewritten into short-answer form; Long-Context is extended while preserving the questionβanswer correspondence; Code queries, I/O conventions, and submission format are unified.
- Verifiability filtering. Keep only items with a unique, automatically checkable answer and a unified extraction format. Long-Context items must show an explicit contextβquestionβanswer link; Code items must run in a sandbox against their test cases.
- Reward reliability check. An LLM judge audits questionβanswer consistency. Math / Knowledge answers are re-solved by several independent models and relabeled by consensus (no consensus β dropped). Long-Context answers are checked by reference comparison plus multi-model quality review. Code test cases, reference solution, and expected outputs are cross-validated; invalid or non-executable tests are removed.
- Difficulty calibration. Repeated rollouts on the RL initialization checkpoint estimate an empirical pass rate per item. Pass rate 1 (already mastered, no gradient) is removed; the learnable band is kept; pass rate 0 with a confirmed-valid label is kept and scheduled with online dynamic sampling. This stage changes difficulty and sampling weight only β never the reference label.
- Formatting & quality review. Export to unified JSONL; check parsability, required fields, unique IDs, extraction rules, and verifier configs; remove exact / near duplicates and items overlapping public evaluation benchmarks; re-check unique answers (Math / Knowledge), context linkage (Long-Context), and test-case executability (Code); record source, filter version, and verifier version.
What each domain does at stages 2β4:
| Domain | Stage 2 Β· standardization | Stage 3 Β· verifiability | Stage 4 Β· reward check |
|---|---|---|---|
| Math / Knowledge | Rewritten as short-answer | Answer extraction and match | Multi-model consensus |
| Long-Context | Context extended, QβA link kept | Explicit contextβanswer link required | Answer compare + model review |
| Code | Unified query and I/O contract | Sandboxed execution against tests | Test / solution / output cross-check |
π¦ Data Format
Each JSONL line is one RL training sample. The release uses five fields: uuid, query, ground_truth, source, and domain.
{
"uuid": "{Domain}_00001",
"query": "β¦the full problem statement shown to the policyβ¦",
"ground_truth": "β¦reference answerβ¦",
"source": "{Original Source} | UltraData-RL-2609",
"domain": "{Domain}"
}
| Field | Type | Description |
|---|---|---|
| uuid | string | Globally unique sample ID (domain prefix + index). |
| query | string | The task the model must answer or complete; for long-context tasks the full context and the question are both in this field. |
| ground_truth | string | object | Reference answer or verifiable target: a string for Math / Knowledge / Long-Context; a test-case object for Code (see below). |
| source | string | UltraData-RL-2609 |
| domain | string | One of Math, Knowledge, Long_Context, Code. |
Code problems are standard-I/O tasks (call_type = std, fn_name = null). Their ground_truth holds two equal-length arrays, inputs and outputs: the i-th entries are the complete stdin and expected stdout of the i-th test case. Verification is execution-based: pipe inputs[i] to the candidate program's stdin and compare stdout with outputs[i] after trailing-whitespace normalization; a submission is correct only if every test case passes. Example with two tests:
{"inputs": ["3\n1 2 3\n", "1\n5\n"], "outputs": ["6\n", "5\n"]}
π Quick Start
from datasets import load_dataset
ds = load_dataset("openbmb/UltraData-RL-2609", "Math", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Knowledge", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Long-Context", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Code", split="train")
print(ds[0]["query"][:300])
print(ds[0]["ground_truth"])
Available configs: Math, Knowledge, Long-Context, Code.
π‘ Intended Uses
- RL / RLVR post-training for MiniCPM-style on-device models.
- Domain slices: mathematical reasoning (
Math), science / knowledge reasoning (Knowledge), long-context multi-hop QA (Long-Context), program synthesis with executable tests (Code). - Mix-ratio studies of RL data versus SFT resources such as UltraData-SFT-2605.
β οΈ Notes and Limitations
- Verifier required for Code. The release includes test cases, not a sandbox. Users must run submissions themselves to compute rewards.
- Static labels. Difficulty filtering and online sampling weights used at construction time are not stored as fields; only the kept items are released.
- Decontamination scope. Screening covered evaluation sets known at construction time; run a new check before introducing a new benchmark.
π Data Sources
- Math: DAPO-Math-17k, DeepScaleR-Preview-Dataset, DeepMath-103K β public verifiable math RL datasets; the union of the three is used.
- Knowledge / STEM: OpenScienceReasoning-2 β scientific reasoning problems, rewritten into short-answer form.
- Long-Context: HotpotQA, Qasper, MuSiQue β multi-hop and document QA extended to longer contexts, plus in-house synthetic data.
- Code: OpenCodeReasoning, OpenCodeReasoning-2, HardTests β competitive programming problems paired with synthesized test cases.
π License and Data Sources
This project is released under the Apache 2.0 license. Upstream datasets are licensed under MIT, CC BY 4.0, and CC BY-SA 4.0, which continue to apply to content derived from them. Apache 2.0 does not override those terms.
No unauthorized unchanged redistribution: Without prior written permission from the original authors (or this organization), any institution, organization, or third-party platform is strictly prohibited from directly reposting, mirroring, re-hosting, or commercially repackaging and republishing any artifacts of this project in any form.
π Citation
If you find UltraData-RL-2609 useful in your research, please consider citing:
@misc{ultradata_rl_2609,
title = {UltraData-RL-2609},
author = {MiniCPM Team},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/openbmb/UltraData-RL-2609}}
}
- Total size
- 188 GB
- Files
- 26
- Last updated
- Sep 11
- Pre-warmed CDN
- US EU US EU