MyeongHoJeong commited on
Commit
21ce2ee
·
verified ·
1 Parent(s): 5629065

Format training dataset list as a table

Browse files
Files changed (1) hide show
  1. README.md +44 -1
README.md CHANGED
@@ -229,7 +229,50 @@ Training data is synthetic and format-augmented decision data plus decision item
229
  datasets (listed below); the JevBench public tiers used only for evaluation carry MIT. Full per-cohort breakdown (row counts, what each covers, licence): [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md#training-data-provenance).
230
 
231
  Public datasets used (train splits where the dataset has one; licence as stated by each dataset; labels come from the
232
- datasets, distractor options are generated by code): SQuAD 2.0 (CC BY-SA 4.0), ARC (CC BY-SA 4.0), BoolQ (CC BY-SA 3.0), CommonsenseQA (MIT), HellaSwag (MIT), Banking77 (CC BY 4.0), Bias in Bios (MIT), Bitext customer support (CDLA-Sharing-1.0), CLINC150 (CC BY 3.0), Amazon Counterfactual (CC BY 4.0), DBpedia-14 (CC BY-SA 3.0), Dolly 15k (CC BY-SA 3.0), GoEmotions (Apache-2.0), MASSIVE (CC BY 4.0), Twitter Financial News Sentiment (MIT), HelpSteer3 (CC BY 4.0), HelpSteer2 (CC BY 4.0), 2WikiMultihopQA (Apache-2.0), HotpotQA (CC BY-SA 4.0), MuSiQue (CC BY 4.0), QASC (CC BY 4.0), DROP (CC BY-SA 4.0), GSM8K (MIT), TempReason (CC BY-SA 3.0), MultiNLI (OANC / CC BY-SA 3.0 / CC BY 3.0), PAWS (Google terms, free for any purpose), PAWS-X (Google terms, free for any purpose), SNLI (CC BY-SA 4.0), WANLI (CC BY 4.0), ContractNLI (CC BY 4.0), CUAD (CC BY 4.0), ShARC (CC BY-SA 3.0), Jailbreak classification (Apache-2.0), Prompt injections (Apache-2.0), Aegis AI Content Safety 2.0 (CC BY 4.0), Jigsaw Toxic Comment Classification (mirror of the Kaggle data) (CC0 (data); comment text CC BY-SA 3.0 (Wikipedia)), Measuring Hate Speech (CC BY 4.0), Image safety classes (MIT). Upstream ids and the cohort each one feeds: [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md#training-data-provenance).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
233
 
234
  An exact-text overlap audit against the public JevBench tiers found 0 exact scenario matches and 181 exact instruction matches — rows in two adequacy-rubric cohorts whose entire instruction field, a generic 58-character adequacy question, is byte-identical to one public hard-tier instruction (0.03 % of the 520,754-row training mixture). These rows are kept and disclosed here rather than regenerated, since the overlap is limited to one rubric question's wording and never touches a scenario or an answer.
235
 
 
229
  datasets (listed below); the JevBench public tiers used only for evaluation carry MIT. Full per-cohort breakdown (row counts, what each covers, licence): [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md#training-data-provenance).
230
 
231
  Public datasets used (train splits where the dataset has one; licence as stated by each dataset; labels come from the
232
+ datasets, distractor options are generated by code):
233
+
234
+ | Dataset | Licence |
235
+ |---|---|
236
+ | SQuAD 2.0 | CC BY-SA 4.0 |
237
+ | ARC | CC BY-SA 4.0 |
238
+ | BoolQ | CC BY-SA 3.0 |
239
+ | CommonsenseQA | MIT |
240
+ | HellaSwag | MIT |
241
+ | Banking77 | CC BY 4.0 |
242
+ | Bias in Bios | MIT |
243
+ | Bitext customer support | CDLA-Sharing-1.0 |
244
+ | CLINC150 | CC BY 3.0 |
245
+ | Amazon Counterfactual | CC BY 4.0 |
246
+ | DBpedia-14 | CC BY-SA 3.0 |
247
+ | Dolly 15k | CC BY-SA 3.0 |
248
+ | GoEmotions | Apache-2.0 |
249
+ | MASSIVE | CC BY 4.0 |
250
+ | Twitter Financial News Sentiment | MIT |
251
+ | HelpSteer3 | CC BY 4.0 |
252
+ | HelpSteer2 | CC BY 4.0 |
253
+ | 2WikiMultihopQA | Apache-2.0 |
254
+ | HotpotQA | CC BY-SA 4.0 |
255
+ | MuSiQue | CC BY 4.0 |
256
+ | QASC | CC BY 4.0 |
257
+ | DROP | CC BY-SA 4.0 |
258
+ | GSM8K | MIT |
259
+ | TempReason | CC BY-SA 3.0 |
260
+ | MultiNLI | OANC / CC BY-SA 3.0 / CC BY 3.0 |
261
+ | PAWS | Google terms, free for any purpose |
262
+ | PAWS-X | Google terms, free for any purpose |
263
+ | SNLI | CC BY-SA 4.0 |
264
+ | WANLI | CC BY 4.0 |
265
+ | ContractNLI | CC BY 4.0 |
266
+ | CUAD | CC BY 4.0 |
267
+ | ShARC | CC BY-SA 3.0 |
268
+ | Jailbreak classification | Apache-2.0 |
269
+ | Prompt injections | Apache-2.0 |
270
+ | Aegis AI Content Safety 2.0 | CC BY 4.0 |
271
+ | Jigsaw Toxic Comment Classification (mirror of the Kaggle data) | CC0 (data); comment text CC BY-SA 3.0 (Wikipedia) |
272
+ | Measuring Hate Speech | CC BY 4.0 |
273
+ | Image safety classes | MIT |
274
+
275
+ Upstream ids and the cohort each one feeds: [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md#training-data-provenance).
276
 
277
  An exact-text overlap audit against the public JevBench tiers found 0 exact scenario matches and 181 exact instruction matches — rows in two adequacy-rubric cohorts whose entire instruction field, a generic 58-character adequacy question, is byte-identical to one public hard-tier instruction (0.03 % of the 520,754-row training mixture). These rows are kept and disclosed here rather than regenerated, since the overlap is limited to one rubric question's wording and never touches a scenario or an answer.
278