Jev-style decision models: open data
Data for Jev-style decision models. Notes say where labels come from: humans, exact rules or an LLM teacher. Last updated 2026-09-24.
Viewer • Updated • 166k • 9.25k • 6Note EVAL · human labels. 22 public classification/NLI/rating sets recast as choice/noul/score. Some configs (e.g. chaosnli, 100 annotators per item) carry human label distributions, so you can test calibration, not only accuracy.
LocalLLaMA/typed-decisions
Viewer • Updated • 3.2k • 27.5k • 108Note EVAL · synthetic, soft labels. The most-used shared benchmark: one state, several typed questions answered together, gold as a probability distribution. 4 workflows, 1.2k train / 400 test. PT/ES translation: telepatia-ai/typed-decisions-pt-es.
tasksource/procedural-typed-decisions
Viewer • Updated • 572k • 1.27k • 3Note EVAL/TRAIN · exact rule labels. Procedurally generated states where every answer is computed from rules stated in the state, so labels have no noise. Good for checking reasoning over state, not guessing.
AndeyTait/JevForge-Mind2Web
Viewer • Updated • 7.01k • 327Note EVAL · human gold (Mind2Web). Web-agent decisions: pick the right page element, plus a yes/no per candidate. Test and OOD splits hold out whole websites.
tasksource/tasksource-jev-typed-decisions
Viewer • Updated • 4.92M • 3.6k • 15Note TRAIN · human labels. 500k rows from tasksource classification and multiple-choice sets, recast as decisions with paraphrased questions. Flat schema: state, question, kind, options, target.
ZefanCai/Open-Jev
Viewer • Updated • 520k • 4.54k • 74Note TRAIN · controlled synthetic, rule labels. 12 frozen configs (routing, extraction, browser actions…) with calibration and OOD splits and detailed provenance. Same flat schema as tasksource-jev.
ZefanCai/Open-Jev-v1.1
Viewer • Updated • 327k • 1.44k • 18Note TRAIN · mixture. The Open-Jev 27B training mix: adds 101k WANLI NLI rows (human labels) and 51k new controlled rows in 8 domains.
SargeDev/jev-distill-corpus-v3
Viewer • Updated • 741k • 1.73k • 26Note TRAIN · LLM-teacher soft labels. 741k templated operational scenarios across 53 domains, for distilling small encoders. Same flat schema. Large, but states are template-generated.
IamBusy/OpenJev-Vision-Research-v0.1
Viewer • Updated • 12.8k • 229Note TRAIN/EVAL · images. The only image decision set so far: 12.8k images, 8.2k of them synthetic scenes with exact labels.
dylantom2012/open-system-one-bench
Viewer • Updated • 10k • 212 • 1Note COMPARISON · per-item predictions from Jev, Laya and others on the same 10k items. Rerun significance tests without API cost.
Luni/laya-jev-benchmark
Updated • 524 • 4Note COMPARISON · Jev vs Laya on benchmarks where Jev has published numbers, run by a third party.
emretheus/jev-rag-benchmark
Preview • Updated • 306Note COMPARISON · Jev as a RAG reranker vs a cross-encoder and no reranking (XQuAD, SciFact), with paired CIs and calibration tables.
nyu-mll/multi_nli
Viewer • Updated • 412k • 18.2k • 121Note SOURCE · human labels. Not Jev-framed. 433k premise/hypothesis pairs: support, contradict or neutral. A common base for a choice question over a passage.
google/boolq
Viewer • Updated • 12.7k • 82.1k • 110Note SOURCE · human labels. Not Jev-framed. 16k yes/no questions answered from a passage. Maps directly to a noul (yes/no) question.
mteb/banking77
Viewer • Updated • 13.1k • 32.3k • 19Note SOURCE · human labels. Not Jev-framed. 13k banking queries, 77 intents. Tests choice over many fine-grained options. Parquet copy of PolyAI/banking77 (the original has no viewer).
fancyzhx/ag_news
Viewer • Updated • 128k • 104k • 198Note SOURCE · human labels. Not Jev-framed. 128k news items in 4 topics. A simple choice question for mixes.
SetFit/sst5
Viewer • Updated • 11.9k • 30k • 21Note SOURCE · human labels. Not Jev-framed. 11k movie-review sentences on a 5-level sentiment scale. Useful for score or ordinal choice questions.