nutrient-document-decision · commercial
A decision model for documents. In the spirit of general decision-making systems like Jev,
nutrient-document-decision reads a document once and answers typed, calibrated questions about
it — yes/no, pick-one, or scored — rather than generating free text. It's built with a particular
focus on document and multimodal understanding: grounding (is a claim actually supported by a
source), document classification (open-vocabulary, image+OCR), and document-split (page-stream
boundary detection) are the tasks we benchmark today, not the limit of what the typed-decision
interface can express — it has also shown promising zero-shot generalization to adjacent judgment
tasks it was never trained on, such as judging which of two OCR outputs is more accurate, and
combined classification-plus-split questions asked together in one pass.
- 🎯 Try it: nutrient-document-decision-demo
- 🏆 Leaderboards: grounding · document classification · open-vocabulary classification · document split
Results
Every task is scored against: this model, the existing Nutrient specialist purpose-built and separately trained for that one task, and a general-purpose vision-language model run zero-shot (no fine-tuning, no task-specific calibration) — the floor any document-AI product has to clear.
Grounding
| nutrient-document-decision | specialist (grounding-en / grounding-multilingual) | general-purpose VLM (zero-shot) | |
|---|---|---|---|
| en AUC | 0.9231 | 0.8815 | 0.8306 |
| multi AUC | 0.9485 | 0.9652 | 0.8748 |
Document classification (image + OCR)
| Track | nutrient-document-decision | specialist (document-classification-v2) | general-purpose VLM (zero-shot) |
|---|---|---|---|
| doclaynet macroF1 | 0.914 | 0.968 | 0.814 |
| forms macroF1 | 0.997 | 1.000 | 1.000 |
| OOD macroF1 (unseen doc types) | 0.632 | 0.946 | 0.795 |
| OOV macroF1 (unseen label wording) | 0.812 | 0.830 | 0.810 |
| Tobacco3482 macroF1 | 0.940 | 0.735 | 0.888 |
Open-vocabulary / figure classification
| Track | nutrient-document-decision | specialist (doc-img-classification) | general-purpose VLM (zero-shot) |
|---|---|---|---|
| broad (top1, ~48-label taxonomy) | 0.911 | 0.880 | 0.744 |
| specialized (in-domain) top1 | 0.827 | 0.713 | 0.492 |
| synonym (label-wording robustness) top1 | 0.833 | 0.730 | 0.740 |
Document split (page-stream segmentation)
| Stream | nutrient-document-decision | specialist (doc-split-v2) | general-purpose VLM (zero-shot) |
|---|---|---|---|
| our200 F1 | 0.962 | 0.944 | 0.792 |
| OpenPSS-short F1 | 0.683 | 0.652 | 0.369 |
| OpenPSS-long F1 | 0.959 | 0.891 | 0.482 |
This model mostly beats or ties the specialist across grounding, document-split, and figure classification, and trails on a few document-classification tracks — notably OOD (unseen document types) and multilingual grounding, where a purpose-built specialist still has an edge. See the leaderboards linked above for the full, independently-reproducible picture against other systems.
Intended use & limits
Built as a decision model for document-heavy workflows (contracts, filings, forms, mixed-format page streams) where a single calibrated pass needs to answer several different typed questions about the same document — not limited to the four tasks benchmarked above. Not a general-purpose chat or reasoning assistant — it reads out typed judgments (yes/no, label choice, numeric score) rather than generating free text, and its judgment on domains very different from its benchmark coverage (e.g. truly novel, unseen document types) should be validated before relying on it unsupervised.
License & data
Commercial. Benchmarked against publicly available document-AI evaluation sets (see the linked leaderboards and benchmark datasets for exact sources and licenses per track).
📩 Get access
nutrient-document-decisionis commercial and its weights are not downloadable here. To run it on-prem — contact Nutrient: nutrient.io/contact-sales.
About the author
This project is maintained and funded by Nutrient - The deterministic document infrastructure enterprises run their highest-stakes workflows on: replayable output, clear exceptions, and full audit trails on the messy, regulated documents where AI alone breaks.