|
Download README.md from nutrientdocs/nutrient-document-decision: direct link, hf CLI and curl.
- Browser
- Download file 5.52 kB
-
https://huggingface.co/nutrientdocs/nutrient-document-decision/resolve/main/README.md
- Command line
-
hf download hf://nutrientdocs/nutrient-document-decision/README.md
-
curl -L -o README.md https://huggingface.co/nutrientdocs/nutrient-document-decision/resolve/main/README.md
5.52 kB
| license: other | |
| license_name: nutrient-commercial | |
| pipeline_tag: image-text-to-text | |
| tags: | |
| - document-ai | |
| - grounding | |
| - document-classification | |
| - document-split | |
| - open-vocabulary | |
| # nutrient-document-decision Β· _commercial_ | |
| **A decision model for documents.** In the spirit of general decision-making systems like Jev, | |
| `nutrient-document-decision` reads a document once and answers typed, calibrated questions about | |
| it β yes/no, pick-one, or scored β rather than generating free text. It's built with a particular | |
| focus on document and multimodal understanding: grounding (is a claim actually supported by a | |
| source), document classification (open-vocabulary, image+OCR), and document-split (page-stream | |
| boundary detection) are the tasks we benchmark today, not the limit of what the typed-decision | |
| interface can express β it has also shown promising zero-shot generalization to adjacent judgment | |
| tasks it was never trained on, such as judging which of two OCR outputs is more accurate, and | |
| combined classification-plus-split questions asked together in one pass. | |
| - π― **Try it:** [nutrient-document-decision-demo](https://huggingface.co/spaces/nutrientdocs/nutrient-document-decision-demo) | |
| - π **Leaderboards:** [grounding](https://huggingface.co/spaces/nutrientdocs/grounding-leaderboard) Β· | |
| [document classification](https://huggingface.co/spaces/nutrientdocs/document-classification-leaderboard) Β· | |
| [open-vocabulary classification](https://huggingface.co/spaces/nutrientdocs/doc-openvocab-leaderboard) Β· | |
| [document split](https://huggingface.co/spaces/nutrientdocs/doc-split-leaderboard) | |
| ## Results | |
| Every task is scored against: this model, the existing Nutrient specialist purpose-built and | |
| separately trained for that one task, and a general-purpose vision-language model run zero-shot | |
| (no fine-tuning, no task-specific calibration) β the floor any document-AI product has to clear. | |
| <!-- RESULTS-TABLE:START --> | |
| ### Grounding | |
| | | **nutrient-document-decision** | specialist ([grounding-en](https://huggingface.co/nutrientdocs/grounding-en) / [grounding-multilingual](https://huggingface.co/nutrientdocs/grounding-multilingual)) | general-purpose VLM (zero-shot) | | |
| |---|---|---|---| | |
| | en AUC | **0.9231** | 0.8815 | 0.8306 | | |
| | multi AUC | 0.9485 | **0.9652** | 0.8748 | | |
| ### Document classification (image + OCR) | |
| | Track | **nutrient-document-decision** | specialist ([document-classification-v2](https://huggingface.co/nutrientdocs/document-classification-v2)) | general-purpose VLM (zero-shot) | | |
| |---|---|---|---| | |
| | doclaynet macroF1 | 0.914 | **0.968** | 0.814 | | |
| | forms macroF1 | 0.997 | **1.000** | **1.000** | | |
| | OOD macroF1 (unseen doc types) | 0.632 | **0.946** | 0.795 | | |
| | OOV macroF1 (unseen label wording) | 0.812 | **0.830** | 0.810 | | |
| | Tobacco3482 macroF1 | **0.940** | 0.735 | 0.888 | | |
| ### Open-vocabulary / figure classification | |
| | Track | **nutrient-document-decision** | specialist ([doc-img-classification](https://huggingface.co/nutrientdocs/doc-img-classification)) | general-purpose VLM (zero-shot) | | |
| |---|---|---|---| | |
| | broad (top1, ~48-label taxonomy) | **0.911** | 0.880 | 0.744 | | |
| | specialized (in-domain) top1 | **0.827** | 0.713 | 0.492 | | |
| | synonym (label-wording robustness) top1 | **0.833** | 0.730 | 0.740 | | |
| ### Document split (page-stream segmentation) | |
| | Stream | **nutrient-document-decision** | specialist ([doc-split-v2](https://huggingface.co/nutrientdocs/doc-split-v2)) | general-purpose VLM (zero-shot) | | |
| |---|---|---|---| | |
| | our200 F1 | **0.962** | 0.944 | 0.792 | | |
| | OpenPSS-short F1 | **0.683** | 0.652 | 0.369 | | |
| | OpenPSS-long F1 | **0.959** | 0.891 | 0.482 | | |
| <!-- RESULTS-TABLE:END --> | |
| This model mostly beats or ties the specialist across grounding, document-split, and figure | |
| classification, and trails on a few document-classification tracks β notably OOD (unseen document | |
| types) and multilingual grounding, where a purpose-built specialist still has an edge. See the | |
| leaderboards linked above for the full, independently-reproducible picture against other systems. | |
| ## Intended use & limits | |
| Built as a decision model for document-heavy workflows (contracts, filings, forms, mixed-format | |
| page streams) where a single calibrated pass needs to answer several different typed questions | |
| about the same document β not limited to the four tasks benchmarked above. Not a general-purpose | |
| chat or reasoning assistant β it reads out typed judgments (yes/no, label choice, numeric score) | |
| rather than generating free text, and its judgment on domains very different from its benchmark | |
| coverage (e.g. truly novel, unseen document types) should be validated before relying on it | |
| unsupervised. | |
| ## License & data | |
| Commercial. Benchmarked against publicly available document-AI evaluation sets (see the linked | |
| leaderboards and benchmark datasets for exact sources and licenses per track). | |
| > ### π© Get access | |
| > | |
| > `nutrient-document-decision` is commercial and its weights are not downloadable here. To run it | |
| > on-prem β **contact Nutrient: [nutrient.io/contact-sales](https://www.nutrient.io/contact-sales/).** | |
| ## About the author | |
| <a href="https://nutrient.io/"> | |
| <img src="https://avatars2.githubusercontent.com/u/1527679?v=3&s=200" height="80" /> | |
| </a> | |
| This project is maintained and funded by [Nutrient](https://nutrient.io/) - The deterministic document infrastructure enterprises run their highest-stakes workflows on: replayable output, clear exceptions, and full audit trails on the messy, regulated documents where AI alone breaks. | |