hung-k-nguyen commited on
Commit
b18c69e
·
0 Parent(s):

Public release

Browse files
Files changed (2) hide show
  1. .gitattributes +35 -0
  2. README.md +103 -0
.gitattributes ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,103 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: nutrient-commercial
4
+ pipeline_tag: image-text-to-text
5
+ tags:
6
+ - document-ai
7
+ - grounding
8
+ - document-classification
9
+ - document-split
10
+ - open-vocabulary
11
+ ---
12
+
13
+ # nutrient-document-decision · _commercial_
14
+
15
+ **A decision model for documents.** In the spirit of general decision-making systems like Jev,
16
+ `nutrient-document-decision` reads a document once and answers typed, calibrated questions about
17
+ it — yes/no, pick-one, or scored — rather than generating free text. It's built with a particular
18
+ focus on document and multimodal understanding: grounding (is a claim actually supported by a
19
+ source), document classification (open-vocabulary, image+OCR), and document-split (page-stream
20
+ boundary detection) are the tasks we benchmark today, not the limit of what the typed-decision
21
+ interface can express — it has also shown promising zero-shot generalization to adjacent judgment
22
+ tasks it was never trained on, such as judging which of two OCR outputs is more accurate, and
23
+ combined classification-plus-split questions asked together in one pass.
24
+
25
+ - 🎯 **Try it:** [nutrient-document-decision-demo](https://huggingface.co/spaces/nutrientdocs/nutrient-document-decision-demo)
26
+ - 🏆 **Leaderboards:** [grounding](https://huggingface.co/spaces/nutrientdocs/grounding-leaderboard) ·
27
+ [document classification](https://huggingface.co/spaces/nutrientdocs/document-classification-leaderboard) ·
28
+ [open-vocabulary classification](https://huggingface.co/spaces/nutrientdocs/doc-openvocab-leaderboard) ·
29
+ [document split](https://huggingface.co/spaces/nutrientdocs/doc-split-leaderboard)
30
+
31
+ ## Results
32
+
33
+ Every task is scored against: this model, the existing Nutrient specialist purpose-built and
34
+ separately trained for that one task, and a general-purpose vision-language model run zero-shot
35
+ (no fine-tuning, no task-specific calibration) — the floor any document-AI product has to clear.
36
+
37
+ <!-- RESULTS-TABLE:START -->
38
+ ### Grounding
39
+
40
+ | | **nutrient-document-decision** | specialist ([grounding-en](https://huggingface.co/nutrientdocs/grounding-en) / [grounding-multilingual](https://huggingface.co/nutrientdocs/grounding-multilingual)) | general-purpose VLM (zero-shot) |
41
+ |---|---|---|---|
42
+ | en AUC | **0.9231** | 0.8815 | 0.8306 |
43
+ | multi AUC | 0.9485 | **0.9652** | 0.8748 |
44
+
45
+ ### Document classification (image + OCR)
46
+
47
+ | Track | **nutrient-document-decision** | specialist ([document-classification-v2](https://huggingface.co/nutrientdocs/document-classification-v2)) | general-purpose VLM (zero-shot) |
48
+ |---|---|---|---|
49
+ | doclaynet macroF1 | 0.914 | **0.968** | 0.814 |
50
+ | forms macroF1 | 0.997 | **1.000** | **1.000** |
51
+ | OOD macroF1 (unseen doc types) | 0.632 | **0.946** | 0.795 |
52
+ | OOV macroF1 (unseen label wording) | 0.812 | **0.830** | 0.810 |
53
+ | Tobacco3482 macroF1 | **0.940** | 0.735 | 0.888 |
54
+
55
+ ### Open-vocabulary / figure classification
56
+
57
+ | Track | **nutrient-document-decision** | specialist ([doc-img-classification](https://huggingface.co/nutrientdocs/doc-img-classification)) | general-purpose VLM (zero-shot) |
58
+ |---|---|---|---|
59
+ | broad (top1, ~48-label taxonomy) | **0.911** | 0.880 | 0.744 |
60
+ | specialized (in-domain) top1 | **0.827** | 0.713 | 0.492 |
61
+ | synonym (label-wording robustness) top1 | **0.833** | 0.730 | 0.740 |
62
+
63
+ ### Document split (page-stream segmentation)
64
+
65
+ | Stream | **nutrient-document-decision** | specialist ([doc-split-v2](https://huggingface.co/nutrientdocs/doc-split-v2)) | general-purpose VLM (zero-shot) |
66
+ |---|---|---|---|
67
+ | our200 F1 | **0.962** | 0.944 | 0.792 |
68
+ | OpenPSS-short F1 | **0.683** | 0.652 | 0.369 |
69
+ | OpenPSS-long F1 | **0.959** | 0.891 | 0.482 |
70
+ <!-- RESULTS-TABLE:END -->
71
+
72
+ This model mostly beats or ties the specialist across grounding, document-split, and figure
73
+ classification, and trails on a few document-classification tracks — notably OOD (unseen document
74
+ types) and multilingual grounding, where a purpose-built specialist still has an edge. See the
75
+ leaderboards linked above for the full, independently-reproducible picture against other systems.
76
+
77
+ ## Intended use & limits
78
+
79
+ Built as a decision model for document-heavy workflows (contracts, filings, forms, mixed-format
80
+ page streams) where a single calibrated pass needs to answer several different typed questions
81
+ about the same document — not limited to the four tasks benchmarked above. Not a general-purpose
82
+ chat or reasoning assistant — it reads out typed judgments (yes/no, label choice, numeric score)
83
+ rather than generating free text, and its judgment on domains very different from its benchmark
84
+ coverage (e.g. truly novel, unseen document types) should be validated before relying on it
85
+ unsupervised.
86
+
87
+ ## License & data
88
+
89
+ Commercial. Benchmarked against publicly available document-AI evaluation sets (see the linked
90
+ leaderboards and benchmark datasets for exact sources and licenses per track).
91
+
92
+ > ### 📩 Get access
93
+ >
94
+ > `nutrient-document-decision` is commercial and its weights are not downloadable here. To run it
95
+ > on-prem — **contact Nutrient: [nutrient.io/contact-sales](https://www.nutrient.io/contact-sales/).**
96
+
97
+ ## About the author
98
+
99
+ <a href="https://nutrient.io/">
100
+ <img src="https://avatars2.githubusercontent.com/u/1527679?v=3&s=200" height="80" />
101
+ </a>
102
+
103
+ This project is maintained and funded by [Nutrient](https://nutrient.io/) - The deterministic document infrastructure enterprises run their highest-stakes workflows on: replayable output, clear exceptions, and full audit trails on the messy, regulated documents where AI alone breaks.