File size: 1,266 Bytes
8036717
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
---
title: Doc-Split Demo
emoji: 📄✂️
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: 6.20.0
app_file: app.py
short_description: Split a merged PDF into its constituent documents
pinned: false
---

# Doc-Split Demo

Upload a PDF that concatenates several documents; the model marks where each new document begins and groups
the pages into documents. Pick **doc-split-v2** (commercial flagship) or the open **doc-split-v1**. Uses the PDF's embedded
text layer where present; scanned pages fall back to the image model. Confidence is calibrated.

Part of the doc-split hub: [leaderboard](https://huggingface.co/spaces/nutrientdocs/doc-split-leaderboard) ·
[flagship](https://huggingface.co/nutrientdocs/doc-split-v2) · [open v1](https://huggingface.co/nutrientdocs/doc-split-v1) ·
[benchmark](https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark).

## Secret required

**doc-split-v2** is commercial and loaded from a **private** repo, so this Space needs an `HF_TOKEN`
**secret** (Settings → Variables and secrets) with read access to it. The weights are used server-side only
and are never downloadable. Without the secret, the v2 model is marked _unavailable_ and the open **doc-split-v1**
still runs. Inference runs on **ZeroGPU**.