LM-Kit Doc Parser 1
lmkit-doc-parser-1.lmk is the single model behind DocumentParser in LM-Kit.NET. One file bundles everything the parser reads a document with, so there is one model to download, ship, verify and load.
It turns PDFs, scans and images into pages of positioned, typed elements (headings, text, tables, formulas, figures, charts with their data) in reading order, and renders them as Markdown, JSON, HTML or DocLang. Everything runs locally, on your hardware.
State-of-the-art on ParseBench. At ParsingEffort.High, LM-Kit Doc Parser 1 reaches an overall ParseBench score of 0.8499, the best result on the benchmark.
This model serves the document parser only. It is not a model to chat with, to generate text with, or to prompt with images directly. Its catalog capability is
DocumentParsingand nothing else.
What is inside
| Component | Role in a parse | Origin | License |
|---|---|---|---|
| Infinity-Parser2 Flash (2B, Q4_K_M, with vision projector) | Reads every page. This is what the file loads as. | infly/Infinity-Parser2-Flash | Apache-2.0 |
| PaddleOCR-VL 1.6 (0.9B, Q4_K_M, with vision projector) | Second reading of figures, charts, text blocks and tables when the first is in doubt. | PaddlePaddle/PaddleOCR-VL-1.6 | Apache-2.0 |
| Docling Layout Heron (RT-DETRv2, 17 classes) | Locates the regions of every page: text, titles, section headers, tables, pictures, formulas, captions, lists, page furniture. | docling-project/docling-layout-heron | Apache-2.0 |
The components are quantized and packed by LM-Kit; their weights are otherwise those of the upstream projects. All credit for the underlying models belongs to their authors. This bundle adds the packaging and the parsing pipeline of LM-Kit.NET around them.
Results
ParseBench scores a parser on five dimensions: tables, charts, content faithfulness (CF), semantic formatting (SF) and visual grounding (VG). Scores run from 0 to 1, higher is better. The same file serves every effort level, so you pick the balance of fidelity and speed per workload.
| Effort | Overall | Tables | Charts | CF | SF | VG | s/page | Pages/min |
|---|---|---|---|---|---|---|---|---|
| High | 0.8499 | 0.9013 | 0.7542 | 0.8979 | 0.8234 | 0.8727 | 3.05 | 19.7 |
| Medium | 0.8404 | 0.8875 | 0.7576 | 0.8891 | 0.8073 | 0.8604 | 2.05 | 29.3 |
| Low | 0.7929 | 0.8332 | 0.7138 | 0.8456 | 0.7632 | 0.8084 | 1.38 | 43.6 |
Medium keeps 99% of the High score in 0.68x of the time, and Low runs at 0.46x of the time of High for the highest throughput.
Use it
Requires LM-Kit.NET. With no argument, the parser downloads this model on first use into the LM-Kit model storage directory and loads it.
using LMKit.Document.Parsing;
using var parser = new DocumentParser(ParsingEffort.Medium);
ParsedDocument document = parser.Parse("report.pdf");
Console.WriteLine(document.ToMarkdown());
Ship the file with your application instead of downloading it:
using var parser = new DocumentParser(ParsingEffort.Medium, @"models\lmkit-doc-parser-1.lmk");
Share the model with other components by passing it loaded. It stays yours; dispose it after the parser:
using LMKit.Model;
using LM model = LM.LoadFromModelID("lmkit-doc-parser:1");
using var parser = new DocumentParser(ParsingEffort.High, model);
ParsingEffort is the one setting: Low, Medium or High trade time for fidelity, and every other decision belongs to the engine. All three levels run on this same file.
Any other model is refused with an InvalidModelException: the constructors accept this model and no other.
To prepare a machine that parses offline:
DocumentParser.GetModelCards(ParsingEffort.High).Single().Download();
Format
| Catalog ID | lmkit-doc-parser:1 |
| File | lmkit-doc-parser-1.lmk |
| Container | LM-Kit .lmk archive of type document-parser, every entry stored uncompressed |
| Size | 2,308,538,941 bytes (2.15 GiB) |
| SHA-256 | 76a4a8582314a2ba061834abbe1006a6e2a8d1004cb9e9dcbbea5192d2156791 |
| Quantization | 4-bit (Q4_K_M) readers; full-precision layout detector |
| Capability | DocumentParsing |
The page reader sits at the top level of the archive, so the file loads as an ordinary LM. The second reader streams from its bytes inside the same file. Neither is extracted to disk or copied into memory. The layout detector is a 168 MB component that its engine reads from a path, so it is unpacked once beside the other stored models.
The name carries a version. A later bundle with different readers or a different layout model ships as lmkit-doc-parser-2.lmk, and -1 keeps working for applications that pin it.
Requirements
- LM-Kit.NET (the release that introduces
DocumentParsing). - Windows x64, Linux x64 or ARM64, or macOS.
- GPU acceleration through CUDA, Vulkan or Metal, or CPU.
License
Apache-2.0, as each component. See the upstream repositories for their notices.
Model tree for lm-kit/lmkit-doc-parser-1.lmk
Base model
PaddlePaddle/PaddleOCR-VL-1.6