Schemer
Extract typed JSON from any text.
On-device structured extraction into a caller-supplied JSON schema.
- GitHub: https://github.com/Desert-Ant-Labs
- Website: https://desertant.com/models/schemer/
| Platforms | |
| Languages | 13 |
| Weights | main |
Install
Hand Schemer a schema and get JSON back that matches it. The schema is the model's input, not a prompt suggestion, and every field is decoded for its type. Schemer does not generate text.
- About 230MB per platform, fully on device
- 13 languages, 0.075 per-language accuracy spread
- Documents up to ~1,100 tokens (roughly 3 pages); short inputs run a faster path
- Offline, on the Neural Engine on Apple devices and on the CPU elsewhere. The model runs once per field, so cost scales with how many fields your schema declares, not with the document alone. On an M1 a short four-field record extracts in about 85ms, a long document in about 1.2s
Three guarantees no generative extractor offers:
- Typed by construction. Labels are always one of your declared choices, numbers arrive clamped to your range, datetimes are ISO 8601, and relative expressions ("tomorrow at 3pm", "i morgen kl 15", misspellings included) resolve against device time.
- Absence detection. Fields the text does not state come back null. Schemer scores 0.91 on this; every LLM and span extractor we benched, at any size, scores between 0.18 and 0.43.
- Verbatim. Extracted strings are substrings of the input, so hallucination auditing is a substring check.
Try it
- Live demo: desert-ant-labs/schemer-demo: paste text, describe your fields, get typed JSON, fully in your browser.
Use
The SDKs download the files for their platform from this repository at a pinned tag and run the whole pipeline on device: Swift on Apple platforms, Linux and Windows, Kotlin on Android, and JavaScript in the browser and in Node. Install lines are in the block above; the full API, options and examples are on the SDK page: https://github.com/Desert-Ant-Labs/desert-ant-core/blob/main/docs/models/schemer.md.
import Schemer
let schemer = Schemer()
let out = try await schemer.extract(from: text, schema: [
.string("title", describe: "the task title"),
.string("guest", describe: "person the meeting is with", nullable: true),
.datetime("start", describe: "start time", nullable: true),
.number("duration_min", describe: "duration in minutes", nullable: true, min: 0, max: 480),
.boolean("is_recurring", describe: "is this a recurring task"),
])
print(out.json)
In JavaScript the same schema is plain JSON:
{
"title": {"type": "string", "describe": "the task title"},
"guest": {"type": "string", "describe": "person the meeting is with", "nullable": true},
"start": {"type": "datetime","describe": "start time", "nullable": true},
"duration_min": {"type": "number", "describe": "duration in minutes", "min": 0, "max": 480, "nullable": true},
"is_recurring": {"type": "boolean", "describe": "is this a recurring task"}
}
Input "Meeting with Sarah tomorow at 3pm for an hour to review the Q3 deck. Recurring weekly." returns (with device time 2026-07-05):
{"title": "review the Q3 deck", "guest": "Sarah",
"start": "2026-07-06T15:00", "duration_min": 60, "is_recurring": true}
Note the typo tolerance, the relative-date resolution, the unit conversion
("an hour" to 60), and the factoring: the title is the purpose of the
meeting, the person lands in guest, and the temporal clutter lands in
the temporal fields.
Inputs and outputs
- Input: a text and a schema. Field types:
string,label(with up to 16values),number(with optionalmin,maxandunit),datetime,boolean,arrayof strings, andarrayof objects (flat properties). Every field takes a shortdescribe;nullable: truetells the model absence is an expected answer. The SDK passes device time as the date relative expressions resolve against; callers can pin it. - Budget: the schema and the text share one 1216-token window; longer inputs are truncated, so plan on roughly 1,100 tokens of text. Inputs of 256 tokens or fewer run on the short shape.
- Output: one value per declared field, in schema order. Strings and
array items are verbatim substrings of the input; labels are one of the
declared
values; numbers are decoded to the field's unit; datetimes are ISO 8601 (YYYY-MM-DDTHH:MM); an absent string, boolean or datetime is null and an absent array is empty. Declarenullable: trueon an optional number, or it composes a value.
Files
Every platform downloads the two shared files plus its own three model files.
| File | Format | Size | Contents |
|---|---|---|---|
schemer-encoder.mlmodelc |
Compiled Core ML (8-bit) | 114MB | Ready to load on Apple platforms (used by the Swift SDK) |
schemer-decode.mlmodelc |
Compiled Core ML (8-bit) | 20MB | Ready to load on Apple platforms (used by the Swift SDK) |
schemer-label.mlmodelc |
Compiled Core ML (8-bit) | 3MB | Ready to load on Apple platforms (used by the Swift SDK) |
schemer-encoder.tflite |
LiteRT / TFLite (int8) | 117MB | Runs on Android, Linux, Windows, Node, and the web (downloaded on demand by the Kotlin and JavaScript SDKs) |
schemer-decode.tflite |
LiteRT / TFLite (int8) | 20MB | Runs on Android, Linux, Windows, Node, and the web |
schemer-label.tflite |
LiteRT / TFLite (int8) | 3MB | Runs on Android, Linux, Windows, Node, and the web |
embeddings.q |
int8 table | 80MB | Shared by every platform |
schemer_tokenizer.bin |
Tokenizer | 11MB | Shared by every platform |
Evaluation
9,021 held-out records across seven usage slices, 13 languages, schemas the model never saw, one order-insensitive, absence-aware scorer for every model. The overall score weights the slices by product usage: short records 0.30, long documents 0.15, long documents with same-type distractors 0.15, reviews and registers 0.15, calendar 0.10, factored domains 0.10, nested schemas 0.05. Every model ran in its documented configuration (tool calling for FunctionGemma, template prompts for NuExtract-2.0, the schema in the system prompt for LFM2-Extract, JSON chat for the instruct models, in-process span extraction for GLiNER2); sizes are measured bytes of the artifact benched.
| Model | Overall | Absence (boolean) | Params | On disk |
|---|---|---|---|---|
| Gemma 4 26B A4B (int4 AWQ) | 0.834 | 0.428 | 26.6B | 17.2GB |
| Claude Haiku 4.5 (API) | 0.816 | 0.431 | n/a | API only |
| Qwen 3.5 9B | 0.812 | 0.434 | 9.1B | 9.1GB |
| Schemer | 0.800 | 0.911 | 211M | 230MB |
| Ministral 3 8B | 0.790 | 0.430 | 8.0B | 17.8GB |
| Gemma 4 E2B (official int4 QAT) | 0.780 | 0.427 | 5.1B | 8.3GB |
| Qwen 3.5 0.8B | 0.651 | 0.351 | 0.87B | 1.75GB |
| NuExtract-2.0-2B | 0.643 | 0.393 | 2.2B | 4.4GB |
| GLiNER2-multi | 0.585 | 0.374 | 307M | 309MB int8* |
| GLiNER2-base | 0.524 | 0.368 | 205M | 834MB |
| LFM2-1.2B-Extract | 0.489 | 0.325 | 1.17B | 2.3GB |
| NuExtract-tiny-v1.5 | 0.334 | 0.222 | 0.5B | 954MB |
| LFM2-350M-Extract | 0.331 | 0.276 | 354M | 1.4GB |
| FunctionGemma 270M (tool calling) | 0.288 | 0.181 | 0.27B | 0.54GB |
*GLiNER2-multi int8 size verified by us by quantizing it and scoring it again (accuracy held within 0.004 of fp32).
The files in this repository were scored themselves, not only the model they came from: the Core ML files score 0.804 overall, record for record within noise of the full-precision model, and the LiteRT files match the Core ML files.
Per slice
| Model | Short | Long doc | Long + distractors | Register | Calendar | Factored | Nested |
|---|---|---|---|---|---|---|---|
| Gemma 4 26B A4B | 0.844 | 0.844 | 0.693 | 0.815 | 0.829 | 0.954 | 1.000 |
| Claude Haiku 4.5 | 0.838 | 0.759 | 0.762 | 0.831 | 0.802 | 0.817 | 1.000 |
| Qwen 3.5 9B | 0.828 | 0.807 | 0.717 | 0.791 | 0.774 | 0.891 | 1.000 |
| Schemer | 0.759 | 0.708 | 0.685 | 0.829 | 0.946 | 0.946 | 0.988 |
| Ministral 3 8B | 0.818 | 0.811 | 0.711 | 0.720 | 0.796 | 0.790 | 0.998 |
| Gemma 4 E2B | 0.809 | 0.785 | 0.705 | 0.773 | 0.719 | 0.759 | 1.000 |
| Qwen 3.5 0.8B | 0.721 | 0.635 | 0.560 | 0.581 | 0.529 | 0.689 | 0.936 |
| NuExtract-2.0-2B | 0.763 | 0.704 | 0.542 | 0.648 | 0.496 | 0.805 | 0.000† |
| GLiNER2-multi | 0.683 | 0.650 | 0.388 | 0.662 | 0.468 | 0.783 | 0.000† |
| GLiNER2-base | 0.619 | 0.587 | 0.371 | 0.612 | 0.362 | 0.668 | 0.000† |
| LFM2-1.2B-Extract | 0.515 | 0.445 | 0.265 | 0.652 | 0.316 | 0.636 | 0.692 |
| NuExtract-tiny-v1.5 | 0.406 | 0.277 | 0.248 | 0.280 | 0.275 | 0.417 | 0.439 |
| LFM2-350M-Extract | 0.370 | 0.318 | 0.219 | 0.299 | 0.327 | 0.367 | 0.503 |
| FunctionGemma 270M | 0.339 | 0.254 | 0.232 | 0.274 | 0.299 | 0.421 | 0.000† |
†The NuExtract-2.0, GLiNER2, and FunctionGemma interfaces cannot express arrays of objects; their nested cells score the resulting empty output. This is an interface limit, not an extraction failure.
Per field type
Bold marks the best score in each row. Schemer's architecture shows plainly: it owns the typed/absence rows and trails larger extraction LLMs on free-text rows.
| Type | Schemer | NuExtract-2.0-2B (2.2B) | Qwen 3.5 0.8B |
|---|---|---|---|
| boolean | 0.911 | 0.393 | 0.351 |
| datetime | 0.769 | 0.641 | 0.652 |
| array | 0.766 | 0.731 | 0.686 |
| number | 0.750 | 0.670 | 0.679 |
| string | 0.648 | 0.774 | 0.716 |
| label | 0.629 | 0.636 | 0.616 |
Per language
Measured on the three slices that cover all 13 languages with identical composition (short records, long documents, long documents with distractors; ~600 records per language).
| Language | Schemer | Qwen 3.5 0.8B | GLiNER2-multi | LFM2-1.2B-Extract |
|---|---|---|---|---|
| French | 0.745 | 0.649 | 0.606 | 0.440 |
| Spanish | 0.734 | 0.662 | 0.610 | 0.435 |
| Italian | 0.731 | 0.653 | 0.603 | 0.452 |
| Portuguese | 0.729 | 0.640 | 0.601 | 0.449 |
| German | 0.728 | 0.635 | 0.591 | 0.418 |
| Dutch | 0.724 | 0.621 | 0.600 | 0.373 |
| Danish | 0.723 | 0.618 | 0.594 | 0.360 |
| English | 0.717 | 0.706 | 0.632 | 0.498 |
| Norwegian | 0.717 | 0.620 | 0.600 | 0.361 |
| Swedish | 0.715 | 0.633 | 0.624 | 0.364 |
| Japanese | 0.707 | 0.614 | 0.384 | 0.405 |
| Chinese | 0.683 | 0.640 | 0.408 | 0.375 |
| Polish | 0.670 | 0.609 | 0.597 | 0.370 |
Schemer scores 0.800 overall; every model above it runs 9.1B to 26.6B parameters or an API. NuExtract-2.0-2B, the closest task-specific extractor, scores 0.643 at ten times the parameters. Schemer wins all 13 languages against the strongest sub-1B LLM, GLiNER2-multi, and LFM2-1.2B-Extract, posts the top calendar score of any model at any size (0.946 next to the 26B's 0.829), and is the only model with reliable absence detection (0.911 next to a 0.18 to 0.43 field).
On long documents Schemer scores 0.708, and 0.685 when the document also holds records of the same type as the target. Both slices are built from the same kind of records as the short slice, so they measure disambiguation inside the input budget, not general document reading. On free prose that mixes many candidates of the same type (several invoices in one thread, several speakers in one transcript) the larger models keep a clear lead, and that is the slice where Schemer is weakest.
Limitations
- Lists of objects extract by segment and recurse in the SDK, not by the model: the text is split into candidate items and each is extracted on its own. The nested figure above comes from a slice built from regular records (itineraries, shopping lists, workout logs), where each item has its own clause. On free prose the common failures are item boundaries: fields mixed between items that share a sentence, or an item missed. A single nested object, and objects inside objects, are not supported.
- Extractive strings only: Schemer will not compose or rewrite text; abstractive titles from messy prose trail extraction LLMs (string 0.648 vs 0.774 for NuExtract-2.0-2B and ~0.86 for the 8B+ class). It does factor titles extractively (purpose clause, people and times split into their fields), but the title is always a substring of the input.
- Judgment labels requiring world knowledge (severity triage) trail the 8B+ class (0.629 vs ~0.82).
- Long prose with many same-type candidates is the weakest slice; see the note under Evaluation. Schemer reads up to about 1,100 tokens and does not chunk longer documents on its own.
- Labels take at most 16 values and always choose one of them.
- Absence is preferred over a wrong value. A field that is stated in the text can still come back null when the phrasing sits far from the field description. Write descriptions the way the text would say it.
- Relative dates resolve through a 13-language lexicon in the SDK; phrasing outside it falls back to the anchor date.
- Not a NER model: it fills the fields you declare; it does not enumerate every entity mention in a text. Use a NER model for that.
- English is the weakest of the 13 languages relative to competitors.
License
Desert Ant Labs Source-Available License. Free for most apps, and a commercial license is required at scale. Full terms are at the link. Licensing: licensing@desertant.com.
Citation
@software{schemer_2026,
title = {Schemer: On-device structured extraction into a caller-supplied JSON schema},
author = {Desert Ant Labs},
year = {2026},
url = {https://huggingface.co/desert-ant-labs/schemer},
}
© 2026 Desert Ant Labs · https://desertant.com
- Downloads last month
- 109