File size: 9,283 Bytes
28c70af
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
# `bankML/schema.rs` — JSON schema to GBNF, as llama-server b11192 builds it

## Summary

`schema.rs` turns a JSON schema into the GBNF text llama-server b11192 would build for it, byte for byte. It is a port
of llama.cpp b11192 (github.com/ggml-org/llama.cpp @ 171e8846b; MIT, © the ggml authors, whose notice the file's
header carries as the licence asks, see LICENSING.md) of:
- `common/json-schema.cpp`: the schema tree, which keywords decide a node's kind, `$ref` resolution, the errors;
- `common/json-schema-to-grammar.cpp` (`common_chat_schema_converter`): rule naming and de-duplication, objects with
  required / optional / additional properties, `_not_strings`, arrays and tuples, integer ranges, the regex → GBNF
  pattern translation, string formats, the primitive rules;
- `common/trie.cpp`, which `_not_strings` walks;
- the parts of nlohmann's `ordered_json` the grammar text depends on: object key order and duplicate keys, the
  integer / float distinction, `dump()`;
- the GBNF the PEG chat parser adds around the schema (`common/chat-auto-parser-generator.cpp`,
  `common/peg-parser.cpp`): the `json-*` rules, `response-format`, `root`, and on the Qwen3 template the `until-13`
  rules for the reasoning block (`until("</think>")`, parser id 13 on the pinned template, with `--reasoning off`).

The grammar text matters because the answer depends on it: the grammar engine ([grammar.md](grammar.md)) masks the
same tokens only if it parses the same rules.

Callers: `grammar.rs` (`from_openai*` and `from_ollama*` convert a schema at request parse time to accept or refuse
it); `native.rs` (`Native::grammar` builds the grammar for the model's template); `serve.rs` (`NativeChat::parse_ctx`
checks that the grammar parses before any token is computed).
The surfaces are `/v1/chat/completions` (`response_format` `json_schema`, `json_object` + `schema`, the top-level
`json_schema`), Ollama's `/api/chat` and `/api/generate` with `format: <schema>`, and the C API's `bankml_chat`.

## Technical usage

```rust
pub enum Value { Null, Bool(bool), Int(i64), UInt(u64), Float(f64), Str(String), Arr(Vec<Value>), Obj(Vec<(String, Value)>) }
impl Value {
    pub fn parse(s: &str) -> Option<Value>
    pub fn from_json(j: &crate::serve::Json) -> Value
    pub fn get(&self, k: &str) -> Option<&Value>
    pub fn is_null(&self) -> bool
    pub fn is_empty(&self) -> bool
    pub fn as_str(&self) -> Option<&str>
    pub fn dump(&self) -> String
}
pub fn json_schema_to_grammar(schema: &Value) -> Result<(String, Vec<String>), String>
pub fn chat_grammar(schema: &Value, t: crate::chat::Template) -> Result<(String, Vec<String>), String>
pub fn format_literal(s: &str) -> String
```

- `Value` is JSON as nlohmann's `ordered_json` sees it: integers apart from floats (`1` vs `1.0`), unsigned apart
  from signed, object keys in first-seen order with the last duplicate's value. `Value::parse` is the exact reading
  (nesting deeper than 512 levels is refused). `from_json` converts bankml's request JSON, whose numbers are f64, so
  a schema written `2.0` reads as `2`; the servers re-read the body with `parse` to keep it a float.
- `json_schema_to_grammar` is llama.cpp's `json_schema_to_grammar(schema)` (`force_gbnf`): the grammar whose root is
  the schema. The second value holds warnings, for example a pattern it reads but cannot translate, which becomes
  "any string" as in llama.cpp.
- `chat_grammar(schema, template)` is what llama-server builds on the jinja path with thinking off: the schema's rules
  and the parser's rules in one converter, with `response-format`, `root`, and on Qwen3 the `until-13` rules. On the
  ChatML templates (SmolLM2-Instruct, `mindx-genN`) there is no reasoning block. Its errors carry llama-server's
  prefix: `Unable to generate parser for this template. Automatic parser generation failed: `.
- `{"type": "object"}` gives exactly each template's JSON-mode grammar (`grammar::JSON_OBJECT_GRAMMAR` /
  `_CHATML`); `grammar.rs` then treats the request as plain JSON mode.

Supported schema shapes include `$ref` into the same document (a `#/…` pointer, such as `#/$defs/…`), `anyOf` /
`oneOf`, `allOf`, `const`, `enum`, `null`, `boolean`, `number`, `integer` with `minimum` / `maximum` (and the
exclusive forms), `string` with `pattern`, `minLength` / `maxLength` and the formats `date`, `time`, `date-time` and
`uuid`, arrays with `items`, `minItems` / `maxItems`, tuples (`prefixItems`), and objects with `properties`,
`required` and `additionalProperties`.

```sh
curl -s 127.0.0.1:PORT/v1/chat/completions -d '{
  "messages": [{"role": "user", "content": "A cat"}],
  "response_format": {"type": "json_schema", "json_schema": {"name": "cat", "schema":
    {"type": "object", "properties": {"name": {"type": "string"}, "age": {"type": "integer", "minimum": 0}},
     "required": ["name"]}}}}'
```

## How it is verified

- Unit tests: `json_reads_and_prints_as_nlohmann`, `the_object_schema_is_the_json_mode_grammar`, `refusals_say_where`,
  and `llama_cpp_test_cases`, which carries llama.cpp's own `tests/test-json-schema-to-grammar.cpp` cases with their
  expected grammars (`testing/json_schema_cases.json`: 81 cases; the test requires at least 80). Each expected
  grammar must come out the same (indentation aside) and parse in `grammar.rs`; each case llama.cpp refuses must be
  refused.
- `oracle_schema_grammars` (`#[ignore]`, in the gate): `testing/schema_oracle.{cpp,py}` calls llama.cpp b11192's own
  `json_schema_to_grammar` and `common_chat_templates_apply` inside the release's `libllama-common.so` (no model) on
  173 schemas: llama.cpp's 81 test cases, Pydantic-shaped schemas like mindX's, edge cases and every refusal path,
  with the template read from each model that carries it. **148 grammars byte-identical bare and on the chat path,
  on each of the three templates; every refusal with llama.cpp's message** (24 bare, 20 chat). Every grammar must
  also parse in `grammar.rs`.
- `oracle_json_schema`, `_ternary`, `_o4` (`native.rs`): `testing/json_schema_oracle.py --record` against
  llama-server b11192, greedy and seeded, from an empty cache, including answers cut by `max_tokens`: Bonsai-8B
  **28 of 28**, Ternary-Bonsai-8B **11 of 11**, Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39 **56 of 56** each.
  Live in the gate through `/v1` (one streamed) and `/api/chat` with `format: <schema>`.

## Design notes

- This is milestone O6b (JSON schemas). JSON mode's grammar was first a constant (0.3.3), checked against the
  grammar llama-server reports; it is now built by the converter, and `the_object_schema_is_the_json_mode_grammar`
  checks the two agree.
- The ChatML templates without reasoning (SmolLM2-Instruct, `mindx-genN`) were added in 0.3.5: their root is the
  generation prompt then the value, with no `until-13` rules, as llama.cpp b11192 emits on both templates
  (`oracle_schema_grammars` records each template from the model that carries it).
- The non-ASCII-before-a-quantifier refusal exists because llama.cpp's grammar for it is not UTF-8 and would misread
  the character; bankml refuses rather than reproduce it.

## Advantages and efficiency

- **Any schema, llama-server's grammar.** mindX's Pydantic-shaped schemas get the exact grammar llama-server would
  use, so the answers are token-identical and checkable, without llama.cpp at run time.
- **Refusals before generation.** A schema is converted when the request is parsed, so a schema b11192 refuses is a
  400 with its message before any token is computed.
- **Bounded on hostile input.** JSON nesting is capped at 512 levels and regex group nesting at 100
  (`MAX_PATTERN_DEPTH`); errors are `Result`s with a path (`JSON schema error at #/properties/a: …`), not panics.
- **Rust practice.** Zero crates (nlohmann's number printing and the trie are written out); no `unsafe`.
- **Next** (docs/TODO.md, docs/OLLAMA.md): tool calls through the template (O6), which build on the schema converter.

## Limitations

- Refused as b11192 refuses, with its message: an unknown type, an empty `enum`, a `$ref` outside the document, a
  pattern no regex reads, and the other paths `oracle_schema_grammars` records.
- One deliberate divergence: a `pattern` with a non-ASCII character before a quantifier (for example `^é+$`).
  llama.cpp splits the character byte by byte into a grammar that is not UTF-8; bankml refuses it and suggests a
  group, `(é)+`.
- Unsupported regex features (for example `\d`) are not refused: the string accepts anything, with a warning, as in
  llama.cpp.
- Known limit, measured: a float in an `enum` or `const` is printed with Rust's shortest round-trip digits, which agree
  with nlohmann's Grisu2 except on rare doubles where Grisu2 is not shortest.
- `chat_grammar` reproduces the jinja path with thinking off on the three reproduced templates only.

## See also

- [../oracles.md](../oracles.md) §5e — the schema oracles
- [../OLLAMA.md](../OLLAMA.md) — O6b, the chat path's wrapping per template
- [../TODO.md](../TODO.md) — 0.4.0, 0.6.0
- [../usage.md](../usage.md) — `/v1` and `/api` request fields
- Sibling pages: [grammar.md](grammar.md), [chat.md](chat.md), [sampler.md](sampler.md), [native.md](native.md),
  [serve.md](serve.md), [ollama.md](ollama.md)