tnh0527 commited on
Commit
09ee555
Β·
verified Β·
1 Parent(s): 8ac074c

Publish gliclass-std-base-v3-daecore-5facet-qint8-v2 (qualified ONNX export and complete attribution)

Browse files
Files changed (3) hide show
  1. MODIFICATIONS.md +26 -0
  2. README.md +70 -28
  3. publication-manifest.json +10 -3
MODIFICATIONS.md ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Modifications from the upstream model
2
+
3
+ `gliclass-std-base-v3-daecore-5facet-qint8-v2` is not an unmodified copy of
4
+ `knowledgator/gliclass-base-v3.0`. It is not endorsed by Knowledgator or by
5
+ Microsoft, whose `deberta-v3-base` backbone the upstream model builds on.
6
+
7
+ What changed:
8
+
9
+ - **Fine-tune.** The upstream GLiClass single-pass classifier was trained
10
+ further on 59,886 labeled passages for five fixed facets (`trap`,
11
+ `decision`, `constraint`, `mechanism`, `procedure`) with the five label
12
+ prompts recorded in `classifier-metadata.json`. Labels were frontier-model
13
+ judgments under a frozen protocol, not human annotations.
14
+ - **Export.** The fine-tuned weights were exported to ONNX (opset 17) as a
15
+ single graph taking `input_ids` and `attention_mask` and emitting five
16
+ logits.
17
+ - **Quantization.** Only the token-embedding `Gather` tables were quantized
18
+ to signed INT8 (per-channel off, reduce-range off); every matrix product
19
+ stays FP32. Twenty-four constant identity nodes were folded. The compression
20
+ was qualified as equivalent to the FP32 export on a 5,298-row panel.
21
+ - **Calibration.** Per-facet temperature scaling and two frozen threshold
22
+ tables (`recall_leaning`, `contract`) travel in `classifier-metadata.json`
23
+ and are part of the artifact's identity.
24
+
25
+ Unchanged: the tokenizer vocabulary and the `<<LABEL>>` / `<<SEP>>` prompt
26
+ convention of the upstream model.
README.md CHANGED
@@ -19,8 +19,8 @@ This model tags passages of working notes and documentation with five
19
  independent facets: `trap`, `decision`, `constraint`, `mechanism`, and
20
  `procedure`. It is a supervised fine-tune of
21
  [`knowledgator/gliclass-base-v3.0`](https://huggingface.co/knowledgator/gliclass-base-v3.0)
22
- trained on 59,886 labeled passages and shipped as a 452.8 MB ONNX bundle
23
- qualified on CPU, CUDA, and DirectML.
24
 
25
  On a frozen 1,300-row held-out panel drawn from three unseen document families,
26
  it reached macro average precision 0.9676 against 0.8526 for the previous
@@ -46,11 +46,16 @@ the model behaves on unrelated real documents has not been measured.
46
  | Serialized artifact | `model.onnx`, 452,812,018 bytes, opset 17 |
47
  | Quantization | signed INT8 on the token-embedding matrix only; the transformer body stays FP32 |
48
  | Threshold tables | two frozen operating views, `contract` and `recall_leaning` |
49
- | Deployment state | promotion-qualified; activation as the shipped default is a separate step |
50
 
51
  The artifact size follows from the quantization choice: the 128k-entry
52
  embedding table, about 98M parameters, is stored at one byte per weight while
53
- the 88M-parameter body stays at four. Full-graph INT8 was not qualified.
 
 
 
 
 
54
 
55
  ## Intended use
56
 
@@ -77,9 +82,11 @@ both training and evaluation.
77
 
78
  **Label provenance.** Every label was produced by a language model applying the
79
  same one-sentence facet definitions the classifier sees. Training rows received
80
- one primary judgment from GPT-5.6 (Sol, medium reasoning), with a frozen
81
- 400-row cross-model audit by GPT-5.6 (Terra, high) and a Sol-high tiebreak only
82
- where the two disagreed semantically. Calibration and promotion panels used two
 
 
83
  independent Sol-high sessions per row plus a third Sol-xhigh session that saw
84
  only the disputed fields. Unresolved fields stay unresolved and are masked per
85
  facet; nothing unresolved becomes a negative.
@@ -96,20 +103,36 @@ agreement with that instrument, not with human ground truth.
96
 
97
  | Lane | Rows | Origin |
98
  |---|---:|---|
99
- | Final training set | 59,886 | 54,803 generated, 5,083 from the operator's own project documentation |
 
 
 
 
 
100
 
101
  Generated documents come from fictional organizations written as complete
102
  documents (procedures, incident reports, decision records, design notes) and
103
  then passed through the production parser. They were preferred over the
104
- operator's real corpus for both training and evaluation because they cover many
105
- more domains, document shapes, and organizational settings; the real corpus is
106
- one person's projects and too narrow to generalize from.
 
 
 
 
 
 
 
 
 
 
107
 
108
  Splits are by whole source family, never by row. The seven families reserved
109
  for calibration and promotion were removed from training along with 3,875
110
  predecessor training rows that shared them. Calibration and promotion families
111
- are disjoint from each other and from training by family, document lineage,
112
- exact passage identity, and lexical near-duplicate screening.
 
113
 
114
  The final fit used weighted binary cross-entropy with a 2.0 multiplier on
115
  mentions-only hard negatives, learning rate 2e-5, batch 4 with 8 accumulation
@@ -168,8 +191,7 @@ is the classifier this model replaces.
168
  | procedure | 0.375 | 0.9578 | 0.8167 | 0.886 | 0.528 | 0.822 | 0.489 |
169
  | macro | 0.631 | 0.9676 | 0.8526 | 0.902 | 0.719 | 0.860 | 0.687 |
170
 
171
- The macro gain is 0.1150. The same candidate gained 0.1200 on the separate
172
- calibration panel, so the two panels agree within 0.005. The one-sided 95%
173
  lower bound on the macro precision gain under the frozen contract thresholds is
174
  0.009; the decision-facet precision gain has a negative lower bound, because the
175
  previous classifier's decision threshold was so strict that it predicted only 26
@@ -278,8 +300,9 @@ Paired macro average-precision differences from this model, with family-grouped
278
  [βˆ’0.0548, βˆ’0.0366]; GPT-5.5 low βˆ’0.0588 [βˆ’0.0666, βˆ’0.0442]; GPT-5.4 low
279
  βˆ’0.0937 [βˆ’0.1015, βˆ’0.0827]. Every interval excludes parity. Raising reasoning
280
  effort improved GPT-5.4 by 0.0509 [0.0456, 0.0589] and GPT-5.5 by 0.0080
281
- [0.0014, 0.0204]; it narrows the gap without closing it. Two unplanned repeats
282
- of the Spark run spanned 0.8086 to 0.8496, a spread of 0.041 that nearly
 
283
  matches the smallest GPT gap and exceeds every one of the four interval widths,
284
  which is the reason single runs are flagged above.
285
 
@@ -313,10 +336,13 @@ precision unless stated.
313
  3. **Context window.** Widening the input from 416 to 768 tokens added +0.0048
314
  [+0.0028, +0.0068] on development and +0.0033 [+0.0011, +0.0065] on a
315
  separate held-out set. A counterfactual-projector branch on minimal pairs
316
- changed nothing (+0.0001) and was dropped.
317
- 4. **Backbone size.** With the full data and a matched recipe on a 5,298-row
318
- eight-family panel, the larger backbone beat the base by 0.0042 average
319
- precision and 0.0152 at 90% recall [+0.0015, +0.0451], but cost 2.6 times
 
 
 
320
  the artifact bytes, 2.4 times the training time, and 1.7 times peak GPU
321
  memory, while the base scored 1.8 times as many rows per second. The base was
322
  selected as the product architecture.
@@ -328,10 +354,18 @@ precision unless stated.
328
  6. **Final fit.** The frozen recipe was trained once on all 59,886 rows,
329
  calibrated, compressed, and read once on the promotion panel.
330
 
331
- What did not help, in one place: more real documents without generated
332
- diversity, the data curriculum, the robust loss, the counterfactual projector,
333
- and the larger backbone at its cost. Better and more diverse examples, boundary
334
- cases, and a wider context did the work.
 
 
 
 
 
 
 
 
335
 
336
  ## Runtime
337
 
@@ -387,7 +421,7 @@ CUDA and 277 MiB on DirectML above a resident session of 657 MiB and 481 MiB res
387
  | Previous-classifier identity | verified 2026-09-04 against the staged bundle's own config |
388
  | Frontier-model comparisons | as recorded; single runs on a subsample that is not independent of the promotion read |
389
  | Serving throughput | 3.74 passages per second, as recorded from the scoring run bound to the compression score receipt |
390
- | Per-passage latency percentiles | not measured |
391
 
392
  ## Limitations, ranked
393
 
@@ -416,8 +450,16 @@ CUDA and 277 MiB on DirectML above a resident session of 657 MiB and 481 MiB res
416
  8. **Serving cost is measured on one machine.** The provider table in the runtime
417
  section comes from a single 8 GB NVIDIA card and a four-core CPU budget; other
418
  hosts will land elsewhere, and cold-start cost is not separated from warm batches.
419
- 9. **Not yet the shipped default.** The artifact is promotion-qualified; the
420
- activation, restamp, and rollback steps are separate operations.
 
 
 
 
 
 
 
 
421
 
422
  ## Reproducibility
423
 
 
19
  independent facets: `trap`, `decision`, `constraint`, `mechanism`, and
20
  `procedure`. It is a supervised fine-tune of
21
  [`knowledgator/gliclass-base-v3.0`](https://huggingface.co/knowledgator/gliclass-base-v3.0)
22
+ trained on 59,886 labeled passages and shipped as a 452.8 MB ONNX graph inside a
23
+ 461.5 MB bundle, qualified on CPU, CUDA, and DirectML.
24
 
25
  On a frozen 1,300-row held-out panel drawn from three unseen document families,
26
  it reached macro average precision 0.9676 against 0.8526 for the previous
 
46
  | Serialized artifact | `model.onnx`, 452,812,018 bytes, opset 17 |
47
  | Quantization | signed INT8 on the token-embedding matrix only; the transformer body stays FP32 |
48
  | Threshold tables | two frozen operating views, `contract` and `recall_leaning` |
49
+ | Deployment state | the installed default since 2026-09-11; published 2026-09-06 |
50
 
51
  The artifact size follows from the quantization choice: the 128k-entry
52
  embedding table, about 98M parameters, is stored at one byte per weight while
53
+ the 88M-parameter body stays at four. Broader quantization was measured, not
54
+ skipped: on the predecessor fit of the same recipe, six arms were scored
55
+ against the FP32 reference on a 5,298-row panel,
56
+ and every one beyond the embedding table cost quality β€” embedding plus attention
57
+ βˆ’0.0013, full per-channel reduced-range βˆ’0.0531, and full per-channel βˆ’0.0599
58
+ macro average precision.
59
 
60
  ## Intended use
61
 
 
82
 
83
  **Label provenance.** Every label was produced by a language model applying the
84
  same one-sentence facet definitions the classifier sees. Training rows received
85
+ two independent judgments each with a higher-effort tiebreak on semantic
86
+ disagreement, except for 3,750 inherited rows that received one primary judgment
87
+ from GPT-5.6 (Sol, medium reasoning) with a frozen 400-row cross-model audit by
88
+ GPT-5.6 (Terra, high) and a Sol-high tiebreak where the two disagreed.
89
+ Calibration and promotion panels used two
90
  independent Sol-high sessions per row plus a third Sol-xhigh session that saw
91
  only the disputed fields. Unresolved fields stay unresolved and are masked per
92
  facet; nothing unresolved becomes a negative.
 
103
 
104
  | Lane | Rows | Origin |
105
  |---|---:|---|
106
+ | Final training set | 59,886 | 54,803 generated; 4,583 real rows from the operator's self-owned corpora and 500 real rows of public-domain US Federal Register text |
107
+
108
+ The 59,886 passages come from 9,376 documents, about 6.4 passages each: 8,212
109
+ generated documents spanning 143 fictional organizations, and 1,164 real
110
+ documents β€” 1,042 from the operator's two workspaces and 122 from the Federal
111
+ Register set.
112
 
113
  Generated documents come from fictional organizations written as complete
114
  documents (procedures, incident reports, decision records, design notes) and
115
  then passed through the production parser. They were preferred over the
116
+ real lane for both training and evaluation because they cover many more domains,
117
+ document shapes, and organizational settings. The real lane is narrower than its
118
+ family count suggests: 315 of its 437 source families are directory groupings inside
119
+ two workspaces belonging to one person, and the other 122 are single Federal
120
+ Register dockets, not 437 organizational settings; the Federal Register rows are
121
+ also regulatory prose from a single genre. Whole-family
122
+ splitting is therefore a strong isolation guarantee in the generated lane, where
123
+ a family is a whole organization, and a weak one in the real lane.
124
+
125
+ Every real row is redistributable β€” the operator's corpora by ruling,
126
+ the Federal Register text as public-domain US government work. The Federal
127
+ Register rows are training supervision only: no panel that calibrates, gates, or
128
+ reports on this model contains them, or any other real row.
129
 
130
  Splits are by whole source family, never by row. The seven families reserved
131
  for calibration and promotion were removed from training along with 3,875
132
  predecessor training rows that shared them. Calibration and promotion families
133
+ are disjoint from each other and from training by family, document lineage, and
134
+ exact passage identity. Assignment was label-blind and checked for support only
135
+ afterward.
136
 
137
  The final fit used weighted binary cross-entropy with a 2.0 multiplier on
138
  mentions-only hard negatives, learning rate 2e-5, batch 4 with 8 accumulation
 
191
  | procedure | 0.375 | 0.9578 | 0.8167 | 0.886 | 0.528 | 0.822 | 0.489 |
192
  | macro | 0.631 | 0.9676 | 0.8526 | 0.902 | 0.719 | 0.860 | 0.687 |
193
 
194
+ The macro gain is 0.1150. The one-sided 95%
 
195
  lower bound on the macro precision gain under the frozen contract thresholds is
196
  0.009; the decision-facet precision gain has a negative lower bound, because the
197
  previous classifier's decision threshold was so strict that it predicted only 26
 
300
  [βˆ’0.0548, βˆ’0.0366]; GPT-5.5 low βˆ’0.0588 [βˆ’0.0666, βˆ’0.0442]; GPT-5.4 low
301
  βˆ’0.0937 [βˆ’0.1015, βˆ’0.0827]. Every interval excludes parity. Raising reasoning
302
  effort improved GPT-5.4 by 0.0509 [0.0456, 0.0589] and GPT-5.5 by 0.0080
303
+ [0.0014, 0.0204]; it narrows the gap without closing it. Three accepted runs
304
+ of the Spark configuration spanned 0.8086 to 0.8496 β€” the reported 0.8086 plus
305
+ repeats at 0.8180 and 0.8496 β€” a spread of 0.041 that nearly
306
  matches the smallest GPT gap and exceeds every one of the four interval widths,
307
  which is the reason single runs are flagged above.
308
 
 
336
  3. **Context window.** Widening the input from 416 to 768 tokens added +0.0048
337
  [+0.0028, +0.0068] on development and +0.0033 [+0.0011, +0.0065] on a
338
  separate held-out set. A counterfactual-projector branch on minimal pairs
339
+ changed nothing against its own replay control (+0.0001) and was dropped; the
340
+ run's window also did not match its draft plan, so it could not have
341
+ authorized a recipe change either way.
342
+ 4. **Backbone size.** On the ModernBERT GLiClass pair, with 44,049 training
343
+ rows and a matched recipe on a 5,298-row eight-family panel, the large arm beat
344
+ the base by 0.0042 average precision and 0.0152 at 90% recall [+0.0015,
345
+ +0.0451] β€” at 95% recall the interval spans zero β€” but cost 2.6 times
346
  the artifact bytes, 2.4 times the training time, and 1.7 times peak GPU
347
  memory, while the base scored 1.8 times as many rows per second. The base was
348
  selected as the product architecture.
 
354
  6. **Final fit.** The frozen recipe was trained once on all 59,886 rows,
355
  calibrated, compressed, and read once on the promotion panel.
356
 
357
+ What did not help: more real documents without generated diversity, the data
358
+ curriculum, the robust loss, the counterfactual projector, and the larger
359
+ backbone at its cost. A distribution-balanced loss, R-Drop, SMART smoothness, a
360
+ three-state auxiliary head, and a 512-token window were also measured and
361
+ rejected, and the experiment log carries each with its numbers. Two further
362
+ branches closed without a scored comparison: a frozen large-NLI recipe, stopped
363
+ after 23.2 hours with no completed epoch, and a decision-only rationale
364
+ auxiliary, which trained four epochs and reduced all three training losses but
365
+ failed artifact assembly on incomplete checkpoint-average candidates, so no
366
+ panel was mounted and no per-row scores were emitted. It is retired as
367
+ inconclusive. Better and more diverse examples, boundary cases, and a wider
368
+ context did the work.
369
 
370
  ## Runtime
371
 
 
421
  | Previous-classifier identity | verified 2026-09-04 against the staged bundle's own config |
422
  | Frontier-model comparisons | as recorded; single runs on a subsample that is not independent of the promotion read |
423
  | Serving throughput | 3.74 passages per second, as recorded from the scoring run bound to the compression score receipt |
424
+ | Per-passage latency percentiles | single-passage p50 per provider only; no p95 or p99 |
425
 
426
  ## Limitations, ranked
427
 
 
450
  8. **Serving cost is measured on one machine.** The provider table in the runtime
451
  section comes from a single 8 GB NVIDIA card and a four-core CPU budget; other
452
  hosts will land elsewhere, and cold-start cost is not separated from warm batches.
453
+ 9. **Second sealed campaign in its family.** An earlier fit was read on a
454
+ sealed panel, returned NO-GO on a per-origin recall gate, and the contract was
455
+ then amended so origin slices are diagnostics rather than independent vetoes.
456
+ That fit's calibration and promotion rows were folded into this model's own
457
+ training set, and this campaign therefore grants its panel promotion authority
458
+ while explicitly declining a claim to a pristine sealed alpha-spending draw.
459
+ The earlier panel carried 333 operator-workspace rows; this one is entirely
460
+ generated, so the per-origin recall that failed on the earlier read cannot be
461
+ measured on this model at all. Resolving evidence would be a promotion panel
462
+ drawn from unspent reserve that includes real rows.
463
 
464
  ## Reproducibility
465
 
publication-manifest.json CHANGED
@@ -1,12 +1,14 @@
1
  {
2
  "schema": "daecore.classifier-publication-manifest",
3
- "staged_at": "2026-09-06T22:38:02+00:00",
4
  "public_repo": "Daecore/gliclass-std-base-v3-5facet-qint8-v2",
5
  "classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
6
  "upstream": {
7
  "model_id": "knowledgator/gliclass-base-v3.0",
8
  "revision": "77a70e6cd52e602ed18184ef37d18bdd3741e3d5"
9
  },
 
 
10
  "files": {
11
  "model.onnx": {
12
  "sha256": "690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c",
@@ -35,8 +37,8 @@
35
  "binding": "nominee receipt runtime_metadata_sha256 (canonical JSON hash)"
36
  },
37
  "README.md": {
38
- "sha256": "e7dc7ef045a4daa57f85e2213da0fe7273322737417bd16147afe8adfc73b01b",
39
- "size": 26973,
40
  "binding": "packaging record"
41
  },
42
  "LICENSE": {
@@ -48,6 +50,11 @@
48
  "sha256": "71c772c9388cdbdf68ca17f9e2f7f85f1ed7c210a326885a5bd64ed06bb44b3c",
49
  "size": 822,
50
  "binding": "packaging record"
 
 
 
 
 
51
  }
52
  }
53
  }
 
1
  {
2
  "schema": "daecore.classifier-publication-manifest",
3
+ "tool_sha256": "ec0f21b2c945a07b0a1f7be9895c91f30a985fb981a5fb83f409604b33879af4",
4
  "public_repo": "Daecore/gliclass-std-base-v3-5facet-qint8-v2",
5
  "classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
6
  "upstream": {
7
  "model_id": "knowledgator/gliclass-base-v3.0",
8
  "revision": "77a70e6cd52e602ed18184ef37d18bdd3741e3d5"
9
  },
10
+ "nominee_receipt_sha256": "d463fc167b2a26f8f5b56e03c18e4c604b231eb25b5093064dd8a9bd7ffa6373",
11
+ "staged_at": "2026-09-18T18:58:21+00:00",
12
  "files": {
13
  "model.onnx": {
14
  "sha256": "690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c",
 
37
  "binding": "nominee receipt runtime_metadata_sha256 (canonical JSON hash)"
38
  },
39
  "README.md": {
40
+ "sha256": "48cc3f3e7659463c7bd6031d923a96f3822e7ebd62a2d5f660d779d95612f76c",
41
+ "size": 29930,
42
  "binding": "packaging record"
43
  },
44
  "LICENSE": {
 
50
  "sha256": "71c772c9388cdbdf68ca17f9e2f7f85f1ed7c210a326885a5bd64ed06bb44b3c",
51
  "size": 822,
52
  "binding": "packaging record"
53
+ },
54
+ "MODIFICATIONS.md": {
55
+ "sha256": "fca339612182863205137d75b60cb23dc921090c0e91a4541a382c3889cd7e95",
56
+ "size": 1399,
57
+ "binding": "packaging record"
58
  }
59
  }
60
  }