code-daemon-ner-v1

A small token-classification (BIO) model that finds software-engineering entities in prose β€” the sentences of a README, a design doc, an issue thread, a commit message. It ships with the UltraCode MCP server as a TensorRT / OpenVINO engine and feeds the documentation knowledge graph: the spans it emits become the concept nodes that link a paragraph of prose to the code entities it talks about.

This is prose NER, not code parsing. Identifiers inside source files are extracted by tree-sitter, exactly and for free. What tree-sitter cannot do is read the sentence "the IVF index rescoring path in faiss_index.zig regressed nDCG@10 by 4 points" and tell you that IVF is an algorithm, faiss_index.zig a file path, and nDCG@10 a metric. That is this model's job.

  • ~117M params β€” XLM-RoBERTa 12 layers / 384 hidden, 250k multilingual SentencePiece vocab (the embedding table dominates the parameter count).
  • 2-input ONNX (input_ids, attention_mask; no token_type_ids) β†’ logits [batch, seq, 17].
  • Max sequence 128 tokens β€” one or two sentences, which is the unit the daemon chunks prose into.
  • Bilingual by construction (English + Russian docs), inherited from the XLM-R backbone.

What this model is for. Entity-level Core F1 is 0.43. That is a useful recall-oriented candidate generator for a knowledge graph, and it is not a labeller to trust unattended. The Evaluation section has the per-type numbers, including the two types that are weak and why.

Label set β€” 8 entity types, 17 BIO classes

Split into a Core tier and a Soft tier, scored separately.

Core β€” what the knowledge graph links on:

  • component β€” a named part of a system, ours or someone's, including model names. DocStore, mMiniLMv2
  • api_endpoint β€” a callable surface: function, method, route, agent tool name. semantic_search, POST /v1/messages
  • file_path β€” a path or a filename. src/semantic/faiss_index.zig, config.toml
  • tool β€” an external executable, runtime, service, language or hardware you invoke. TensorRT, rclone, SQL, Tesla T4

Soft β€” useful, tolerated at lower precision:

  • algorithm β€” a named method or procedure. BM25, beam search, Louvain
  • data_structure β€” a named container or layout. CSR, Prolly tree, arena
  • config β€” a knob, flag, env var or setting. journal_mode=MEMORY, --epochs
  • metric β€” a measurement or its unit. nDCG@10, macro-F1, p95 latency

The component / tool / api_endpoint boundary costs annotators the most, so it is fixed by rule rather than by feel, and the rules are applied in order β€” that ordering is what resolves names fitting two types. A model is a component (a part of a system); an engine, runtime, language or piece of hardware is a tool (you invoke it); an agent tool name is an api_endpoint (a callable surface, not a program). About a dozen names are genuinely polysemous and take their label from the sentence.

How it was made

There is no off-the-shelf corpus for this label set, so the training labels were manufactured by weak supervision β€” several independent labellers vote per chunk and the votes are reconciled into a single BIO sequence β€” followed by supervised fine-tuning of the backbone and a temperature calibration pass.

The calibration is the part you need at inference: a single scalar, shipped as temperature.json (T = 0.665). Divide the logits by it before the softmax if you consume probabilities rather than argmax.

Evaluation is human-labelled and disjoint from the training labels β€” 1 500 hand-reviewed sentences, normalised so that no surface name carries two labels and no span overlaps another.

Evaluation

Measured on the 1 500-row human gold set, entity-level (exact span and type, seqeval convention), macro-averaged over the bare types β€” the metric NER results are normally reported in.

Split macro-F1 micro-F1
Core 0.4264 0.4502
Soft 0.2072 0.1910

Per type, with precision and recall, because they differ a lot and the difference is the story:

Type Tier F1 P R
api_endpoint Core 0.521 0.409 0.720
file_path Core 0.497 0.391 0.682
tool Core 0.425 0.421 0.429
component Core 0.263 0.218 0.330
metric Soft 0.273 0.404 0.206
algorithm Soft 0.239 0.193 0.312
data_structure Soft 0.198 0.169 0.239
config Soft 0.119 0.166 0.093

What this says, plainly:

  • Recall runs well ahead of precision on the strongest types, and that is a property of the training labels rather than of the model: api_endpoint and file_path are labelled in the training corpus at roughly 2.6Γ— the density the human gold uses, and the model over-predicts them by about the same factor. Over-labelled in, over-predicted out. If you use it as a candidate generator and filter afterwards, this is the behaviour you want; as a final labeller it is not.
  • component is the weakest Core type (0.263). Unlike tool it is not under-labelled β€” its training density already exceeds the gold's β€” so density is not what limits it. It is genuinely the hardest of the four: what counts as a named part of a system is contextual in a way a path or a call is not.
  • config is the weakest overall (0.119, recall 0.093). Its labels are inconsistent between the two halves of the training corpus β€” dense in documentation, nearly absent in Q&A β€” and averaging two incompatible labelling regimes yields noise. A known, unfixed defect.
  • Multi-token spans are weak. The I- classes are near-empty in training, so most predictions are effectively single-token. Only ~4% of Core gold spans are multi-word, so this costs less than it sounds, but do not expect long names to come back whole.

For direction rather than level: the model this replaces scores 0.077 Core on the same ruler, and the previous build of this one scored 0.323. All three were measured identically.

Against a zero-shot NER β€” the alternative to fine-tuning

The honest comparison for a private label set is not another fine-tune, it is a zero-shot model you hand the label names to. GLiNER is that family.

What matches and what does not. No public model exists at this size for these eight labels β€” that absence is why this model was trained β€” so the comparators match on task, gold and ruler but not on size (289 M and 909 M against 117 M) or on runtime. Read the speed column below as an order of magnitude. Same 1 500-row gold, same ruler, same whitespace-token grid; each baseline swept over two label phrasings Γ— two thresholds and reported at its best of the four:

model params macro micro
this model 117 M 0.426 0.450
gliner-multitask-v1.0 909 M 0.284 0.296
gliner_multi-v2.1 289 M 0.155 0.193

Precision / recall, same order: 0.38 / 0.55 Β· 0.23 / 0.41 Β· 0.17 / 0.23.

Throughput over the same 1 500 chunks, in chunks tagged per second:

model chunks/s where
this model 68 CPU, onnxruntime
gliner-multitask-v1.0 10 GPU, PyTorch
gliner_multi-v2.1 39 GPU, PyTorch

This model is faster on a CPU than either baseline on a GPU.

The gap is not uniform across types, and the shape of it is the useful part:

type this multitask multi
api_endpoint 0.521 0.382 0.006
file_path 0.497 0.300 0.248
tool 0.425 0.312 0.286
component 0.263 0.143 0.080

Where a label name carries its own meaning β€” file_path, tool β€” zero-shot gets within ~1.5Γ— of this model. Where the label is a project convention, it collapses: api_endpoint here means "a callable surface, including an agent tool name", which no amount of prompting recovers, and one of the two baselines scores 0.006 on it.

So the thing being bought by fine-tuning is the label convention, not the architecture. If your entity types are ordinary ones, a zero-shot model may serve you better than this one; if they are yours, they have to be trained in.

Measured 2026-09-11 on a laptop RTX 5060; harness and raw JSON ship in the UltraCode repo (models/_distill_shared/bench_vs_generic.py).

Inside a real index

Measured live in the UltraCode daemon, TensorRT, sharing the GPU with the other workers as it does in production. Per 32-chunk batch:

stage ms share
tokenize (5 threads) 15 75 %
GPU inference 3 15 %
total 20

The model is not the bottleneck β€” the tokenizer is. The GPU could run ~10 700 chunks/s; the pipeline feeds it ~1 600. If you embed this model yourself, parallelise tokenisation before reaching for a faster engine.

Intended use

Extract engineering entities from documentation-style prose to build or enrich a knowledge graph, to tag issues, or to route search. Use it as a recall-oriented candidate generator behind a rule or human filter, not as a labeller you trust unattended.

Out of scope: extracting symbols from source code (use a parser), general-domain NER (use a general-domain model), and any high-stakes automatic decision.

import onnxruntime as ort, numpy as np, json
from transformers import AutoTokenizer

tok    = AutoTokenizer.from_pretrained(".")            # bundled XLM-R SentencePiece
labels = json.load(open("code-daemon-ner-v1_label_map.json"))["id2label"]
T      = json.load(open("temperature.json"))["T"]
sess   = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])

def tag(text, max_len=128):
    enc = tok(text, return_tensors="np", truncation=True, max_length=max_len,
              return_offsets_mapping=True, return_token_type_ids=False)
    logits = sess.run(None, {"input_ids":      enc["input_ids"].astype(np.int64),
                             "attention_mask": enc["attention_mask"].astype(np.int64)})[0][0]
    probs  = np.exp(logits / T) / np.exp(logits / T).sum(-1, keepdims=True)   # calibrated
    for (a, b), row in zip(enc["offset_mapping"][0], probs):
        lab = labels[str(int(row.argmax()))]
        if lab != "O" and b > a:
            yield text[a:b], lab, float(row.max())

for span, label, p in tag("The IVF rescoring path in src/semantic/faiss_index.zig cost 4 nDCG@10 points."):
    print(f"{span!r:40} {label:18} {p:.2f}")

The numbers above were measured at sequence length 128 with identifier pre-splitting (camelCase / snake_case split before tokenizing), which is how the model was trained and how it is deployed. Scoring it at a different sequence budget, or without the pre-split, gives a different and lower result.

What's in this repo

Flat layout, named per runtime Γ— GPU arch Γ— OS (single-profile β€” no length buckets):

  • TensorRT code-daemon-ner-v1_{win_x64,linux_x64}_trt11.0_sm_{75,80,86,89,120}.engine and code-daemon-ner-v1_linux_x64_trt11.0_sm_90.engine β€” INT8 Q/DQ, for Turing, Ampere (A100 / A30 and RTX 30xx), Ada, Hopper (H100 / H200, Linux only) and Blackwell. A serialized engine is keyed on {GPU arch Γ— OS Γ— TensorRT version} with no compat fallback, so any other machine uses the OpenVINO or ONNX path.
  • OpenVINO 2026.4 code-daemon-ner-v1_ov2026.4_{cpu,igpu_lnl}_int8_b32_s128.{xml,bin} and ..._npu_int4_b8_s128.{xml,bin} β€” Intel CPU / iGPU / NPU.
  • Tokenizer β€” sentencepiece.bpe.model (XLM-R SP; the daemon loads it natively) plus tokenizer_config.json, which is what AutoTokenizer.from_pretrained reads. The fast tokenizer ships as code-daemon-ner-v1_tokenizer.json.
  • Labels & calibration β€” code-daemon-ner-v1_label_map.json (17 classes, label2id / id2label), temperature.json (a single scalar under the key T), manifest.json (build metadata), eval_metrics.json (per-type and per-class F1). The label map and the fast tokenizer carry the model-id prefix because that is the name the daemon resolves; the snippet above uses the same names.
  • ONNX β€” model.onnx FP32 dynamic-shape (the build source, and standalone onnxruntime / optimum use), model_static.onnx (seq pinned to 128, the engine-build input), and model_int8qdt.onnx (INT8 Q/DQ).
  • Raw weights β€” model.safetensors + config.json, the FP32 weights in the XLMRobertaForTokenClassification layout with the 17 labels in the config: AutoModelForTokenClassification loads them with no missing or unexpected keys, and its per-token logits match the ONNX to 1e-5. Calibrate them with temperature.json exactly as in the snippet above. The Apple (MLX) build is prepared from this pair.

On Apple Silicon the daemon runs this model through MLX (model_gpu_mlx{VER}/, prepared from the raw weights above), the only GPU lane there.

Apple Neural Engine (Core ML)

coreml_ane/embed.mlpackage/ is a Core ML multifunction package that runs this token classifier on the Apple Neural Engine. It is a bundle, not a file β€” four entries (Manifest.json, shapes.json, Data/com.apple.CoreML/model.mlmodel, Data/com.apple.CoreML/weights/weight.bin) that must keep their relative paths.

One compiled function per shape, named b<batch>_s<seq>: b32 s128. Fixed shapes are not a simplification. ct.EnumeratedShapes converts and runs, and measures 233 emb/s against 2 873 on the same encoder, because the dynamic ops it injects push the graph off the Neural Engine. shapes.json lists what was compiled, so a caller can ask instead of assuming.

Weights are fp16. The package is built from the MLX safetensors beside it, so a model is ANE-ready exactly when it is MLX-ready β€” there is no second set of source weights. Load it with MLComputeUnits.cpuAndNeuralEngine: plain .all lets Core ML place the graph on the GPU instead, which measured 809 emb/s against the ANE's 2 873.

A shape the package does not carry is not an error β€” the caller is expected to fall back to the MLX graph, which takes any shape. That is what makes the fixed-shape package safe to ship alongside model_gpu_mlx*/ rather than instead of it.

License

Released under the MIT license.

Backbone: nreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large β€” a re-upload of Microsoft's mMiniLMv2 (microsoft/unilm, MIT). The re-upload repository itself declares no license.

Neither the training corpus nor the human gold set (1 500 rows, evaluation-only) ships in this repo. Not legal advice.

Attribution

Backbone: nreimers/mMiniLMv2-L12-H384-distilled-from-XLMR-Large (itself distilled from XLM-RoBERTa-Large).

Downloads last month
97
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for faxenoff/code-daemon-ner-v1

Quantized
(3)
this model