zlm-v1-iab-domain-classifier

A 149M-parameter ModernBERT model that classifies a web destination into IAB content and audience categories from its URL alone — no page fetch, no page text.

Code: github.com/zerogpu/zlm-v1-iab-domain-classifier · Benchmark dataset: ZeroGPU/iab-url-gold-784

Given nothing but a domain such as espn.com, the model returns standard IAB Tech Lab content and audience taxonomy categories. On a 784-URL benchmark labeled independently by GPT-5.5 and excluded from training, it is more accurate than GPT-5.4-nano on the same task while running ~50× faster (~36 ms vs ~1,900 ms median per URL).

Typical uses: contextual ad targeting, brand-safety screening, and audience enrichment in settings where only the URL is available — including dead, parked, or un-crawlable domains where there is no page to read.

The on-device build (int8, transformers.js layout) that ZeroGPU runs on edge devices in production is published as its own repository, ZeroGPU/zlm-v1-iab-domain-classifier-edge.

Model description

  • Backbone: Alibaba-NLP/gte-modernbert-base (ModernBERT, 149M parameters), fine-tuned end-to-end.
  • Heads: two multi-label MLP heads over the encoder — IAB Content (704 labels) and IAB Audience (1,567 labels).
  • Loss: Asymmetric Loss (ASL), sequence length 64 (the rendered URL string is short).
  • Calibration: per-head decision thresholds tuned on the 784-URL gold set (content default 0.60, audience default 0.50, with per-tier-1 overrides in thresholds.json; see Evaluation). Calibration lifted content precision from ~0.26 to ~0.42 and is a large part of the accuracy win.
  • Formats: fp32 ONNX (onnx/model.onnx), int8-quantized ONNX (onnx/model_quantized.onnx), the original PyTorch checkpoint (model.safetensors). The on-device build (per-channel int8, single concatenated output, transformers.js layout) is published separately; see On-device build.

Input format

The model is trained and served on a deterministic string rendering of the URL, not the raw URL. Reproduce it byte-for-byte at inference time:

"<host> | <segmented words> | tld:<tld>"

Examples:

https://www.sportsnews.co.uk/foo  ->  "sportsnews.co.uk | sports news | tld:co.uk"
espn.com                          ->  "espn.com | espn | tld:com"

The transformation: lowercase; strip scheme/www./port/path/query; split the host on ./-/digits; peel the TLD (public suffix) off as its own tld: token; word-segment glued labels (sportsnews → sports news, via wordninja) so the word-piece encoder sees real words. A reference implementation (url_text.py) is included in this repository.

Usage (Python, onnxruntime)

import json
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer
from url_text import url_to_text  # reference implementation in this repo

sess = ort.InferenceSession("onnx/model_quantized.onnx")
tok = Tokenizer.from_file("tokenizer.json")
meta = json.load(open("multi_head_classifier_metadata.json"))
thresholds = json.load(open("thresholds.json"))

text = url_to_text("https://www.espn.com")   # "espn.com | espn | tld:com"
enc = tok.encode(text)
outputs = sess.run(None, {
    "input_ids": np.array([enc.ids], dtype=np.int64),
    "attention_mask": np.array([enc.attention_mask], dtype=np.int64),
})

# One output per head (see meta["categories"] for order); apply sigmoid and the
# calibrated per-head threshold, then map indices to names via meta["label_mappings"].
def sigmoid(x):
    return 1.0 / (1.0 + np.exp(-x))

for head, logits in zip(meta["categories"], outputs):
    probs = sigmoid(logits[0])
    cut = thresholds[head]["_default"]
    labels = meta["label_mappings"]["category_labels"][head]
    picked = [(labels[i], float(p)) for i, p in enumerate(probs) if p >= cut]
    print(head, sorted(picked, key=lambda t: -t[1])[:6])

The same model runs in Node.js via onnxruntime-node (the tokenizer files are standard Hugging Face tokenizers JSON). For the best accuracy, load onnx/model.onnx (fp32) instead of the int8 file; see Quantized variants.

On-device build

The on-device build ZeroGPU serves to edge devices (browsers, Node.js / Docker workers, Android) in production is published as its own repository, ZeroGPU/zlm-v1-iab-domain-classifier-edge, with usage examples for onnxruntime and transformers.js. It differs from the files here in four ways:

  • One output instead of two. logits has shape [batch, 2271]: the 704 content logits followed by the 1,567 audience logits, with head-prefixed labels (content|Sports), so a standard text-classification pipeline can run it.
  • Per-channel int8 quantization (the int8 graph here is per-tensor); accuracy in Quantized variants.
  • 64-token input window, the length the model was trained with.
  • Head-level thresholds only: content 0.60, audience 0.50; the per-tier-1 overrides in thresholds.json are not applied on device.

Evaluation

Task: URL-only multi-label IAB content classification. Benchmark set: 784 held-out URLs the model never saw in training, spanning the full breadth of IAB tier-1 categories. Benchmark rows were excluded from the training corpus by both id and text. The set is published at ZeroGPU/iab-url-gold-784. The same 784 rows were also the validation set: they selected the best training epoch and the decision thresholds. The thresholded score below is therefore a mild upper bound on fully unseen data. Ground truth: every URL independently labeled by GPT-5.5 — a stronger model than either system under test, and not the producer of the training labels, so it does not favor the fine-tuned model. Baseline: GPT-5.4-nano prompted to do the same URL-to-IAB task. Metric: content micro-F1; both systems scored identically against the gold labels. The results below were measured on the server-side model; the on-device build's accuracy is in Quantized variants.

Headline result

This model (calibrated) GPT-5.4-nano
Content micro-F1 0.3845 0.3526
Precision 0.4247 0.3332
Recall 0.3512 0.3744
Median latency / URL ~36 ms ~1,900 ms
Off-taxonomy labels 0 (closed label set) possible (free-form)

The model wins by +0.032 F1 (≈ +9%), driven by a large precision gain at roughly equal recall: when it assigns a category, it is right more often.

Headline F1 Precision / recall / F1 Latency

Quantized variants

The headline numbers come from the fp32 model. The int8 files trade some recall for size; all variants still score above GPT-5.4-nano. Content micro-F1 at the 0.60 default threshold on the same 784 URLs, one URL at a time without padding:

File Size Precision Recall Content F1 Same top label as fp32
onnx/model.onnx (fp32) 605 MB 0.4243 0.3508 0.3841 —
onnx/model_quantized.onnx (int8, per-tensor) 152 MB 0.4362 0.3309 0.3763 73%
On-device build, ZeroGPU/zlm-v1-iab-domain-classifier-edge (int8, per-channel) 153 MB 0.4185 0.3242 0.3654 76%
GPT-5.4-nano — 0.3332 0.3744 0.3526 —

(The fp32 ONNX graph reproduces the PyTorch checkpoint's 0.3845 to within 0.0004.) Run the int8 files on single URLs or tightly padded batches: dynamic int8 quantization computes activation ranges over the whole input, padding included, and in padded batches of 64 both int8 files fell to about 0.35 F1.

Encoder ablation

The same fine-tune recipe was A/B-tested on two backbones (uncalibrated, F1@6 on the same 784-row gold set):

Backbone Gold F1@6 Precision Recall
all-MiniLM-L6-v2 0.2623 0.2066 0.3589
gte-modernbert-base 0.3332 0.2625 0.4561

MiniLM plateaued well below the nano baseline; ModernBERT cleared it once calibrated.

Encoder A/B

Per-category results (IAB tier-1)

The model beats GPT-5.4-nano in 18 of the 26 well-populated tier-1 categories, with the largest wins in commercially important verticals (Shopping, Technology & Computing, Business & Finance, Sports, Movies). It trails mostly in small, ambiguous, or naming-driven categories where a bare domain under-determines the topic.

Per-category delta

Full per-tier-1 table (784-row gold set)
Tier-1 category Rows Model F1 Nano F1 Δ
Pets 3 0.4706 0.1600 +0.3106
Healthy Living 10 0.4444 0.2529 +0.1916
Shopping 49 0.5525 0.4038 +0.1487
Unknown 8 0.2308 0.0976 +0.1332
Movies 11 0.7213 0.5902 +0.1311
Technology & Computing 62 0.5514 0.4228 +0.1286
Science 6 0.3333 0.2128 +0.1206
Television 9 0.5000 0.3939 +0.1061
Sports 28 0.4795 0.3756 +0.1039
Personal Finance 37 0.3012 0.2098 +0.0914
Business and Finance 108 0.3697 0.2861 +0.0836
Home & Garden 22 0.5000 0.4492 +0.0508
Travel 21 0.4189 0.3708 +0.0481
Careers 15 0.3670 0.3279 +0.0391
Hobbies & Interests 10 0.5172 0.4789 +0.0384
Video Gaming 19 0.3740 0.3357 +0.0383
Sensitive Topics 20 0.2500 0.2192 +0.0308
Education 54 0.4160 0.3883 +0.0277
Fine Art 7 0.4878 0.4762 +0.0116
Music and Audio 8 0.3333 0.3226 +0.0108
News and Politics 15 0.1600 0.1509 +0.0091
Automotive 37 0.3084 0.3493 −0.0409
Family and Relationships 9 0.2593 0.3030 −0.0438
Food & Drink 29 0.2588 0.3163 −0.0575
Style & Fashion 73 0.3629 0.4275 −0.0645
Books and Literature 9 0.4828 0.5714 −0.0887
Medical Health 38 0.2500 0.3410 −0.0910
Events and Attractions 50 0.1193 0.2787 −0.1594
Religion & Spirituality 13 0.3889 0.5610 −0.1721
Real Estate 4 0.1600 0.3529 −0.1929

Per-category F1

Training data

The model was trained on ~1 million domains whose IAB categories were labeled by a stronger teacher LLM that could see the live page (title and description) at labeling time. The model distills that page-aware knowledge into the URL string — at run time, from the URL alone, it reproduces judgments that normally require fetching the page. This privileged-information distillation is why it can exceed a same-class LLM that only sees the bare URL. The independently labeled 784-URL gold set was kept out of training and used for epoch selection, threshold calibration and the benchmark.

Limitations and intended use

  • Absolute F1 is modest (~0.38) because the task is hard: a bare domain often genuinely under-determines the topic. Treat outputs as probabilistic enrichment signals, not ground truth about a site.
  • Accuracy is not uniform across categories — see the per-category table. It is weakest on Events & Attractions, Religion & Spirituality, Real Estate, and Medical Health.
  • The model can only emit labels from the fixed IAB taxonomy snapshot it was trained on (704 content + 1,567 audience labels).
  • Trained predominantly on ASCII/English-tokenizable domains; non-Latin or punycode-heavy domains will see degraded quality.
  • The taxonomy includes sensitive categories (e.g. health, religion, adult content). Do not use the model to infer sensitive attributes of individual people; it is built for classifying web destinations for contextual advertising use cases, and predictions in sensitive categories should be reviewed before being acted on.

Changelog

  • 2026-09-30 — The on-device bundle moved to its own repository, ZeroGPU/zlm-v1-iab-domain-classifier-edge (same files, at the repository root), and edge/ was removed from this repository. URLs pinned to earlier commits of this repository keep working.

  • 2026-09-29 — Added edge/, the on-device bundle ZeroGPU serves in production (per-channel int8, concatenated logits output, transformers.js layout, 64-token window). Weights are unchanged, and every root file is untouched, so existing code keeps working. Card corrections: training used 64-token sequences (not 512); the 784-URL gold set also served as the validation set; added accuracy figures for the int8 files.

  • 2026-08-04 — Initial release.

License and attribution

Citation

@misc{zerogpu2026iabdomainclassifier,
  title  = {zlm-v1-iab-domain-classifier: URL-only IAB classification with a fine-tuned ModernBERT},
  author = {ZeroGPU},
  year   = {2026},
  url    = {https://huggingface.co/ZeroGPU/zlm-v1-iab-domain-classifier}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zgpu-ai/zlm-v1-iab-domain-classifier

Quantized
(14)
this model
Quantizations
1 model

Space using zgpu-ai/zlm-v1-iab-domain-classifier 1

Evaluation results

  • Content micro-F1 (calibrated) on 784-URL held-out gold benchmark (independent GPT-5.5 labels)
    self-reported
    0.385
  • Content micro-precision on 784-URL held-out gold benchmark (independent GPT-5.5 labels)
    self-reported
    0.425
  • Content micro-recall on 784-URL held-out gold benchmark (independent GPT-5.5 labels)
    self-reported
    0.351