Instructions to use ZeroGPU/zlm-v1-moderation-edge with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use ZeroGPU/zlm-v1-moderation-edge with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-classification', 'ZeroGPU/zlm-v1-moderation-edge');
zlm-v1-moderation-edge
A 139M-parameter DeBERTa moderation model that flags unsafe English text and scores it on the 13 categories of the OpenAI moderation taxonomy, calibrated and small enough to run on the device (int8 ONNX, 185 MB).
For each input the model returns a binary unsafe verdict plus a calibrated score for
each category: hate, hate/threatening, harassment, harassment/threatening,
self-harm, self-harm/intent, self-harm/instructions, sexual, sexual/minors,
violence, violence/graphic, illicit, illicit/violent. On a 4,277-text held-out
test split it reaches binary F1 0.899 vs 0.853 for OpenAI's omni-moderation (see
Evaluation for how the gold labels were made), and it runs locally with no
network call.
This is the exact bundle ZeroGPU runs on edge devices (browsers, Node.js / Docker workers, Android) in production. It is laid out for transformers.js and also runs directly in onnxruntime.
Model description
- Base model:
KoalaAI/Text-Moderation, a DeBERTa-base encoder (12 layers, hidden size 768, 50,265-token byte-level BPE vocabulary), fine-tuned end to end. - Heads: the encoder output is mean-pooled over the attention mask and feeds two
linear heads: a binary head (
unsafe) and a 13-way category head. The binary head is trained on every example; the category loss is applied only where category labels exist.flaggedcomes from the binary head, not from OR-ing the categories. - Parameters: 138.6M.
- Output: one
logitstensor of shape[batch, 14]: index 0 isunsafe, indices 1โ13 are the categories in the order listed above (postprocess.jsonโlabels). - Post-processing (in
postprocess.json, reproduced in the usage code below):- sigmoid on all 14 logits;
- clamp every sub-category to at most its parent (
hate/threateningโคhate, โฆ); - isotonic calibration per head (piecewise-linear curves), then clamp again;
- thresholds:
flagged = unsafe โฅ 0.4835, and a per-category cut for each category; - reconcile with
flagged: when not flagged, no category is set; when flagged and no category crosses its cut, the highest-scoring category is set; a set sub-category also sets its parent.
- Input window: 192 tokens, the length the model was evaluated at and is served with on devices in production (the architecture allows up to 512). In production it is served to devices with at least 4 GB of RAM.
- Quantization: int8 dynamic quantization (QUInt8, per-channel), except the feed-forward output projections of encoder layers 0โ5, which stay fp32. Early DeBERTa layers carry activation outliers that full int8 quantization turns into verdict flips, most visibly on very short inputs; keeping those six projections in fp32 (+42 MB) removes most of them.
Files
| File | Purpose |
|---|---|
onnx/model_quantized.onnx |
int8 graph (with the fp32 layers above), 185 MB. Inputs input_ids, attention_mask (int64); output logits [batch, 14] |
postprocess.json |
Labels, hierarchy, isotonic calibration curves, thresholds, max_input_length |
config.json |
DeBERTa config with the 14-label id2label, problem_type: multi_label_classification |
tokenizer.json, tokenizer_config.json |
DeBERTa byte-level BPE tokenizer (cased) |
tokenizer.json truncates at 192 tokens, the production input window. postprocess.json
also lists a looser max_input_length of 384; the examples below stay at 192, the tested
configuration. The repository contains no PyTorch weights.
Usage
The raw sigmoid scores are not the model's decision. Apply postprocess.json as below
to get the calibrated scores and verdicts ZeroGPU serves. Classify one text at a time
without padding: with dynamic int8 quantization, padding shifts the scores.
Python (onnxruntime)
import json
import numpy as np
import onnxruntime as ort
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer
root = snapshot_download("ZeroGPU/zlm-v1-moderation-edge")
post = json.load(open(f"{root}/postprocess.json", encoding="utf-8"))
tok = Tokenizer.from_file(f"{root}/tokenizer.json")
tok.no_padding()
tok.enable_truncation(max_length=192) # production input window
sess = ort.InferenceSession(f"{root}/onnx/model_quantized.onnx", providers=["CPUExecutionProvider"])
CATS, HIER, THR, CAL = post["categories"], post["hierarchy"], post["thresholds"], post["calibration"]
def clamp_to_parent(scores): # a sub-category never scores above its parent
for sub, parent in HIER.items():
scores[sub] = min(scores[sub], scores[parent])
def moderate(text):
ids = tok.encode(text).ids # one text at a time, unpadded
logits = sess.run(["logits"], {
"input_ids": np.array([ids], dtype=np.int64),
"attention_mask": np.ones((1, len(ids)), dtype=np.int64),
})[0][0]
p = 1.0 / (1.0 + np.exp(-logits.astype(np.float64))) # 1. sigmoid; p[0] = "unsafe", p[1:] = CATS
scores = dict(zip(CATS, map(float, p[1:])))
clamp_to_parent(scores) # 2. hierarchy clamp (raw)
unsafe = float(np.interp(p[0], CAL["binary"]["x"], CAL["binary"]["y"])) # 3. isotonic calibration
scores = {c: float(np.interp(v, CAL[c]["x"], CAL[c]["y"])) for c, v in scores.items()}
clamp_to_parent(scores) # ... and clamp again after calibration
flagged = unsafe >= THR["binary"] # 4. thresholds
categories = {c: flagged and scores[c] >= THR[c] for c in CATS}
if flagged and not any(categories.values()): # 5. flagged <=> at least one category
categories[max(CATS, key=scores.get)] = True
for sub, parent in HIER.items():
if categories[sub]:
categories[parent] = True
return {"flagged": flagged, "unsafe_score": unsafe, "categories": categories, "category_scores": scores}
for text in ["What's a good recipe for banana bread?",
"I will find where you live and make you regret ever talking to me."]:
r = moderate(text)
print(r["flagged"], round(r["unsafe_score"], 3), [c for c, on in r["categories"].items() if on])
Output:
False 0.225 []
True 1.0 ['harassment', 'harassment/threatening', 'illicit']
JavaScript (transformers.js v3 โ browser, Node.js, workers)
import { AutoModelForSequenceClassification, AutoTokenizer, Tensor } from '@huggingface/transformers';
const repo = 'ZeroGPU/zlm-v1-moderation-edge';
const post = await (await fetch(`https://huggingface.co/${repo}/resolve/main/postprocess.json`)).json();
const tokenizer = await AutoTokenizer.from_pretrained(repo);
const model = await AutoModelForSequenceClassification.from_pretrained(repo, { dtype: 'q8' });
const { categories: CATS, hierarchy: HIER, thresholds: THR, calibration: CAL } = post;
const interp = (v, { x, y }) => { // piecewise-linear, clamped at the ends (numpy.interp)
if (v <= x[0]) return y[0];
if (v >= x[x.length - 1]) return y[y.length - 1];
let i = 1; while (x[i] < v) i++;
return y[i - 1] + ((v - x[i - 1]) * (y[i] - y[i - 1])) / (x[i] - x[i - 1]);
};
const clampToParent = (s) => { for (const [sub, parent] of Object.entries(HIER)) s[sub] = Math.min(s[sub], s[parent]); };
// Cut to the 192-token production window and keep the closing [SEP]
// (transformers.js' own `truncation` option drops that last token, which shifts the scores).
function encode(text, maxLength = 192) {
let ids = Array.from(tokenizer(text).input_ids.data, Number);
if (ids.length > maxLength) ids = [...ids.slice(0, maxLength - 1), ids.at(-1)];
const tensor = (data) => new Tensor('int64', BigInt64Array.from(data, BigInt), [1, data.length]);
return { input_ids: tensor(ids), attention_mask: tensor(ids.map(() => 1)) };
}
async function moderate(text) {
const { logits } = await model(encode(text));
const p = Array.from(logits.data, (z) => 1 / (1 + Math.exp(-z))); // p[0] = "unsafe", p[1..] = CATS
const scores = Object.fromEntries(CATS.map((c, i) => [c, p[i + 1]]));
clampToParent(scores);
const unsafe = interp(p[0], CAL.binary);
for (const c of CATS) scores[c] = interp(scores[c], CAL[c]);
clampToParent(scores);
const flagged = unsafe >= THR.binary;
const categories = Object.fromEntries(CATS.map((c) => [c, flagged && scores[c] >= THR[c]]));
if (flagged && !CATS.some((c) => categories[c])) categories[CATS.reduce((a, b) => (scores[b] > scores[a] ? b : a))] = true;
for (const [sub, parent] of Object.entries(HIER)) if (categories[sub]) categories[parent] = true;
return { flagged, unsafe_score: unsafe, categories, category_scores: scores };
}
console.log(await moderate('I will find where you live and make you regret ever talking to me.'));
Evaluation
Test set: 4,277 held-out English texts (1,425 unsafe, 2,852 safe), drawn from the
same sources and labelling pipeline as the training data; every positive in the test set
is real text. Gold labels come mostly from the same LLM policy judge that labelled the
training data, so this benchmark measures agreement with that judge's policy on this
distribution: it is fair for "does this match the target policy better than
omni-moderation", not a neutral third-party referee. Baseline: OpenAI
omni-moderation scored at its documented 0.5 cut. The numbers below were measured on the
fp32 model; see int8 parity for the on-device file.
Binary safe / unsafe
| Model | Precision | Recall | F1 | AUC |
|---|---|---|---|---|
| zlm-v1-moderation-edge | 0.8878 | 0.9109 | 0.8992 | 0.9843 |
| OpenAI omni-moderation | 0.8817 | 0.8267 | 0.8533 | 0.9671 |
Per category
| Category | n | Precision | Recall | F1 | AUC | omni F1 | omni AUC | ฮ F1 |
|---|---|---|---|---|---|---|---|---|
| hate | 390 | 0.743 | 0.779 | 0.761 | 0.974 | 0.830 | 0.991 | โ0.069 |
| hate/threatening | 300 | 0.679 | 0.670 | 0.674 | 0.941 | 0.649 | 0.988 | +0.026 |
| harassment | 367 | 0.641 | 0.681 | 0.661 | 0.953 | 0.581 | 0.929 | +0.080 |
| harassment/threatening | 298 | 0.635 | 0.641 | 0.638 | 0.954 | 0.601 | 0.953 | +0.036 |
| self-harm | 369 | 0.934 | 0.805 | 0.865 | 0.981 | 0.763 | 0.993 | +0.101 |
| self-harm/intent | 299 | 0.773 | 0.843 | 0.806 | 0.978 | 0.810 | 0.991 | โ0.003 |
| self-harm/instructions | 308 | 0.887 | 0.763 | 0.820 | 0.977 | 0.690 | 0.989 | +0.130 |
| sexual | 373 | 0.744 | 0.777 | 0.760 | 0.977 | 0.794 | 0.989 | โ0.034 |
| sexual/minors | 304 | 0.651 | 0.355 | 0.460 | 0.891 | 0.670 | 0.995 | โ0.210 |
| violence | 612 | 0.725 | 0.843 | 0.779 | 0.968 | 0.753 | 0.963 | +0.026 |
| violence/graphic | 300 | 0.673 | 0.707 | 0.689 | 0.966 | 0.101 | 0.937 | +0.588 |
| illicit | 632 | 0.752 | 0.802 | 0.776 | 0.967 | 0.584 | 0.963 | +0.193 |
| illicit/violent | 517 | 0.824 | 0.758 | 0.790 | 0.974 | 0.531 | 0.963 | +0.258 |
The model has the higher F1 in 9 of 13 categories. omni-moderation has the higher AUC in 7 of 13: much of its F1 gap comes from its fixed 0.5 cut rather than from worse ranking. These per-category figures are raw threshold crossings (step 4 above, before reconciliation). Reconciliation only removes category positives on texts the binary head did not flag, so with it per-category precision is at or above these figures and recall at or below. The evaluation used a 192-token window.
sexual/minors is the weakest category (recall 0.355): its training positives are
verified synthetic examples only, while its test set is real text.
int8 parity
The on-device int8 file against the fp32 model, 225 inputs run one at a time (175 of them one- or two-word inputs), after full post-processing:
| Metric | Value |
|---|---|
flagged agreement |
0.973 |
| Per-category agreement | 0.982 |
| Mean |ฮ probability| | 0.007 |
flagged flips |
6 of 225 |
Five of the six flips are texts the fp32 model flags and int8 clears (for example the
single words book, car and cat); int8 flags one word fp32 does not (shoot).
Latency
int8 model, onnxruntime 1.30, one text, CPU (Intel Core Ultra 9 275HX), median of 50 runs:
| Threads | 16 tokens | 64 tokens | 128 tokens | 192 tokens |
|---|---|---|---|---|
| 1 | 17.6 ms | 55.5 ms | 114.8 ms | 194.9 ms |
| 4 | 10.4 ms | 22.0 ms | 39.5 ms | 83.2 ms |
Training
- About 72,000 English texts (71,829 rows): texts from public safety and toxicity datasets, labelled into the 13 categories by a teacher LLM acting as a policy judge (judging intent, so politely phrased harmful requests count), combined with the source datasets' own labels mapped onto the taxonomy.
- Rare sub-categories (
violence/graphic,hate/threatening,harassment/threatening,self-harm/intent,self-harm/instructions,sexual/minors) were topped up with LLM-generated examples, each independently verified by a second LLM call;sexual/minorsexamples are non-explicit. - The safe class was topped up with benign comments to two safe texts per unsafe one.
- A label is positive at judge score โฅ 0.5;
unsafeis positive when any category is. Sub-category labels imply their parent. - Loss: binary cross-entropy on the
unsafehead plus masked binary cross-entropy on the category head. Thresholds were chosen per head on the validation split (bootstrap median of the F1-optimal cut), and the isotonic calibration curves were fit on the same split; calibration is monotone, so it does not change any decision.
The training data is not published.
Limitations and intended use
- English only. Non-English input scores near the decision threshold regardless of content: in our checks a Spanish threat was not flagged and a benign German question was. Translate first.
- Very short inputs are unreliable. One- and two-word inputs land on a score plateau
just above the
unsafecut: 13 of 110 common benign single words (such asok,theandwalk) are flagged by this int8 file. Do not use it to moderate isolated words, or require a minimum length. - Benign texts can be flagged, most often as
self-harm(we have seen a half-marathon training plan, a baking recipe and a composting guide flagged) and asviolence/illicit(crime fiction). sexual/minorsrecall is low (0.355 on the test set); do not rely on this model alone for child-safety enforcement.- Benchmark gold labels come from the same judge as the training labels (see Evaluation), and the test set is in-distribution.
- Use it as one signal in a moderation pipeline, with human review for consequential decisions.
License and attribution
- License: CodeML OpenRAIL-M 0.1, inherited from the base model
KoalaAI/Text-Moderation. The terms, including use restrictions, are inLICENSE. - Category taxonomy follows the OpenAI moderation categories.
Citation
@misc{zerogpu2026moderationedge,
title = {zlm-v1-moderation-edge: calibrated on-device text moderation with a fine-tuned DeBERTa},
author = {ZeroGPU},
year = {2026},
url = {https://huggingface.co/ZeroGPU/zlm-v1-moderation-edge}
}
- Downloads last month
- 1,251
Model tree for ZeroGPU/zlm-v1-moderation-edge
Base model
KoalaAI/Text-ModerationEvaluation results
- Binary F1 on Held-out moderation test split, 4,277 English texts (LLM-judge gold)self-reported0.899
- Binary precision on Held-out moderation test split, 4,277 English texts (LLM-judge gold)self-reported0.888
- Binary recall on Held-out moderation test split, 4,277 English texts (LLM-judge gold)self-reported0.911
- Binary AUC on Held-out moderation test split, 4,277 English texts (LLM-judge gold)self-reported0.984