toxicity-classifier / README.md
spamon's picture
Upload README.md with huggingface_hub
f0df930 verified
|
Raw History Blame Contribute Delete
1.23 kB
metadata
license: openrail++
base_model: textdetox/bert-multilingual-toxicity-classifier
pipeline_tag: text-classification
tags:
  - toxic
  - onnx
  - int8
language:
  - en
  - de
  - fr
  - es
  - it
  - ru
  - uk
  - ar
  - hi
  - ja
  - zh
  - he
  - am
  - tt

Ordicio toxicity classifier (ONNX, int8)

An int8 ONNX export of textdetox/bert-multilingual-toxicity-classifier for on-device inference in the Ordicio apps (ONNX Runtime on Android / iOS, onnxruntime-web in browsers).

  • manifest.json: the current version, file hashes, tokenizer settings, labels and threshold.
  • v1/model.onnx: weights quantized to int8 (dynamic quantization), inputs input_ids, attention_mask, token_type_ids as int32 (cast to int64 inside the graph), output logits (index 0 neutral, 1 toxic).
  • v1/vocab.txt: the WordPiece vocabulary of the base model (cased, no lowercasing).

The int8 model agrees with the fp32 model on 98.5 % of decisions (3600 samples of textdetox/multilingual_toxicity_dataset in 12 languages). For accuracy per language, see the base model's card: German results are weak (F1 0.52).

The use restrictions of the openrail++ license of the base model apply to this export.