attempt_vanilla: a day of GLiNER PII fine-tuning on one MI300X

This repo is the cleaned-up record of an old experiment (24 Feb 2026). The goal was to continue fine-tuning two knowledgator GLiNER checkpoints in token-level mode on a ~1.03M-example synthetic PII corpus, using a single AMD MI300X (192 GB, ROCm 7.1).

It is published as an experiment log, not a production model. No F1 or precision/recall was computed during these runs; F1 for neighbouring checkpoints from a later sweep is under Evaluation. Benchmark the kept weights yourself before using them.

Related repos

Repo Role
arthrod/gliner_eval_folder Feb-2026 eval sweep (18 models Γ— 9 sets) that scored checkpoints 1000 and 8000 of this run (rows arthrod/attempt_vanilla:…)
arthrod/gliner_review_comparison The same sweep's predictions, one split per model
arthrod/v1-deberta_large-token_level-1-7 Standalone upload from the same day, formerly arthrod/gliner-pii-large-full-v1 (weights + ONNX)
arthrod/gliner-opf-ptbr-pii-v1, arthrod/gliner-opf-ptbr-pii-bench-v1 Later PT-BR PII models and benchmark
knowledgator/gliner-pii-large-v1.0 Base model

What's in the repo

Path What it is
gliner_logs/gliner-pii-large-full-v1/checkpoint-7500/ Main model. knowledgator/gliner-pii-large-v1.0 (DeBERTa-v3-large encoder) fine-tuned for 7,500 steps. Lowest eval loss of the run.
gliner_logs/gliner-pii-large-full-v1/checkpoint-8000/trainer_state.json Full log history of that run (it was stopped by hand at step ~8.2k of 15k).
gliner_logs/gliner-multitask-v1.0/checkpoint-1000/ Short 1,000-step top-up of knowledgator/gliner-multitask-v1.0 (DeBERTa-v2-xlarge encoder).
gliner_logs/h-gliner-token-full-v1/checkpoint-17000/ Config + trainer state only of a failed run (loss was exactly 0 for 17k steps, see below). Weights deliberately not kept.
configs/, backups/ Training YAMLs (vanilla_g = multitask, vanilla_h = DeBERTa-v3-small token run, vanilla_f backup = pii-large).
runs/ Raw stdout of the bring-up and main training launches.
wandb/ Offline W&B run directories (configs, summaries, output logs, .wandb files).
scripts/, bench/, data/ Training, evaluation, ONNX-conversion and benchmarking scripts used around these runs.

Optimizer states, RNG/scheduler states, intermediate checkpoints, and the vendored GLiNER / flash-attention build trees were removed before publishing.

Quick start

GLiNER 0.2.25's from_pretrained has no subfolder argument, so download the checkpoint folder first:

from huggingface_hub import snapshot_download
from gliner import GLiNER

path = snapshot_download(
    "arthrod/attempt_vanilla",
    allow_patterns=["gliner_logs/gliner-pii-large-full-v1/checkpoint-7500/*"],
)
model = GLiNER.from_pretrained(f"{path}/gliner_logs/gliner-pii-large-full-v1/checkpoint-7500")

text = "Contact John Smith at john.smith@example.com or +1 (555) 010-0199."
labels = ["person", "email", "phone number"]
for ent in model.predict_entities(text, labels, threshold=0.5):
    print(ent["text"], "->", ent["label"], round(ent["score"], 2))

Trained with gliner==0.2.25 and transformers==5.1.0.

Runs

All three runs used the same data: data/train.json (~1,026,320 samples) and data/eval.json, cut down to 50,000 samples after the first full eval (490k samples) proved far too slow. The data is in-house synthetic PII and is not included here.

1. gliner-pii-large-full-v1 (the main run)

  • Start: knowledgator/gliner-pii-large-v1.0, encoder microsoft/deberta-v3-large
  • span_mode: token_level, max_len: 768, max_width: 100, max_types: 30
  • lr encoder 5e-6, lr others 7e-6, linear schedule, 10% warmup, effective batch 184, bf16
  • Focal loss alpha=0.75, gamma=0, loss_reduction: sum (so loss values scale with batch size)
  • Planned 15,000 steps; stopped by hand at ~8,200 (1.43 epochs, ~6.5 h)
step 500 1000 2000 3000 4000 5000 6000 7000 7500 8000
eval loss 936.8 653.7 535.6 478.0 477.4 451.4 438.9 440.0 417.7 449.8

Train loss went from ~11,600 (step 10) to ~257 (step 8,000).

2. gliner-multitask-v1.0

  • Start: knowledgator/gliner-multitask-v1.0, encoder microsoft/deberta-v2-xlarge, max_len: 1024
  • 1,000 steps, batch 40, same LRs, ran concurrently on the same GPU as run 3
  • Train loss 4,015 β†’ 114; eval loss 167.7 (step 500) β†’ 149.2 (step 1000)

3. h-gliner-token-full-v1: the "loss = 0" trap

A from-scratch token-level GLiNER on microsoft/deberta-v3-small (configs/vanilla_h.yaml). Every logged train loss, eval loss and grad-norm was exactly 0.0 from step 10 to step 17,000. The main attempt burned about 2 h 47 m (17.5k steps, ~35 evals) before a DataLoader worker crash ended it. The root cause was never confirmed. Plausible suspects are prev_path: "none" passed as a string and a hard-coded class_token_index. Lesson: assert loss > 0 and grad_norm > 0 in the first few dozen steps.

Bring-up notes (from runs/ and wandb/)

Five launches failed in the first ~45 minutes before training ran:

  1. KeyError: -1 in get_negatives (data-format mismatch)
  2. UniEncoderTokenGLiNER has no gradient_checkpointing_enable
  3. onnxruntime missing
  4. ROCm torch/torchvision mismatch (operator torchvision::nms does not exist)
  5. "bf16/gpu not supported" (wrong virtualenv)

The final stack was Python 3.14, gliner 0.2.25, flashDeBERTa kernels and flash-attention 2.8.3 built for gfx9. A batch-size sweep (32 β†’ 128 β†’ 200 β†’ 184) settled on 184 at ~2.2–2.5 s/it; 200 slowed to ~9.9 s/it. When the concurrent multitask run finished, the remaining run's eval throughput jumped from ~290 to ~965 samples/s.

Evaluation (from the later sweep)

No task metric was computed during training. A later sweep, arthrod/gliner_eval_folder (summary config), scored two checkpoints of gliner-pii-large-full-v1 with exact-span F1 at threshold 0.5. These are checkpoints 1000 and 8000. Their weights are not kept here; checkpoint 7500, which is kept, was not in the sweep.

eval set (5,000 samples unless noted) ckpt-1000 F1 ckpt-8000 F1 arthrod/gliner-pii-large-full-v1 F1
arthrod__gliner_canonical_ptbr_pii_v1 0.691 0.730 0.730
nvidia__Nemotron_PII 0.601 0.616 0.627
arthrod__ptbr_pii_val_19k (19,164) 0.367 0.385 0.383
urchade__synthetic_pii_ner_mistral_v1 0.263 0.247 0.247
urchade__pile_mistral_v0.1 0.133 0.129 0.129
knowledgator__gliner_multilingual_synthetic 0.121 0.113 0.113
Ihor__gliner_post_train β€” 0.131 0.131
knowledgator__GLINER_multi_task_synthetic_data 0.096 0.081 0.081

The Hub model (now arthrod/v1-deberta_large-token_level-1-7) scores the same as ckpt-8000 on six of the eight shared sets (all but Nemotron-PII and the 19k val set), which suggests it is that checkpoint or one very close to it. Label sets differ per eval set, so compare within a row only.

Limitations

  • No task metrics during training. Eval loss is not comparable across runs (different batch sizes, max_types, sum reduction). scripts/eval_models_on_eval_json.py (exact-span F1 against nvidia/gliner-PII) was written but never run against these checkpoints. The only F1 numbers come from the later sweep above, which did not include the kept checkpoint 7500.
  • Trained on synthetic data only. Expect gaps on real-world documents, and don't rely on it as the sole safeguard for PII redaction.
  • English-centric encoders (DeBERTa-v3 / v2).

License

Apache-2.0, matching the knowledgator base models and the GLiNER library.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for arthrod/attempt_vanilla

Finetuned
(3)
this model