vi-ner-videberta

ViDeBERTa token-classification model for Vietnamese bank-statement OCR. It extracts PERSON, ORGANIZATION, and ADDRESS; the address label is deliberately retained to disambiguate names inside administrative addresses (for example, a person-like street or ward name) from actual people and organizations.

Evaluation

Trainer token/window BIO metrics

These legacy Trainer summaries are computed over tokenized windows. They are reported separately from final decoded document-span exact-match metrics.

Metric Validation Test
precision 0.895407 -
recall 0.908239 -
f1 0.901777 -
accuracy 0.984443 -

Per-entity F1

Entity Validation Test
PERSON 0.891723 -
ORGANIZATION 0.889339 -
ADDRESS 0.922978 -

The training run did not record a public dashboard URL. Local TensorBoard event files are included under tensorboard/ when present.

Training data

Split sizes: test: 29,587, train: 331,468, validation: 31,809.

Source Rows
curated_ambiguity 54
curated_statement_patterns 10
llm_statement_synthetic 1,955
llm_statement_wave03 324
llm_statement_wave04 410
llm_statement_wave05 63
llm_statement_wave06 58
llm_statement_wave07 102
llm_statement_wave09 32
llm_statement_wave10 33
llm_statement_wave11 29
llm_statement_wave12 32
llm_statement_wave13 24
llm_statement_wave14 16
llm_statement_wave15 3
llm_statement_wave16 15
llm_statement_wave17 10
llm_statement_wave18 8
llm_statement_wave19 11
llm_statement_wave20 3
llm_statement_wave21 12
llm_statement_wave22 4
llm_statement_wave23 12
llm_statement_wave24 3
llm_statement_wave25 11
llm_statement_wave26 1
llm_statement_wave27 12
llm_statement_wave28 1
llm_statement_wave29 2
llm_statement_wave30 1
llm_statement_wave31 6
llm_statement_wave33 2
llm_statement_wave34 4
llm_statement_wave35 2
llm_statement_wave37 3
llm_statement_wave38 5
llm_statement_wave39 13
llm_statement_wave40 4
llm_statement_wave42 4
llm_statement_wave44 3
llm_statement_wave45 3
llm_statement_wave46 2
masothue_llm_wave03 107
masothue_llm_wave04 72
masothue_statement_augmented 34,998
masothue_synthetic 119,794
meddies_pii_vi 15,212
news_ner 2,973
pap_ner 34,103
phoner_covid19 9,955
private_statement_csv_llm 7,778
private_statement_ocr_llm 5,769
statement_hard_cases 344
vietnamnet_gold 8,396
vlsp2016 16,763
vlsp2021_ndtands 19,329
wikiann_vi 27,364

This checkpoint was trained on a mixed-source dataset that includes restricted or non-redistributable examples. The source rows are not uploaded with this model. Review every upstream license and usage condition before commercial use or redistribution.

Reproducibility

  • Base checkpoint: Fsoft-AIC/videberta-base
  • Base revision: d72e04d6e5065b2babb6bccad6f86960701f0572
  • Seed: 20260927
  • Labels: O, B-PERSON, I-PERSON, B-ORGANIZATION, I-ORGANIZATION, B-ADDRESS, I-ADDRESS

Intended use and limitations

The intended unit is one OCR-extracted bank-statement transaction line. OCR errors, abbreviations, unseen bank formats, and names that are also locations can reduce accuracy. Predictions should support downstream analysis, not replace review for legal, compliance, identity, or payment decisions.

Downloads last month
66
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dathuynh1108/vi-ner-videberta

Finetuned
(16)
this model