wikibert-ViFactCheck-GE

This model is TurkuNLP/wikibert-base-vi-cased fine-tuned for VFC-GE on ViFactCheck using the claim paired with gold evidence.

Evaluation protocol

  • Dataset size: 7,232 examples.
  • Shared fixed stratified splits for FC and GE: 5,785 train / 723 development / 724 test.
  • Labels: Supported, Refuted, and Not Enough Information.
  • Fine-tuning seeds: [42, 22, 202].
  • Training: 3 epoch(s), AdamW, learning rate 2e-05, weight decay 0.01, warmup ratio 0.1.
  • Effective train batch size: 8 (hard-validated against every published run).
  • Maximum sequence length: 256.
  • Input mode: raw Vietnamese claim and passage.
  • The claim is always preserved; only the second sequence (gold evidence) is truncated when the pair exceeds the encoder limit.
  • Topic, author, outlet, URL and other source metadata are excluded from model inputs.
  • No class weighting, resampling, retrieval model, sentence ranking, test-time model selection or external evidence is used.
  • Checkpoints are selected by development Macro-F1. The representative published checkpoint is seed 202, selected only by development Macro-F1.

Results

Test metrics are reported as mean ± sample standard deviation over seeds [42, 22, 202].

Metric Mean ± std
Test Macro-F1 0.7859 ± 0.0078
Test accuracy 0.7864 ± 0.0084
Test macro precision 0.7862 ± 0.0073
Test macro recall 0.7870 ± 0.0084
Development Macro-F1 0.7795 ± 0.0127

Per-seed results

seed dev_macro_f1 test_macro_f1 test_accuracy micro_batch_size gradient_accumulation_steps
22.000000 0.783415 0.794564 0.795580 8.000000 1.000000
42.000000 0.765365 0.779342 0.779006 8.000000 1.000000
202.000000 0.789837 0.783924 0.784530 8.000000 1.000000

Label mapping

{
  "0": "supported",
  "1": "refuted",
  "2": "not_enough_information"
}

Usage

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "BaoNhan/wikibert-ViFactCheck-GE"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=False)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

claim = "Thông tin này đã được cơ quan chức năng xác nhận."
evidence = "Bài báo cung cấp bằng chứng liên quan đến phát biểu trên."
inputs = tokenizer(
    claim,
    evidence,
    return_tensors="pt",
    truncation="only_second",
    max_length=256,
)
with torch.no_grad():
    probabilities = model(**inputs).logits.softmax(dim=-1)[0]
predicted_id = int(probabilities.argmax())
print(model.config.id2label[predicted_id], probabilities.tolist())

Files

  • aggregate_metrics.json: aggregate metrics and training manifest.
  • artifacts/per_seed_results.csv: one row per fine-tuning seed.
  • artifacts/seed_*_confusion_matrix.csv: confusion matrix for each seed.
  • artifacts/seed_*_classification_report.json: per-class metrics.
  • artifacts/seed_*_test_predictions.csv: IDs, gold/predicted labels and probabilities; raw claims and passages are excluded.

Limitations

ViFactCheck supplies the correct source article and therefore does not evaluate open-web evidence retrieval. VFC-FC can truncate relevant information in long articles and jointly measures verification plus robustness to irrelevant context. VFC-GE uses oracle gold evidence and must not be presented as a realistic end-to-end deployment setting. This model is a research classifier, not an automated arbiter of truth, and may produce confidently incorrect predictions.

Dataset citation

@inproceedings{hoa2025vifactcheck,
  title={ViFactCheck: A New Benchmark Dataset and Methods for Multi-domain News Fact-Checking in Vietnamese},
  author={Hoa, Tran Thai and Duy, Tran Quang and Tran, Khanh Quoc and Nguyen, Kiet Van},
  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
  volume={39},
  number={1},
  pages={308--316},
  year={2025},
  doi={10.1609/aaai.v39i1.32008}
}
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BaoNhan/wikibert-ViFactCheck-GE

Finetuned
(5)
this model

Dataset used to train BaoNhan/wikibert-ViFactCheck-GE