CafeBERT — ViClickbait-2025

Fine-tuned from uitnlp/CafeBERT for binary Vietnamese clickbait detection.

Experimental setup

  • Input: headline paired with lead paragraph; no URL, source, category, publish time, image, or engagement metadata.
  • Fixed 80/10/10 split using StratifiedGroupKFold with seed 42.
  • Fine-tuning seeds: [42, 22, 202]; 3 epochs per seed.
  • Development Macro-F1 selects checkpoints and representative seed.
  • Weighted cross-entropy from training-label frequencies: True.
  • Effective batch size: 8; max length: 256.

Results

Metric Mean ± sample std
Test Macro-F1 0.8047 ± 0.0060
Test accuracy 0.8236 ± 0.0074
Dev Macro-F1 0.8256 ± 0.0053

Representative seed: 22, selected only by development Macro-F1.

Per-seed

seed dev_macro_f1 test_macro_f1 test_accuracy
22 0.8312 0.8043 0.8246
42 0.8249 0.7989 0.8158
202 0.8207 0.8108 0.8304

Labels

  • 0: non-clickbait
  • 1: clickbait

Usage

from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "BaoNhan/cafebert-ViClickbait-2025"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
title = "Tiêu đề bài báo"
lead = "Đoạn dẫn của bài báo"
inputs = tokenizer(title, lead, return_tensors="pt", truncation=True, max_length=256)
prediction = model(**inputs).logits.argmax(dim=-1).item()
print(model.config.id2label[prediction])

Dataset

Limitations

The dataset is small, temporally bounded to 2023–2025, and collected from eight Vietnamese news platforms. Results may not transfer to social media, other publishers, or emerging clickbait styles.

Downloads last month
4
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support