HAV classifier (xlm-roberta-large)

Binary classifier for social media posts: HAV (High Analytical Value) vs LAV (Low Analytical Value). Fine-tuned from FacebookAI/xlm-roberta-large.

A post is HAV if a researcher could quote it on its own as a genuine example of a person's feeling, opinion or experience about a product, brand, category or concept, and it is specific enough to stand alone. It is LAV if it isn't a genuine consumer voice (help/how-to, promotion or spam, noise, off-topic) or can't be used on its own (a conversation fragment, or a generic one-liner). Any sentiment can be HAV.

Texts longer than 256 tokens are truncated, as they were in training. With a threshold of 0.5 this is the same as taking the most likely label, so pipeline('text-classification', model=repo) gives the same labels.

Training data

Social media posts from social-listening exports, spanning several brands, product categories and platforms. Categories include: Beauty, Beverages, Tech, Clothing, Personal Finance, Fitness, etcPlatforms include: Instagram, Reddit, Twitter, Facebook, Youtube, TikTok, etc. 14,982 labelled posts (41% HAV).

Labels were produced by an LLM labeller (claude-sonnet-5-5) following a written rubric built around the question "Could a researcher quote this on its own as a real example of a person's feeling, opinion or experience about a topic?". The model therefore learns the labeller's judgement, including its mistakes. Training data was pre-cleaned from spam and posts too short or long.

Training procedure

Full fine-tuning of FacebookAI/xlm-roberta-large with a linear classification head and cross-entropy loss (Hugging Face Trainer), max sequence length 256, seed 42.

hyperparameter value
learning_rate 1.12e-05
num_epochs 3
batch_size 16
weight_decay 0.01
warmup_ratio 0.06
lr_scheduler linear
class_weight None
label_smoothing 0.0

Evaluation

5-fold cross-validation over all 14,982 labelled posts. Each post is scored by a fold model that did not train on it; numbers are the mean ± SD across the 5 folds. The released model is trained on all 14,982 posts with the same settings, so these numbers estimate its performance.

metric 0.500 (chosen) 0.873 (tuned)
HAV precision 0.802 ± 0.017 0.860 ± 0.013
HAV recall 0.843 ± 0.013 0.750 ± 0.004
HAV F1 0.822 ± 0.003 0.801 ± 0.004
HAV F0.5 0.809 ± 0.011 0.836 ± 0.009
LAV precision 0.885 ± 0.006 0.838 ± 0.001
LAV recall 0.852 ± 0.018 0.914 ± 0.009
accuracy 0.848 ± 0.006 0.846 ± 0.004
macro F1 0.845 ± 0.005 0.838 ± 0.004
ROC AUC 0.931 ± 0.004 0.931 ± 0.004
average precision 0.899 ± 0.006 0.899 ± 0.006
flagged HAV % 43.6 36.1

The threshold tuned on out-of-sample predictions (highest HAV precision with HAV recall ≥ 0.75) was 0.873. It was judged too conservative on real-world data, so 0.5 is published instead. Both are shown in the evaluation table.

Versions

The previous model in this repo is kept under the tag modernbert-base-v1. Load it with from_pretrained("sharecreative/high-analytical-value-v1", revision="modernbert-base-v1"). It uses its own threshold, as described in its own README at that tag.

Downloads last month
11
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sharecreative/high-analytical-value-v1

Finetuned
(1032)
this model