Model Card for BSC-EDU Regression Classifier

A multilingual text-embedding regression model for assigning a continuous educational-value score to web documents. The model is designed for quality-based data filtering, producing scores in the range [0, 4].

Model Details

Model Description

The BSC-EDU classifier is a lightweight multilingual regression model based on a text-embedding encoder. It is fine-tuned to predict a continuous educational-value score for text documents. The resulting score can be used to rank or filter documents according to their predicted educational value.

The model was trained using examples in Spanish, Catalan, and Basque, with 500,000 examples per language. The first 512 tokens from each example were used. The model's predictions are continuous rather than restricted to discrete classes, allowing flexible score thresholds for downstream data filtering.

  • Developed by: Barcelona Supercomputing Center (BSC)
  • Funded by: []
  • Shared by: []
  • Model type: Multilingual text-embedding regression model / text quality classifier
  • Language(s) (NLP): 74 languages supported by the base model.
  • License: Apache 2.0
  • Finetuned from model: snowflake-arctic-embed-l-v2.0 multilingual text-embedding model

Citation

BibTeX:

@inproceedings{bsc-edu,
  title = {BSC-EDU},
  note = {}
}
Downloads last month
79
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for langtech-languagemodeling/bsc-edu-annotator

Finetuned
(35)
this model