Model Card for BSC-EDU Regression Classifier
A multilingual text-embedding regression model for assigning a continuous educational-value score to web documents. The model is designed for quality-based data filtering, producing scores in the range [0, 4].
Model Details
Model Description
The BSC-EDU classifier is a lightweight multilingual regression model based on a text-embedding encoder. It is fine-tuned to predict a continuous educational-value score for text documents. The resulting score can be used to rank or filter documents according to their predicted educational value.
The model was trained using examples in Spanish, Catalan, and Basque, with 500,000 examples per language. The first 512 tokens from each example were used. The model's predictions are continuous rather than restricted to discrete classes, allowing flexible score thresholds for downstream data filtering.
- Developed by: Barcelona Supercomputing Center (BSC)
- Funded by: []
- Shared by: []
- Model type: Multilingual text-embedding regression model / text quality classifier
- Language(s) (NLP): 74 languages supported by the base model.
- License: Apache 2.0
- Finetuned from model: snowflake-arctic-embed-l-v2.0 multilingual text-embedding model
Citation
BibTeX:
@inproceedings{bsc-edu,
title = {BSC-EDU},
note = {}
}
- Downloads last month
- 79
Model tree for langtech-languagemodeling/bsc-edu-annotator
Base model
Snowflake/snowflake-arctic-embed-l-v2.0