|
Download README.md from msmth/Source-1: direct link, hf CLI and curl.
- Browser
- Download file 22.1 kB
-
https://huggingface.co/msmth/Source-1/resolve/main/README.md
- Command line
-
hf download hf://msmth/Source-1/README.md
-
curl -L -o README.md https://huggingface.co/msmth/Source-1/resolve/main/README.md
22.1 kB
| license: apache-2.0 | |
| base_model: jhu-clsp/mmBERT-base | |
| base_model_relation: finetune | |
| # library_name is left out on purpose: Source-1 runs only through the included source1.py, and a | |
| # "library_name: transformers" entry would offer AutoModel and pipeline snippets that load the backbone without | |
| # the trained heads and return meaningless scores. | |
| pipeline_tag: text-classification | |
| inference: false | |
| language: | |
| - en | |
| - ar | |
| - az | |
| - bg | |
| - bn | |
| - ca | |
| - cs | |
| - da | |
| - de | |
| - el | |
| - es | |
| - et | |
| - fa | |
| - fi | |
| - fil | |
| - fr | |
| - gu | |
| - he | |
| - hi | |
| - hr | |
| - hu | |
| - id | |
| - it | |
| - ja | |
| - ka | |
| - kk | |
| - kn | |
| - ko | |
| - lt | |
| - lv | |
| - ml | |
| - mr | |
| - ms | |
| - nl | |
| - "no" | |
| - pl | |
| - pt | |
| - ro | |
| - ru | |
| - sk | |
| - sl | |
| - sq | |
| - sr | |
| - sv | |
| - sw | |
| - ta | |
| - te | |
| - th | |
| - tr | |
| - uk | |
| - ur | |
| - vi | |
| - zh | |
| tags: | |
| - data-filtering | |
| - pretraining-data | |
| - quality-classifier | |
| - data-quality | |
| - data-curation | |
| - text-quality | |
| - multilingual | |
| - modernbert | |
| - mmbert | |
| datasets: | |
| - HuggingFaceFW/fineweb-2 | |
| - HuggingFaceFW/fineweb | |
| - HuggingFaceFW/finepdfs | |
| - HuggingFaceFW/finewiki | |
| - HPLT/HPLT2.0_cleaned | |
| - allenai/c4 | |
| - wikimedia/wikipedia | |
| - wikimedia/wikisource | |
| - HuggingFaceFW/fineweb-edu-llama3-annotations | |
| - HuggingFaceTB/finemath | |
| - open-web-math/open-web-math | |
| - codeparrot/github-code-clean | |
| - bigcode/commitpackft | |
| - HuggingFaceTB/cosmopedia | |
| - google/civil_comments | |
| - PleIAs/common_corpus | |
| - OpenAssistant/oasst2 | |
| - CohereLabs/aya_dataset | |
| - HuggingFaceFW/finetranslations | |
| - HuggingFaceTB/smollm-corpus | |
| - HuggingFaceTB/issues-kaggle-notebooks | |
| - allenai/WildChat-1M | |
| - codeparrot/github-code | |
| - bigcode/commitpack | |
| - OxAISH-AL-LLM/wiki_toxic | |
| - databricks/databricks-dolly-15k | |
| - Hello-SimpleAI/HC3 | |
| - ai4bharat/sangraha | |
| - liamdugan/raid | |
| - PleIAs/Post-OCR-Correction | |
| - joelniklaus/eurlex_resources | |
| - common-pile/arxiv_abstracts_filtered | |
| - common-pile/arxiv_papers_filtered | |
| - common-pile/biodiversity_heritage_library | |
| - common-pile/biodiversity_heritage_library_filtered | |
| - common-pile/caselaw_access_project | |
| - common-pile/data_provenance_initiative_filtered | |
| - common-pile/doab_filtered | |
| - common-pile/foodista_filtered | |
| - common-pile/github_archive | |
| - common-pile/libretexts_filtered | |
| - common-pile/library_of_congress | |
| - common-pile/library_of_congress_filtered | |
| - common-pile/news_filtered | |
| - common-pile/oercommons_filtered | |
| - common-pile/peS2o_filtered | |
| - common-pile/pre_1929_books | |
| - common-pile/pre_1929_books_filtered | |
| - common-pile/pressbooks_filtered | |
| - common-pile/project_gutenberg_filtered | |
| - common-pile/public_domain_review_filtered | |
| - common-pile/python_enhancement_proposals_filtered | |
| - common-pile/regulations_filtered | |
| - common-pile/stackexchange | |
| - common-pile/ubuntu_irc_filtered | |
| - common-pile/uk_hansard_filtered | |
| - common-pile/usgpo | |
| - common-pile/usgpo_filtered | |
| - common-pile/uspto_filtered | |
| - common-pile/wikiteam | |
| - common-pile/wikiteam_filtered | |
| # Source-1 | |
| Source-1 scores how useful a text is for training language models. It returns 13 scores and labels, such as | |
| educational value, spam, toxicity and topic, plus one overall score and a keep-or-drop decision. It reads 53 | |
| languages, up to 8,192 tokens at a time. It has 307M parameters and is free to use under Apache-2.0. | |
|  | |
| ## Highlights | |
| The numbers are rank agreement with an independent proprietary LLM grader: how closely a model puts texts in the same | |
| order as the grader does (1.0 means the same order, 0 means no link). The grader's grades were never trained on, and on | |
| the main test set it was given the same instructions as Source-1's teacher. | |
| - **Beats each of the 16 public quality scorers we tested, including FineWeb-Edu and propella-1, on every test set it | |
| was run on** (the 11 English-only scorers were not run on the 12-language exam). On the main test set (495 held-out | |
| texts in 53 languages): 0.90 vs 0.76 for the best of them, propella-1 4B, a model with 13x more parameters. | |
| - **Nearly matches its teacher**, the open-weight 27B LLM that labeled its training data, with 1/88 of its parameters: | |
| 0.90 vs 0.91 on the main test set. | |
| - **Also ahead when judged on educational value alone**, the thing most public scorers were built for. | |
| *Caveat: the grader scored with Source-1's own rubric (its scoring guide), so this test plays to Source-1's | |
| strengths. Even a scorer that matched the grader's educational-value scores exactly would reach only 0.87 here. See | |
| [Limitations](#limitations) and [EVALUATION.md](EVALUATION.md).* | |
| ## Quick start | |
| Source-1 runs through the included `source1.py`. Do not load it with transformers' `pipeline` or `AutoModel`: they | |
| skip the trained scoring heads and give meaningless scores. | |
| ```bash | |
| pip install -U huggingface_hub # provides the hf command | |
| hf download msmth/Source-1 --revision v1.0.0 --local-dir Source-1 --exclude "model.fp32.safetensors" | |
| cd Source-1 | |
| pip install -r requirements.txt | |
| python source1.py --model . --input examples/sample.jsonl --output scores.jsonl --device cpu # reproduces examples/expected_output.jsonl | |
| ``` | |
| ```python | |
| from source1 import Source1 | |
| model = Source1.from_pretrained(".") | |
| doc = model.score("Photosynthesis is how plants turn light, water and carbon dioxide into sugar and oxygen.") | |
| print(doc["overall"], doc["keep"]) # a 0-5 score and the keep/drop decision | |
| ``` | |
| ### For AI agents and scripts | |
| - **Load it only through `source1.py`.** `pipeline("text-classification", ...)` and `AutoModel...` classes load the | |
| backbone without Source-1's 13 trained heads and return meaningless `LABEL_0` / `LABEL_1` scores. | |
| - **Check the setup** with the `examples/sample.jsonl` command above: its output must equal | |
| `examples/expected_output.jsonl`. | |
| - **Output:** one JSON object per document (the 13 fields, `overall`, `keep`, `drop_reasons`). **Exit codes:** 0 done; | |
| 2 bad arguments or an unreadable input, before the model loads; 1 a bad record during a run. | |
| - No `trust_remote_code` and no prompts. Loading from a local folder makes no network calls. | |
| <details> | |
| <summary>More usage: many texts, precision, options, speed, files</summary> | |
| `source1.py` needs only `torch`, `transformers`, `safetensors` and `tokenizers` (no `trust_remote_code`). The model | |
| is published on the Hugging Face Hub as `msmth/Source-1`. | |
| ```python | |
| from source1 import Source1 | |
| model = Source1.from_pretrained(".") # a local directory or a Hub repo id; bfloat16 weights by default | |
| doc = model.score(open("article.txt", encoding="utf-8").read(), title="Optional title") | |
| print(doc["overall"], doc["keep"]) # 0-5 score and the keep/drop decision | |
| print(doc["educational_value"], doc["spam_seo"], doc["format"]) | |
| # Many documents at once: plain strings, or dicts with "text" and optionally "title". | |
| results = model.score_batch(["First document ...", {"text": "Second document ...", "title": "A title"}]) | |
| # The full-precision copy of the weights (model.fp32.safetensors), computing in float32: | |
| model_fp32 = Source1.from_pretrained(".", precision="fp32", dtype="fp32") | |
| ``` | |
| ```bash | |
| python source1.py --model . --input docs.jsonl --text-field text --output scores.jsonl --device cuda | |
| python source1.py --model . --input page.txt --device cpu --precision fp32 --dtype fp32 | |
| ``` | |
| - **Output.** One flat dict per document with the 13 fields, `overall`, `keep` and `drop_reasons`. It also gives the | |
| length in tokens, the number of chunks (`parts`), whether a chunk was cut, and per-chunk results (`chunks`). Empty | |
| text (or only spaces, zero-width or control characters) gives `keep` false, `drop_reasons` `["empty text"]` and | |
| `null` for `overall` and every field, so leave those records out before sorting by `overall`. | |
| - **Precision.** By default it loads `model.safetensors` (bfloat16, 0.6 GB). For the full float32 weights (1.2 GB), | |
| remove `--exclude` from the download and pass `precision="fp32"`. It computes in bfloat16 on NVIDIA Ampere or newer | |
| GPUs and in float32 elsewhere. float16 is not supported. In bfloat16, scores can shift by up to about 0.04 | |
| depending on which texts share a batch; pass `dtype="fp32"` or `batch_tokens=1` to avoid that. | |
| - **Options.** `drop_line` sets your own drop rule, such as `"toxicity >= 4 or spam_seo >= 3 or boilerplate >= 4.5"`. | |
| `max_chunks` scores only some chunks of very long documents, for speed. `revision` pins a Hub version (for example `revision="v1.0.0"`). The command | |
| line reads JSONL, JSON or plain text (also gzip, bzip2 or xz compressed), and `python source1.py --help` lists | |
| every flag. | |
| - **Examples and speed.** `examples/` holds six documents and their expected CPU output. A GPU can differ by a few | |
| hundredths. Source-1 scores about 30 chunks (50,000 tokens) per second on one RTX 3090 in bfloat16. | |
| - **On a CPU.** By default it scores 16,384 tokens per batch there and uses about 3 GB of RAM. On 4 threads it reads | |
| about 900 tokens per second on typical chunks and about 600 on full-length ones (about 12 seconds per | |
| 7,000-token chunk). Use `--max-chunks` for long documents. | |
| **Files** | |
| | file | contents | | |
| |---|---| | |
| | `model.safetensors` | the fine-tuned mmBERT-base backbone in bfloat16 (default) | | |
| | `model.fp32.safetensors` | the same backbone in float32 (load with `precision="fp32"`) | | |
| | `config.json` | backbone configuration (ModernBERT) | | |
| | `heads.safetensors` | the 13 scoring heads | | |
| | `source1.json` | rubric, head layout, pooling, maximum length and text normalization | | |
| | `calibration.json` | the drop line and the quality-score offsets, with how they were chosen | | |
| | `tokenizer.json`, `tokenizer_config.json` | mmBERT's tokenizer (same vocabulary and merges, re-saved) | | |
| | `source1.py` | standalone loader, Python API and command line | | |
| | `requirements.txt` | `torch`, `transformers`, `safetensors`, `tokenizers`, `huggingface_hub` | | |
| | `examples/` | six sample inputs (`sample.jsonl`) and their expected command-line output (`expected_output.jsonl`) | | |
| | `EVALUATION.md` | the full evaluation, the rubric anchors and the training details | | |
| | `images/` | the benchmark chart above | | |
| | `LICENSE`, `NOTICE`, `AUTHORS` | license text, third-party notices and credits, authors | | |
| | `CREDITS_BOOKS.tsv` | per-work credits for the open-books part of the training data, for books that are not public domain or CC0 (part of `NOTICE`) | | |
| </details> | |
| ## What you get | |
| | field | kind | what it measures | | |
| |---|---|---| | |
| | `format` | label | kind of text (10 types) | | |
| | `topic` | label | subject (15 topics) | | |
| | `content_type` | label | plain text, code or math | | |
| | `educational_value` | quality, 0-5 | teaches something useful | | |
| | `reasoning_depth` | quality, 0-5 | explains why, step by step | | |
| | `writing_quality` | quality, 0-5 | clear and well organized | | |
| | `information_density` | quality, 0-5 | real content, not filler | | |
| | `reliability` | quality, 0-5 | careful and trustworthy | | |
| | `spam_seo` | red flag, 0-5 | ads, SEO spam and scams | | |
| | `boilerplate` | red flag, 0-5 | menus, templates, link lists | | |
| | `toxicity` | red flag, 0-5 | hate, harassment, explicit content | | |
| | `code_quality` | code only, 0-5 | quality of the code | | |
| | `math_quality` | math only, 0-5 | quality of the math | | |
| Source-1 returns these 13 fields. Higher is better for quality scores and worse for red flags; the code and math | |
| scores are `null` when they do not apply. You also get: | |
| - `overall`: one 0-5 score, the quality scores minus penalties for red flags. | |
| - `keep`: false (drop) if `toxicity >= 4` or `spam_seo >= 3.5` or `boilerplate >= 4.5`. It checks only these red | |
| flags, so also set a threshold on `overall`. | |
| Long documents are split into chunks. Each chunk is scored, the results are combined into one, and `keep` is decided | |
| on the combined scores (each chunk's own result is in `chunks`). | |
| <details> | |
| <summary>All fields in detail, the scoring formula, the drop line and long documents</summary> | |
| **Label values.** | |
| - `format`: tutorial, reference, news, forum_qa, academic, fiction, code_file, product_page, blog_opinion, other. | |
| - `topic`: science, technology, programming, math, health, finance, history, politics_law, society, | |
| philosophy_religion, arts_entertainment, literature, sports, lifestyle, other. | |
| - `content_type`: plain_text, text_with_code, code_only, math_heavy. | |
| **Scores.** Each 0-5 score is a decimal such as 2.73: the model estimates how likely each level from 0 to 5 is, and | |
| the score is the average level weighted by those chances. `code_quality` applies when `content_type` is | |
| text_with_code or code_only, or `format` is code_file. `math_quality` applies when `content_type` is math_heavy or | |
| `topic` is math. Otherwise they are `null`. For a split document they average the chunks where they apply, so they can | |
| be set even when the document's `content_type` is plain_text. What each level means is in `source1.json` and | |
| [EVALUATION.md Appendix A](EVALUATION.md#appendix-a-rubric-anchors), with the extra rules that the teacher and the | |
| held-out grader were given. | |
| **Formula.** | |
| ``` | |
| quality = 0.30*educational_value + 0.20*reasoning_depth + 0.15*writing_quality | |
| + 0.20*information_density + 0.15*reliability | |
| if code_quality or math_quality applies: | |
| quality = 0.8*quality + 0.2*mean(the gated scores that apply) | |
| penalty = 0.5*max(0, spam_seo - 1) + 0.4*max(0, boilerplate - 1) + 0.8*max(0, toxicity - 1) | |
| overall = clip(quality - penalty, 0, 5) | |
| drop if toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5 (keep = false) | |
| ``` | |
| **The drop line.** The line in `calibration.json` is a bit stricter on spam than the rubric's own rule | |
| (`spam_seo >= 4`). Both keep ads and promotional pages, which the rubric scores `spam_seo` 3. The line was chosen by a | |
| fixed rule on the teacher's labels. Stricter lines and what they cost are in | |
| [EVALUATION.md](EVALUATION.md#the-drop-line). You do not have to use `keep`: ranking by `overall`, or your own rules | |
| on single fields, may work better for you. | |
| `keep` applies only this red-flag line, so very short or degenerate text (a single word, an emoji, one letter | |
| repeated) can still be kept. Combine it with a threshold on `overall`. A drop line can also use `overall`, `tokens` | |
| and `parts` (the last two for whole documents only), for example | |
| `"toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5 or overall < 1"`. | |
| **Long documents.** A document longer than about 7,800 tokens is split into balanced chunks at natural breaks. Each | |
| chunk is scored with a one-line header, as in training (source type, title if given, "Part i of n"). The document gets | |
| the label that covers the most tokens and the token-weighted average of each score. `toxicity` takes the maximum, and | |
| the code and math scores average only the chunks where they apply. `overall` and `keep` are then recomputed from these | |
| combined scores, so a document can be kept even when some of its chunks would be dropped. Each chunk's own scores and | |
| `keep` are in `chunks`. | |
| </details> | |
| ## Good for / Not for | |
| Good for: | |
| - Filtering, ranking, weighting and mixing pretraining text in its 53 languages, using any of its fields. | |
| - Checking a corpus: how much of it is spam, boilerplate, code, math or fiction. | |
| - Research on data quality and on teaching small models to copy LLM judgments. | |
| Not for: | |
| - Judging people, job applications or student work. | |
| - Fact-checking or content moderation: `reliability` judges care, not truth, and `toxicity` is a rough data filter. | |
| - Deciding whether text is licensed or legal to use. | |
| - Other languages, non-text input, or generating text. | |
| ## Limitations | |
| - **Home ground.** The grader used Source-1's own rubric. Models from the grader's family also helped shape that | |
| rubric, pick the teacher and tune the drop rule. The test texts come from the same kinds of sources as the training | |
| data. Whether filtering with Source-1 trains better models is untested. | |
| - **Two of the three test sets also helped pick the teacher**, by how well it agreed with the grader on them, so they | |
| may flatter Source-1. The main test set was built after that and was not used to pick the model. | |
| - **propella-1 has no single score.** We combine four of its ratings. If we also count its red-flag ratings, | |
| Source-1's lead shrinks by about half but remains ([details](EVALUATION.md#how-propella-1-is-read)). | |
| - **Scores run a bit high.** On English text it rates educational value, reasoning, writing and reliability about | |
| 0.3 points above the grader, as its teacher does. | |
| - **Its keep/drop rule catches about 2 of every 3 texts the grader drops.** A stricter rule catches more but also | |
| drops more good text. | |
| - **Weaker outside web text** (books, chats, code, synthetic text). The code and math scores are the least reliable. | |
| - **No reproduction kit.** The test data, the grader's labels and the metrics script are not included. | |
| More in [EVALUATION.md](EVALUATION.md#limitations). | |
| ## Training | |
| - Fine-tuned from [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base), with 13 small scoring heads added. | |
| - About 173,000 text chunks in 53 languages, each labeled by an open-weight 27B LLM teacher (no human labels). | |
| - License and safety filters were applied: see [NOTICE](NOTICE) and [EVALUATION.md](EVALUATION.md#appendix-b-training-in-detail). | |
| - **Development note.** AI assistants helped write the rubric and the code. Grades from proprietary LLMs, including | |
| the grader's model family, helped choose the rubric, the teacher and its setup, and the drop rule. LLM reviews also | |
| helped shape the license filters. None of these grades or reviews was ever used as a label or training example | |
| ([full account](EVALUATION.md#independence-from-development)). | |
| <details> | |
| <summary>Training details</summary> | |
| - **Base model.** [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) (ModernBERT architecture, 8,192-token | |
| context), all weights fine-tuned, with 13 linear heads on the mean-pooled hidden states: 307M parameters in total. | |
| - **Data.** 172,895 chunks (350M tokens, 151,281 documents) in 53 languages, 38.9% English: filtered and unfiltered | |
| web text, PDFs, wikis, math, permissively licensed code, conversations, synthetic text and openly licensed books, | |
| after license filtering and a safety filter. 9,744 chunks for validation and 9,553 for testing, split by document. | |
| - **Labels.** Every training label comes from an open-weight 27B LLM teacher scoring each chunk against the 13-field | |
| rubric. No human labels were used, and no output of a proprietary model was used as a label or training target. | |
| - **Recipe.** 2 epochs (5,320 steps of 131,072 tokens), AdamW, learning rate 5e-5 (heads 10x), bfloat16 mixed | |
| precision, about 1.9 hours on one H100 NVL. The learning rate and language mix came from a short sweep read on | |
| validation agreement with the teacher; the checkpoint is the final step, which also had the best validation score. | |
| Data sources, filtering, recipe and development details: | |
| [EVALUATION.md](EVALUATION.md#appendix-b-training-in-detail) and | |
| [Independence from development](EVALUATION.md#independence-from-development). | |
| </details> | |
| ## License and credits | |
| - Source-1: [Apache License 2.0](LICENSE). Copyright 2026 The Source-1 Authors (see `AUTHORS`). | |
| - Base model: [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) by the mmBERT authors, MIT License. | |
| - Training data credits and third-party notices: [NOTICE](NOTICE) and `CREDITS_BOOKS.tsv`. | |
| <details> | |
| <summary>License details</summary> | |
| [`NOTICE`](NOTICE) holds the full third-party notices and data credits. In short: | |
| - **Base model.** Fine-tuned from [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) by the mmBERT authors at | |
| Johns Hopkins University (Marone et al., 2025), MIT License; all encoder weights were further trained and 13 scoring | |
| heads added. mmBERT's tokenizer is based on the Gemma 2 tokenizer by Google. | |
| - **Labels.** Produced by an open-weight 27B LLM released under Apache-2.0, self-hosted; no teacher weights are here. | |
| - **Training data.** Public sources under their own terms, credited in `NOTICE` and `CREDITS_BOOKS.tsv`: among them | |
| data under the ODC Attribution License (FineWeb, FineWeb-2, FinePDFs, C4, FineMath and others; Common Crawl data | |
| through them was subject to the Common Crawl Terms of Use), Wikipedia-family text (CC BY-SA, GFDL or CC BY; with | |
| thanks to its volunteer editors), Common Pile v0.1, HPLT 2.0, EU publications, and Parliamentary information | |
| licensed under the Open Parliament Licence v3.0. Books include World Bank publications under CC BY 3.0 IGO; the World | |
| Bank and the other publishers do not endorse this model. Code is limited to permissive licenses. | |
| - **Share-alike text.** About 12.6% of training documents carry CC BY-SA or GFDL licenses. Source-1 is a classifier: | |
| it outputs scores, not text, and no training text is distributed with it. The weights are released under Apache-2.0 | |
| with attribution, as comparable quality classifiers are, on the view that such a scorer is not an adaptation of the | |
| text it was trained on. Copyleft code was removed anyway. | |
| - Upstream license metadata can be wrong. If you find a source that should not be here, please tell us (below). | |
| </details> | |
| ## Contact | |
| Questions, corrections and removal requests: open a discussion in the | |
| [Community tab](https://huggingface.co/msmth/Source-1/discussions). If your request involves personal information, open | |
| a discussion without the details and we will arrange a private way to reach us. We review every request and, where the | |
| content is in our training data, exclude it from future versions; published weights cannot be changed. | |
| ## Citation | |
| <details> | |
| <summary>BibTeX</summary> | |
| ```bibtex | |
| @misc{source1_2026, | |
| title = {Source-1: a multilingual 13-field scorer for pretraining data}, | |
| author = {{The Source-1 Authors}}, | |
| year = {2026}, | |
| howpublished = {\url{https://huggingface.co/msmth/Source-1}} | |
| } | |
| ``` | |
| Please also cite mmBERT: | |
| ```bibtex | |
| @misc{marone2025mmbertmodernmultilingualencoder, | |
| title = {mmBERT: A Modern Multilingual Encoder with Annealed Language Learning}, | |
| author = {Marc Marone and Orion Weller and William Fleshman and Eugene Yang and Dawn Lawrie and Benjamin Van Durme}, | |
| year = {2025}, | |
| eprint = {2509.06888}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CL}, | |
| url = {https://arxiv.org/abs/2509.06888} | |
| } | |
| ``` | |
| </details> | |