--- license: mit base_model: FacebookAI/xlm-roberta-large pipeline_tag: text-classification language: - multilingual tags: - news - topic-classification - multi-label - xlm-roberta library_name: transformers --- # SNT News Classifier v0.6 Multi-label news topic classifier: **12 top-level (L1) and 71 sub-level (L2) categories**, built on `xlm-roberta-large` with two independent sigmoid heads. Both levels are genuinely multi-label — an article about a trade deal can be `world` + `money_and_business` + `politics` at the same time. Per-class decision thresholds (tuned on a held-out validation split) ship inside `config.json`; `predict_labels()` applies them and falls back to argmax so no article is ever left unlabeled. Built by [Sweenk](https://sweenk.com) to categorize its news feed; released so others can use and scrutinize it. **Links:** [GitHub — model & training code](https://github.com/sweenk/snt-model) · [GitHub — data pipeline](https://github.com/sweenk/snt-data) ## At a glance - **Multi-label at both levels** — a story can be `world` + `money_and_business` + `politics` at once, each above its own per-class tuned threshold. - **0.847 L1 macro-F1** on 26,412 held-out articles; all 12 top-level categories clear a 0.65 per-class floor. - **Multilingual encoder** — `xlm-roberta-large` (100 languages), fine-tuned on ~264K news articles. - **Honest about its labels** — the gold is LLM-teacher-generated, not human-annotated, and the evaluation section says exactly what that means. ## Quick start ```python from transformers import AutoModel, AutoTokenizer model = AutoModel.from_pretrained("sweenk/snt-classifier", trust_remote_code=True) tok = AutoTokenizer.from_pretrained("sweenk/snt-classifier") enc = tok("OpenAI raises $6.6B. The startup announced its latest funding round...", return_tensors="pt", truncation=True, max_length=512) print(model.predict_labels(**enc)) # [{'l1': [{'key': 'money_and_business', 'p': 0.99}, {'key': 'tech_and_ai', 'p': 0.98}], # 'primary_l1': 'money_and_business', # 'l2': [{'key': 'companies_and_industries', 'p': 0.92}, ...]}] ``` Input convention: `"{title}\n\n{body}"`, truncated at 512 tokens. The classifier was trained on title+body; titles alone work but body text improves routing (the training prompt explicitly prioritizes body over headline). ## Taxonomy — 12 L1 / 71 L2 | L1 category | # L2 | L2 sub-categories | |---|---|---| | `sports` | 12 | `american_football`, `baseball`, `basketball`, `college_sports`, `combat_sports`, `golf`, `hockey`, `motorsports`, `olympics`, `other_sports`, `soccer`, `tennis` | | `politics` | 6 | `elections_and_campaigns`, `government_and_policy`, `immigration_and_borders`, `political_figures_and_scandals`, `social_issues_and_activism`, `state_and_local_politics` | | `world` | 4 | `geopolitics_and_diplomacy`, `humanitarian_crises`, `terrorism_and_security`, `war_and_conflict` | | `entertainment_and_pop_culture` | 7 | `books_and_arts`, `celebrities_and_gossip`, `gaming`, `internet_culture_and_creators`, `media_and_journalism`, `movies_and_tv`, `music` | | `money_and_business` | 8 | `companies_and_industries`, `cost_of_living`, `crypto_and_fintech`, `housing_and_real_estate`, `macro_economy_and_rates`, `markets_and_investing`, `personal_finance`, `work_and_careers` | | `crime_and_justice` | 4 | `courts_and_trials`, `crime_and_policing`, `scams_and_fraud`, `true_crime` | | `tech_and_ai` | 5 | `artificial_intelligence`, `big_tech_and_startups`, `cybersecurity_and_privacy`, `gadgets_and_apps`, `screen_time_and_digital_life` | | `science_and_space` | 4 | `archaeology_and_history`, `psychology_and_behavior`, `scientific_discoveries`, `space_and_astronomy` | | `health_and_wellness` | 5 | `fitness_and_exercise`, `medical_and_public_health`, `mental_health`, `nutrition_and_diet`, `sleep_and_longevity` | | `lifestyle` | 8 | `education_and_schools`, `faith_and_spirituality`, `fashion_and_beauty`, `food_and_drink`, `home_and_garden`, `parenting_and_family`, `relationships_and_dating`, `travel_and_places` | | `weather_and_environment` | 5 | `climate_change`, `disasters_and_accidents`, `energy_and_climate_solutions`, `nature_and_wildlife`, `severe_weather` | | `human_stories` | 3 | `animals_and_pets`, `good_news_and_kindness`, `offbeat_and_unusual` | The full machine-readable taxonomy (`l1_keys`, `l2_keys`, `l2_parent`, per-class thresholds) is in `config.json`. ## Evaluation — our own numbers, stated plainly Held-out test split: **26,412 articles** (a 10% slice of the labeled corpus, stratified by primary L1). "Tuned" = per-class thresholds optimized on the *validation* split, then applied unchanged to test. | Metric | @0.5 threshold | tuned thresholds | |---|---|---| | L1 macro F1 | 0.817 | **0.847** | | L1 micro F1 | 0.831 | **0.857** | | L1 primary accuracy (argmax) | 0.833 | — | | L2 macro F1 | 0.647 | 0.685 | | L2 micro F1 | 0.753 | 0.754 | ### Per-class L1 F1 (test) | L1 | F1 @0.5 | F1 tuned | threshold | |---|---|---|---| | `sports` | 0.941 | 0.954 | 0.850 | | `politics` | 0.845 | 0.859 | 0.750 | | `world` | 0.783 | 0.827 | 0.850 | | `entertainment_and_pop_culture` | 0.887 | 0.898 | 0.850 | | `money_and_business` | 0.786 | 0.825 | 0.950 | | `crime_and_justice` | 0.823 | 0.857 | 0.950 | | `tech_and_ai` | 0.802 | 0.857 | 0.950 | | `science_and_space` | 0.802 | 0.833 | 0.950 | | `health_and_wellness` | 0.821 | 0.860 | 0.900 | | `lifestyle` | 0.865 | 0.867 | 0.700 | | `weather_and_environment` | 0.816 | 0.857 | 0.950 | | `human_stories` | 0.632 | 0.671 | 0.900 | ### What you should know before trusting these numbers Read this section — it is the honest part. - **The gold labels are model-assisted, not human-annotated.** The corpus (~278K articles from HuffPost archives, CommonCrawl News, daily.dev, and Sweenk production) was labeled by a mechanical migration from an earlier taxonomy plus multiple passes of a Claude Sonnet teacher with a rule-based prompt, spot-audited by humans (QA gates at 70–87% agreement on sampled batches). Test F1 therefore measures agreement with an LLM teacher, not with human ground truth. - **Class imbalance is real (~21x).** `politics`/`lifestyle`/`entertainment` have ~44-47K training rows; `science_and_space` ~2.2K and `tech_and_ai` ~3.7K. Training compensates with per-class `pos_weight` (clamped at 10) and caps the majority classes at 25K primary-label train rows so they don't swamp the rare ones. The weakest class here is `human_stories` (F1 0.671) — the fuzziest category by construction. - **Labels were corrected over time, so F1 is not comparable across releases.** v0.5.1 re-labeled ~1,500 systematically mislabeled rows with route-by-cause rules (accidents by cause, terror → crime, pharma earnings → money); on those rows agreement with the corrected gold went 25.9% → 75.3%. Because the gold labels themselves changed between releases, aggregate F1 deltas are not like-for-like — treat each release's numbers as self-referential. - **Multilingual ability is inherited, not measured.** The encoder is XLM-R, but nearly all training articles are English. Expect degraded (unquantified) quality on non-English news. - **L3 (named topics / entities) is not part of this model** — Sweenk handles that downstream with a separate extraction step. ## Architecture `xlm-roberta-large` encoder → CLS pooling → dropout(0.1) → two parallel linear heads (L1: 12 logits, L2: 71 logits), both sigmoid. Trained 3 epochs, BCE loss with per-class `pos_weight` (L1) and loss weights L1:1.0 / L2:2.0, bf16 autocast, gradient checkpointing. Inference upcasts logits to fp32 before sigmoid (bf16 sigmoid saturates above logit ~6.2, which collapses co-confident multi-label pairs). ## Versions | Version | What changed | |---|---| | v0.5 | First multi-label release (12 L1 / 71 L2, dual sigmoid heads) | | v0.5.1 | Corrective retrain: route-accidents-by-cause, terror→crime, earnings→money, govt-personnel→politics, wildlife→weather, body-over-headline; ~1,500 corrected labels; re-tuned thresholds | | v0.6 | Retrain on the corrected corpus + ~23K newly teacher-labeled articles (Sweenk production + daily.dev science/tech); majority-class capping (25K/class) on top of `pos_weight`; re-tuned thresholds | ## License & attribution Model weights: MIT. Base model: [FacebookAI/xlm-roberta-large](https://huggingface.co/FacebookAI/xlm-roberta-large) (MIT). The training corpus contains article text from public news sources and is **not** redistributed with this model.