Lynceus detection (EuroBERT-610m)
Version française : README.fr.md
A pre-filter, not a verdict. A low score does not mean a page is sound: it means the page has not been analysed. This model judges neither a page, nor a source, nor a person.
The pre-filter of Project Lynceus. It reads a French or English web page (up to 8,192 tokens) and tells whether it is worth sending to the full analysis, the one that spots rhetorical techniques, quotes the passages and explains them. It gives no verdict and says nothing about the reliability of a source.
Its role in Lynceus is described in the target architecture, §8. To run it without PyTorch, see "ONNX" below.
What it outputs
For a page: a score between 0 and 1, the probability that it uses at least one technique of the Lynceus taxonomy (31 techniques). Below the threshold, no full analysis.
What a tool using it must show below the threshold: "not analysed", never "no technique". The model misses 2 to 6 % of pages that do use techniques; showing an absence would be an implicit verdict. And the pre-filter must never block an analysis a reader asks for.
Only the detection head is published. The category and technique heads trained with it are not reliable enough to be distributed (category right 8 times out of 10, one technique out of two wrong).
How good it is
Measured on 117 pages read by a human and never seen in training (the banc of the Lynceus silver dataset).
| ROC AUC | |
|---|---|
| against the blind human reading (the independent figure) | 0.90 |
| against the corrected reference (human reading amended after seeing the annotator panel: it favours the panel) | 0.96 |
This bench is for development. It is not the Lynceus test set, which stays human, blind and separate.
As a pre-filter, at threshold 0.05: 94 to 98 % of pages using techniques are sent to analysis, for about half of all pages. In real browsing, where such pages are probably rarer than in this bench (40 %), the share of analyses avoided would be larger.
What it misses most often: a single technique in a text with no other signal (a missing source in a personal testimony), a rumour on a satirical site, a religious text making health claims. Serious-sounding pseudo-science, long its weak spot, is better detected since counter-examples were added (debunks and analyses on the same topics).
Intended use, and what not to do with it
- To decide whether a page should be analysed, in a Lynceus instance or a similar tool, before a more expensive call.
- Never as a verdict shown to a reader, nor as a score of a page.
- Never to judge, rank, filter or profile sources, outlets or people. Lynceus describes techniques in a page; a high score does not mean an outlet lies, and a low score does not mean it is reliable. Any moderation, censorship or blocklist use goes against the very purpose of the project.
How it was trained
- Data: 1,936 pages of the Lynceus silver dataset, annotated by a panel of three open-weight models (GLM-5.3, DeepSeek V4 Pro, MiniMax M3) following the Lynceus annotation guide. The 117 bench pages are held out.
- Model: EuroBERT-610m, mean pooling over tokens, three heads trained together (category, techniques, detection), 4 epochs, 8,192 tokens, learning rate 3e-5.
Loading it
The folder holds the encoder (encodeur/, with its tokenizer), the detection head (detection.pt) and the settings (config.json).
import json, torch
from transformers import AutoModel, AutoTokenizer
reglages = json.load(open("config.json"))
tokenizer = AutoTokenizer.from_pretrained("encodeur", trust_remote_code=True)
encodeur = AutoModel.from_pretrained("encodeur", trust_remote_code=True).eval()
detection = torch.nn.Linear(encodeur.config.hidden_size, 1)
detection.load_state_dict(torch.load("detection.pt"))
def score(texte: str) -> float:
e = tokenizer(texte, truncation=True, max_length=reglages["longueur"], return_tensors="pt")
with torch.no_grad():
etats = encodeur(**e).last_hidden_state
masque = e["attention_mask"].unsqueeze(-1).to(etats.dtype)
moyenne = (etats * masque).sum(1) / masque.sum(1)
return torch.sigmoid(detection(moyenne)).item()
The expected input is the Markdown extracted from the page (trafilatura, or Readability then Turndown), as in training. transformers must stay at version 4 for the EuroBERT code.
Do not fix the tokenizer. Recent transformers versions warn about an incorrect pattern in this tokenizer and suggest fix_mistral_regex=True. The model was trained without that fix, which changes how numbers are split in most pages: load the tokenizer as it is.
ONNX, without PyTorch
The onnx/ folder holds the same model as a single graph (encoder, mean pooling, head, sigmoid) in two forms, with their thresholds in config.json (key onnx). The tokenizer is the one in encodeur/.
detection-fp16.onnx(1.2 GB): exactly the quality of the PyTorch model. The form to prefer, and the one for a GPU or a browser (WebGPU).detection-int8.onnx(610 MB): three times faster on a CPU, but less precise, and its loss varies from one training run to another.
Measured on the 117 bench pages, against the corrected reference except for the "blind" column. The threshold is the highest one that keeps 94 % of the pages using techniques; precision is the share of pages sent to analysis that really use some (43 % without a pre-filter).
| File | Tokens read | ROC AUC | blind | Threshold | Precision | Pages sent | Time per page, 2 cores | 6 cores |
|---|---|---|---|---|---|---|---|---|
| fp16 | 512 | 0.939 | 0.916 | 0.051 | 0.85 | 47 % | ~3 s | ~1.1 s |
| fp16 | 1,024 | 0.951 | 0.911 | 0.133 | 0.84 | 48 % | ~7 s | ~2.4 s |
| fp16 | 2,048 | 0.956 | 0.913 | 0.338 | 0.89 | 45 % | ~15 s | ~6 s |
| int8 | 512 | 0.921 | 0.913 | 0.003 | 0.70 | 57 % | ~1.1 s | ~0.4 s |
| int8 | 1,024 | 0.941 | 0.899 | 0.017 | 0.73 | 55 % | ~2.7 s | ~1 s |
- Reading 512 tokens is almost enough: whether a page uses a technique is mostly decided in its opening. Beyond 2,048 tokens, CPU time explodes for no gain.
- Times are medians for a page that fills the length read, one page at a time, computation limited to 2 or 6 cores of a recent x86 CPU. They vary with the processor.
- Thresholds are tuned on these same 117 pages (50 using techniques, so one page is worth 2 points of recall): they are optimistic. Confirm them on other pages before relying on them.
import json
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("encodeur", trust_remote_code=True)
session = ort.InferenceSession("onnx/detection-fp16.onnx", providers=["CPUExecutionProvider"])
seuil = json.load(open("config.json"))["onnx"]["formes"]["detection-fp16.onnx"]["longueurs"]["512"]["seuil"]
def score(texte: str, longueur: int = 512) -> float:
e = tokenizer([texte], truncation=True, max_length=longueur, return_tensors="np")
sortie = session.run(["score"], {"input_ids": e["input_ids"].astype(np.int64),
"attention_mask": e["attention_mask"].astype(np.int64)})
return float(sortie[0][0])
to_analyse = score(text) >= seuil # otherwise: "not analysed", never "no technique"
Exported with kit/exporter_onnx.py from the lynx-corpus repository.
Limitations
- Trained on annotations written by language models: it inherits their shared biases.
- Evaluated against a single human reader, on 117 pages.
- French and English only; pages of 200 to 60,000 characters.
- The dataset composition follows a search plan, not a random sample of the web.
Reporting an error or a contrary use
An error, a misjudged page, a use contrary to this card: open a discussion on this repository, or an issue on Project-Lynceus.
Licence
Weights under Apache 2.0, like EuroBERT-610m they derive from. The training annotations are published under CC BY-SA 4.0; their sole rights holder, the author of this model, grants the Apache 2.0 licence to the weights derived from them. This will no longer hold if annotations from other contributors ever enter training.
- Downloads last month
- 7
Model tree for nashicloud/lynceus-detection
Base model
EuroBERT/EuroBERT-610m