Text Classification
Transformers
Safetensors
English
modernbert
ai-text-detection
idea-provenance
text-embeddings-inference
Instructions to use rishanthrajendhran/IdeaLens-ModernBERT-L-RolesOnly with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rishanthrajendhran/IdeaLens-ModernBERT-L-RolesOnly with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="rishanthrajendhran/IdeaLens-ModernBERT-L-RolesOnly")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens-ModernBERT-L-RolesOnly") model = AutoModelForSequenceClassification.from_pretrained("rishanthrajendhran/IdeaLens-ModernBERT-L-RolesOnly", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 4,668 Bytes
29ca5d7 33f8de3 b1fa184 29ca5d7 b1fa184 29ca5d7 b1fa184 29ca5d7 b1fa184 29ca5d7 b1fa184 29ca5d7 b1fa184 29ca5d7 b1fa184 d0d1063 b1fa184 d0d1063 b1fa184 d0d1063 b1fa184 d0d1063 b1fa184 574d180 b1fa184 b4df431 b1fa184 33f8de3 b1fa184 d0f4e10 b1fa184 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 | ---
license: cc-by-nc-sa-4.0
base_model: answerdotai/ModernBERT-large
library_name: transformers
language:
- en
tags:
- ai-text-detection
- idea-provenance
datasets:
- rishanthrajendhran/WildOutlines
---
# IdeaLens-ModernBERT-L-RolesOnly
IdeaLens-ModernBERT-L-RolesOnly is an idea-level detector: it judges **whose ideas a document contains**, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with [IdeaLens](https://huggingface.co/rishanthrajendhran/IdeaLens) and trained on the same data.
| | |
|---|---|
| Model | ModernBERT-large with a sequence-classification head |
| Reads | only the outline's sequence of role labels, one `[Role]` per line, with the content removed |
| Training data | [WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines), train split |
| Output | P(human); a document is flagged as AI when P(human) is below a cut |
| Default cut | 0.05082 (global, 1% false-positive rate) |
| Hardware | any GPU; a CPU works for small jobs |
## Usage
The [idealens](https://github.com/RishanthRajendhran/IdeaLens) package ([PyPI](https://pypi.org/project/idealens/)) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo.
```bash
pip install "idealens[hf]"
idealens run docs.jsonl -o scores.jsonl --model IdeaLens-ModernBERT-L-RolesOnly
```
Input is JSONL with a `text` field per document. To score outlines you already have, use `idealens score outlines.jsonl -o scores.jsonl --model IdeaLens-ModernBERT-L-RolesOnly`. Score outlines as extracted; the paraphrasing step is only for training data.
In Python, step by step:
```python
import idealens as il
texts = [open("document.txt").read()]
formats = il.classify(texts) # one of the eight formats per document
outlines = il.extract(texts, formats) # role-labelled outlines
with il.Detector("IdeaLens-ModernBERT-L-RolesOnly") as det:
records = det.score_outlines(outlines, format=formats)
r = records[0]
print(r["p_human"], r["verdict"]["ai"]) # P(human); flagged at the 1% global cut?
print(outlines[0].render()) # the outline that was scored
```
Or in one call: `records = il.run(texts, det)`. `classify` and `extract` use Gemini 3.7 Flash, the extractor the thresholds were fitted with (`GEMINI_API_KEY`); pass `provider=idealens.providers.make(...)` to use Vertex, OpenAI, Anthropic, OpenRouter or a local server. The [package README](https://github.com/RishanthRajendhran/IdeaLens#ways-to-use-idealens) covers the other ways to run it.
## Thresholds
`thresholds.json` holds this model's cuts at 0.1%, 0.5%, 1%, 2%, 5%, 10% and 20% false-positive rates, fitted on the 80,000 human
documents of WildOutlines' `calibration` split: one global cut per rate, plus per-format and per-topic cuts. The package applies them. A cut
fitted for one model does not transfer to another model's scores. For documents unlike English web text, fit cuts on
human documents from your own domain with `idealens calibrate`.
## Related
- [IdeaLens](https://huggingface.co/rishanthrajendhran/IdeaLens): the main idea-level detector, with full documentation
- [ProseLens](https://huggingface.co/rishanthrajendhran/ProseLens): its prose-level counterpart
- [idealens](https://github.com/RishanthRajendhran/IdeaLens): the Python package that runs these models
- [WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines): the training corpus
- [IdeaLens collection](https://huggingface.co/collections/rishanthrajendhran/idealens-6abee785ce6196fc0be9200f): every model and dataset in one place
## License
[CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/): free to share and adapt for non-commercial
purposes, with attribution and under the same license. Built on [answerdotai/ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large) (Apache 2.0).
## Citation
```bibtex
@article{idealens2026,
title = {IdeaLens: Detecting AI Ideas in Long-form Writing},
author = {Rajendhran, Rishanth and Choi, Minjoon and Russell, Jenna and Namuduri, Ramya and B{\"o}l{\"o}ni-Turgut, Deniz and Karpinska, Marzena and Wieting, John and Iyyer, Mohit},
journal = {arXiv preprint arXiv:2610.06778},
year = {2026},
eprint = {2610.06778},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.06778}
}
```
|