|
Download README.md from rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem: direct link, hf CLI and curl.
- Browser
- Download file 4.27 kB
-
https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem/resolve/main/README.md
- Command line
-
hf download hf://rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem/README.md
-
curl -L -H "Authorization: Bearer $HF_TOKEN" -o README.md https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem/resolve/main/README.md
4.27 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| tags: | |
| - ai-text-detection | |
| - idea-provenance | |
| datasets: | |
| - rishanthrajendhran/WildOutlines | |
| extra_gated_prompt: "Access is granted individually. Please say who you are and what you intend to use the weights for." | |
| # IdeaLens-LogisticClassifier-PerItem | |
| IdeaLens-LogisticClassifier-PerItem is an idea-level detector: it judges **whose ideas a document contains**, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with [IdeaLens](https://huggingface.co/rishanthrajendhran/IdeaLens) and trained on the same data. | |
| | | | | |
| |---|---| | |
| | Model | logistic regression over OpenAI `text-embedding-3-large` embeddings | | |
| | Reads | each outline item, embedded on its own; the item scores are pooled to the document by their mean log-odds | | |
| | Training data | [WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines), train split | | |
| | Output | P(human); a document is flagged as AI when P(human) is below a cut | | |
| | Default cut | 0.20413 (global, 1% false-positive rate) | | |
| | Hardware | CPU only; embedding calls need `OPENAI_API_KEY` | | |
| ## Usage | |
| The [idealens](https://github.com/RishanthRajendhran/IdeaLens) package ([PyPI](https://pypi.org/project/idealens/)) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. | |
| ```bash | |
| pip install "idealens[openai]" | |
| idealens run docs.jsonl -o scores.jsonl --model IdeaLens-LogisticClassifier-PerItem | |
| ``` | |
| Input is JSONL with a `text` field per document. To score outlines you already have, use `idealens score outlines.jsonl -o scores.jsonl --model IdeaLens-LogisticClassifier-PerItem`. Score outlines as extracted; the paraphrasing step is only for training data. | |
| In Python, step by step: | |
| ```python | |
| import idealens as il | |
| texts = [open("document.txt").read()] | |
| formats = il.classify(texts) # one of the eight formats per document | |
| outlines = il.extract(texts, formats) # role-labelled outlines | |
| with il.Detector("IdeaLens-LogisticClassifier-PerItem") as det: | |
| records = det.score_outlines(outlines, format=formats) | |
| r = records[0] | |
| print(r["p_human"], r["verdict"]["ai"]) # P(human); flagged at the 1% global cut? | |
| print(r["item_p_human"]) # each item's own P(human) | |
| print(outlines[0].render()) # the outline that was scored | |
| ``` | |
| Or in one call: `records = il.run(texts, det)`. `classify` and `extract` use Gemini 3.7 Flash, the extractor the thresholds were fitted with (`GEMINI_API_KEY`); pass `provider=idealens.providers.make(...)` to use Vertex, OpenAI, Anthropic, OpenRouter or a local server. The [package README](https://github.com/RishanthRajendhran/IdeaLens#ways-to-use-idealens) covers the other ways to run it. | |
| ## Thresholds | |
| `thresholds.json` holds this model's cuts at 0.1%, 0.5%, 1%, 2% and 5% false-positive rates, fitted on the 80,000 human | |
| documents of WildOutlines' `calibration` split: one global cut per rate, plus per-format and per-topic cuts. The package applies them. A cut | |
| fitted for one model does not transfer to another model's scores. For documents unlike English web text, fit cuts on | |
| human documents from your own domain with `idealens calibrate`. | |
| ## Related | |
| - [IdeaLens](https://huggingface.co/rishanthrajendhran/IdeaLens): the main idea-level detector, with full documentation | |
| - [ProseLens](https://huggingface.co/rishanthrajendhran/ProseLens): its prose-level counterpart | |
| - [idealens](https://github.com/RishanthRajendhran/IdeaLens): the Python package that runs these models | |
| - [WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines): the training corpus | |
| - [IdeaLens collection](https://huggingface.co/collections/rishanthrajendhran/idealens-6abee785ce6196fc0be9200f): every model and dataset in one place | |
| ## Citation | |
| ```bibtex | |
| @article{idealens2026, | |
| title = {IdeaLens: Detecting AI Ideas in Long-form Writing}, | |
| author = {Anonymous}, | |
| journal = {arXiv preprint arXiv:TBD}, | |
| year = {2026}, | |
| url = {https://arxiv.org/abs/TBD} | |
| } | |
| ``` | |