Text Generation
Transformers
Safetensors
English
nemotron_h
ai-text-detection
idea-provenance
conversational
Instructions to use rishanthrajendhran/IdeaLens-NoParaphrase with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rishanthrajendhran/IdeaLens-NoParaphrase with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="rishanthrajendhran/IdeaLens-NoParaphrase") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens-NoParaphrase") model = AutoModelForCausalLM.from_pretrained("rishanthrajendhran/IdeaLens-NoParaphrase", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use rishanthrajendhran/IdeaLens-NoParaphrase with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "rishanthrajendhran/IdeaLens-NoParaphrase" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/IdeaLens-NoParaphrase", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/rishanthrajendhran/IdeaLens-NoParaphrase
- SGLang
How to use rishanthrajendhran/IdeaLens-NoParaphrase with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "rishanthrajendhran/IdeaLens-NoParaphrase" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/IdeaLens-NoParaphrase", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "rishanthrajendhran/IdeaLens-NoParaphrase" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rishanthrajendhran/IdeaLens-NoParaphrase", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use rishanthrajendhran/IdeaLens-NoParaphrase with Docker Model Runner:
docker model run hf.co/rishanthrajendhran/IdeaLens-NoParaphrase
|
Download README.md from rishanthrajendhran/IdeaLens-NoParaphrase: direct link, hf CLI and curl.
- Browser
- Download file 15.8 kB
-
https://huggingface.co/rishanthrajendhran/IdeaLens-NoParaphrase/resolve/main/README.md
- Command line
-
hf download hf://rishanthrajendhran/IdeaLens-NoParaphrase/README.md
-
curl -L -o README.md https://huggingface.co/rishanthrajendhran/IdeaLens-NoParaphrase/resolve/main/README.md
15.8 kB
| license: other | |
| license_name: openmdw-1.1 | |
| license_link: LICENSE | |
| base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | |
| library_name: transformers | |
| language: | |
| - en | |
| tags: | |
| - ai-text-detection | |
| - idea-provenance | |
| datasets: | |
| - rishanthrajendhran/WildOutlines | |
| # IdeaLens-NoParaphrase | |
| IdeaLens-NoParaphrase is an ablation of IdeaLens: the same backbone, documents and labels, trained on the outlines | |
| as extracted, without the paraphrasing step that removes the documents' wording. It reads a role-labelled outline and | |
| returns P(human), the probability that the document's ideas are human. | |
| It is `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` fine-tuned with LoRA (rank 64) on the outlines of 1M English | |
| web documents ([WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines), `outline` field). | |
| **Try it in your browser:** the [IdeaLens & ProseLens demo](http://ideadetector.ai/) scores your own text with both detectors, | |
| no installation or keys needed. | |
| ## Results | |
| This model is not reported in the paper. It is released for comparison with IdeaLens. | |
| ## Usage | |
| Scoring a document takes two steps: | |
| 1. **Extract an outline.** An LLM writes the outline from the document, its format's role vocabulary and six worked | |
| examples (the paper uses Gemini 3.7 Flash). The [idealens](https://github.com/RishanthRajendhran/IdeaLens) package ships the prompts, role vocabularies and | |
| worked examples, and runs this step. Score the outline as extracted; the paraphrasing step is only for training data. | |
| 2. **Score the outline** with this model, as below. | |
| ### Quick start: everything with the idealens package | |
| [idealens](https://github.com/RishanthRajendhran/IdeaLens) ([PyPI](https://pypi.org/project/idealens/)) classifies each document's format, extracts its outline with the prompt, role | |
| vocabulary and worked examples IdeaLens-NoParaphrase was trained with, scores the outline on vLLM and applies the thresholds in this | |
| repo: | |
| ```bash | |
| pip install "idealens[vllm]" | |
| export GEMINI_API_KEY=... # or --provider vertex | openai | anthropic | openrouter | compatible | |
| idealens run docs.jsonl -o scores.jsonl --model IdeaLens-NoParaphrase --dry-run # price the LLM calls first | |
| idealens run docs.jsonl -o scores.jsonl --model IdeaLens-NoParaphrase | |
| ``` | |
| Input is JSONL with a `text` field per document. The same steps in Python: | |
| ```python | |
| import idealens as il | |
| texts = [open("document.txt").read()] | |
| formats = il.classify(texts) # one of the eight formats per document | |
| outlines = il.extract(texts, formats) # role-labelled outlines | |
| with il.Detector("IdeaLens-NoParaphrase") as det: # vLLM, with this repo's thresholds | |
| records = det.score_outlines(outlines, format=formats) | |
| r = records[0] | |
| print(r["p_human"], r["verdict"]["ai"], r["verdict"]["cut"]) # P(human); flagged at the 1% global cut? | |
| print(outlines[0].render()) # the outline that was scored | |
| ``` | |
| `classify` and `extract` call Gemini 3.7 Flash, the extractor the thresholds were fitted with; pass | |
| `provider=idealens.providers.make("openai", "gpt-6-sol")` (or Vertex, Anthropic, OpenRouter, a local server) to use | |
| another. Each record also carries verdicts at every calibrated false-positive rate under the global, per-format and | |
| per-topic schemes. The [package README](https://github.com/RishanthRajendhran/IdeaLens#ways-to-use-idealens) covers batch jobs, scoring outlines you already have and calibrating | |
| on your own data. | |
| ### Mix and match: extract with the package, score with your own code | |
| ```python | |
| import idealens as il | |
| text = open("document.txt").read() | |
| outline = il.extract([text], il.classify([text]))[0].render() # one "[Role] content" line per item | |
| print(p_human(outline)) # p_human (transformers) or p_human_batch (vLLM), defined below | |
| ``` | |
| The rest of this section runs the model directly. | |
| ### Load the merged model (66 GB download) | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| tok = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens-NoParaphrase") | |
| model = AutoModelForCausalLM.from_pretrained("rishanthrajendhran/IdeaLens-NoParaphrase", dtype=torch.bfloat16, device_map="auto").eval() | |
| ``` | |
| The weights take 59 GiB of GPU memory, and each input adds more; see *Hardware requirements*. | |
| ### Or apply the adapter to the base model (3 GB download) | |
| `adapter/` holds the LoRA adapter as trained, in the layout of the Tinker training service. If you already have the | |
| base model, `load_adapter.py` merges the adapter into it in memory. The resulting weights are bit-identical to the | |
| merged model's: | |
| ```python | |
| import importlib.util | |
| from huggingface_hub import hf_hub_download | |
| path = hf_hub_download("rishanthrajendhran/IdeaLens-NoParaphrase", "load_adapter.py") | |
| spec = importlib.util.spec_from_file_location("load_adapter", path) | |
| la = importlib.util.module_from_spec(spec); spec.loader.exec_module(la) | |
| model, tok = la.load_model() # base model + adapter/, then la.p_human(model, tok, outline) | |
| ``` | |
| Do not load `adapter/` with `peft.PeftModel`. In transformers, Nemotron fuses the Mamba gate and x projections into | |
| one `in_proj` and stores each layer's 128 routed experts as a single 3D tensor, so PEFT has nowhere to attach most of | |
| the adapter and skips it without a warning; the model then scores close to the base model. `tinker-cookbook`'s | |
| `weights.build_hf_model` can also merge the adapter into full weights. | |
| ### Score an outline | |
| Write the outline one item per line, as `[Role] content`, or take it from `il.extract` as above. IdeaLens-NoParaphrase compares the next-token probabilities of `human` and `ai`: | |
| ```python | |
| SYSTEM = "Given a role-labelled outline of a document, answer with one word: human if the source document was human-written, ai if it was AI-generated." | |
| SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n" | |
| HUMAN, AI = 50755, 2464 # token ids of "human" and "ai" | |
| @torch.no_grad() | |
| def p_human(outline): | |
| ids = tok.encode(f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{outline}{SUFFIX}", | |
| add_special_tokens=False) | |
| logits = model(torch.tensor([ids], device=model.device)).logits[0, -1].float() | |
| return torch.softmax(logits[[HUMAN, AI]], -1)[0].item() | |
| outline = ("[Central Development] A town's water supply fails after a drought, and residents organise to share wells.\n" | |
| "[Background Context] The reservoir has been shrinking for three summers.\n" | |
| "[Open Question] Whether the council will fund a new pipeline remains undecided.") | |
| print(p_human(outline)) | |
| ``` | |
| Build the prompt string exactly as above rather than through the chat template. | |
| ### Score with vLLM | |
| For many inputs, vLLM is about 15 times faster than the code above and fits much longer inputs on one 80 GB GPU. | |
| With vLLM 0.21 (install `xgrammar==0.2.1`; later releases require transformers < 5), reusing `SYSTEM`, `SUFFIX`, | |
| `HUMAN` and `AI` from above: | |
| ```python | |
| import math, os | |
| os.environ.setdefault("VLLM_USE_FLASHINFER_SAMPLER", "0") # FlashInfer kernels compile CUDA code and need nvcc | |
| os.environ.setdefault("VLLM_USE_FLASHINFER_MOE_FP16", "0") | |
| os.environ.setdefault("VLLM_USE_DEEP_GEMM", "0") # H100 warmup crashes when DeepGEMM is not installed | |
| from transformers import AutoTokenizer | |
| from vllm import LLM, SamplingParams | |
| tok = AutoTokenizer.from_pretrained("rishanthrajendhran/IdeaLens-NoParaphrase") | |
| llm = LLM(model="rishanthrajendhran/IdeaLens-NoParaphrase", dtype="bfloat16", max_num_seqs=256, enable_prefix_caching=False, | |
| max_logprobs=20, enable_flashinfer_autotune=False, seed=0) | |
| sp = SamplingParams(max_tokens=1, temperature=0.0, logprobs=20) | |
| def p_human_batch(outlines): | |
| prompts = [{"prompt_token_ids": tok.encode(f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n{x}{SUFFIX}", | |
| add_special_tokens=False)} for x in outlines] | |
| out = [] | |
| for r in llm.generate(prompts, sp, use_tqdm=False): | |
| lp = r.outputs[0].logprobs[0] # the top 20 next-token log-probabilities | |
| out.append(1 / (1 + math.exp(lp[AI].logprob - lp[HUMAN].logprob))) | |
| return out | |
| # when done: without this, vLLM 0.21 keeps a script running after its last line | |
| llm.llm_engine.engine_core.shutdown() | |
| ``` | |
| `max_num_seqs=256` keeps every running sequence's Mamba state in memory; vLLM's H100 default (1,024) does not fit | |
| beside the weights. If `human` or `ai` is missing from the top 20 (rare), score the prompt followed by each label | |
| token with `SamplingParams(max_tokens=1, prompt_logprobs=0)` and read the last prompt log-probability of each. | |
| Scores agree with the training-time scores to about 0.001 in P(human) on average; A100 and H100 GPUs differ by as much. | |
| ### Thresholds | |
| IdeaLens-NoParaphrase flags a document as having AI ideas when P(human) is below a cut. Each cut is set so | |
| that a given share of human documents is flagged (the false-positive rate, FPR), measured on the 80,000 human | |
| documents in WildOutlines's `calibration` split (10,000 per format). The paper's operating point is the global | |
| cut at 1% FPR. These cuts were fitted on scores from the merged weights in this repo, computed locally in bf16. | |
| | FPR | 0.1% | 0.5% | 1% | 2% | 5% | | |
| |---|---:|---:|---:|---:|---:| | |
| | Global cut | 0.00538 | 0.03359 | 0.09947 | 0.50391 | 0.94738 | | |
| Per-format cuts give each format its own operating point. They need the document's format, which the paper | |
| assigns with WebOrganizer's annotation prompt run on Gemini 3.7 Flash; the calibration documents use the formats | |
| recorded in WildOutlines. Each is the | |
| format's own quantile, shrunk toward the global cut with weight n / (n + 2500); at 0.1% FPR 10,000 documents | |
| per format are too few, so there is no per-format cut. A document outside these eight formats has no | |
| per-format cut; do not fall back to the global cut for it. | |
| | Format | 0.5% | 1% | 2% | 5% | | |
| |---|---:|---:|---:|---:| | |
| | Nonfiction Writing | 0.01682 | 0.04034 | 0.14608 | 0.74359 | | |
| | Knowledge Article | 0.02455 | 0.05113 | 0.16982 | 0.79109 | | |
| | Personal Blog | 0.10743 | 0.47575 | 0.79001 | 0.97165 | | |
| | News Article | 0.02812 | 0.07552 | 0.33610 | 0.86126 | | |
| | Academic Writing | 0.17216 | 0.61199 | 0.85898 | 0.98434 | | |
| | User Reviews | 0.15260 | 0.35787 | 0.76215 | 0.97357 | | |
| | Personal About Page | 0.05767 | 0.23257 | 0.65484 | 0.94379 | | |
| | Creative Writing | 0.36622 | 0.60222 | 0.80796 | 0.96994 | | |
| `thresholds.json` holds every cut at full precision, plus per-topic cuts for the WebOrganizer topics with enough calibration documents. | |
| These rates hold for English web documents like the training data. For another domain, fit the cut on | |
| human documents from that domain. | |
| IdeaLens-NoParaphrase scores 91% of human calibration documents above 0.99, so its cuts at higher FPRs sit close | |
| to 1 and small shifts in score move the realised FPR a long way. | |
| ## Hardware requirements | |
| Measured with transformers 5.15 in bf16 on NVIDIA H100 80GB GPUs (our other runs used A100 80GB), with transformers' | |
| PyTorch implementation of the Mamba layers (no fused Mamba kernels installed). We have not tried CPU-only inference. | |
| | | Merged model | Adapter route (`load_adapter.py`) | | |
| |---|---|---| | |
| | Download | 65.8 GB | 65.8 GB base model + 3.1 GB adapter | | |
| | Peak CPU RAM while loading | 60 GiB | 60 GiB | | |
| | GPU memory once loaded | 58.8 GiB | 58.8 GiB (66 GiB during the ~10 s it takes to apply the adapter) | | |
| GPU memory then grows with the length of the input, by about 4.2 MiB per token at typical lengths, scoring one | |
| input at a time: | |
| | Input tokens | 500 | 1,000 | 2,000 | 4,000 | 8,000 | | |
| |---|---:|---:|---:|---:|---:| | |
| | Peak GPU memory, one 80 GB GPU | 61.0 GiB | 63.1 GiB | 67.3 GiB | 75.7 GiB | does not fit | | |
| | Peak memory per GPU, two 80 GB GPUs (`device_map="auto"`) | | | | 47.4 GiB | 63.7 GiB | | |
| | Seconds per input, H100 | 0.18 | 0.34 | 0.66 | 1.32 | 2.70 | | |
| Inputs of 6,000 tokens do not fit on one 80 GB GPU and 12,000 do not fit on two; lowering the Mamba chunk size from | |
| 128 to 64 did not change either limit. | |
| IdeaLens-NoParaphrase reads outlines, which are short. The outlines in WildOutlines's calibration split average about 640 | |
| tokens with the prompt, and the longest is under 3,800, so one 80 GB GPU (A100 80GB or H100 80GB) is enough. Outline | |
| extraction runs through an LLM API and needs no local GPU. | |
| ## Intended use and limitations | |
| - IdeaLens-NoParaphrase estimates the provenance of a document's ideas. It should not be the sole basis for decisions about a person's work. | |
| - It was trained on English web documents of at least 500 words in eight long-form formats (Nonfiction Writing, Knowledge Article, Personal Blog, News Article, Academic Writing, User Reviews, Personal About Page, Creative Writing). | |
| - Its training labels come from the Pangram prose detector, applied to whole documents. They record who wrote the prose, and because these outlines keep some of the documents' wording, the model can learn from that wording as well as from the ideas. | |
| - Errors in outline extraction carry into the score. | |
| ## Related models | |
| | Model | Backbone | Reads | | |
| |---|---|---| | |
| | [IdeaLens](https://huggingface.co/rishanthrajendhran/IdeaLens) | Nemotron-3.5-Lightning-30B-A3B, LoRA | outline | | |
| | [ProseLens](https://huggingface.co/rishanthrajendhran/ProseLens) | Nemotron-3.5-Lightning-30B-A3B, LoRA | document text | | |
| | [IdeaLens-NoParaphrase](https://huggingface.co/rishanthrajendhran/IdeaLens-NoParaphrase) (this model) | Nemotron-3.5-Lightning-30B-A3B, LoRA | outline, trained without paraphrasing | | |
| | [IdeaLens-Qwen3.5-9B](https://huggingface.co/rishanthrajendhran/IdeaLens-Qwen3.5-9B) | Qwen3.5-9B, classification head | outline | | |
| | [IdeaLens-ModernBERT-L](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L) | ModernBERT-large | outline | | |
| | [ProseLens-ModernBERT-L](https://huggingface.co/rishanthrajendhran/ProseLens-ModernBERT-L) | ModernBERT-large | document text | | |
| | [IdeaLens-LogisticClassifier](https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier) | logistic regression over text-embedding-3-large | outline | | |
| | [IdeaLens-ModernBERT-L-NoParaphrase](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-NoParaphrase) | ModernBERT-large | outline, trained without paraphrasing | | |
| | [IdeaLens-ModernBERT-L-RolesOnly](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-RolesOnly) | ModernBERT-large | role labels only | | |
| | [IdeaLens-Qwen3.5-9B-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-Qwen3.5-9B-PerItem) | Qwen3.5-9B, classification head | single outline items, pooled | | |
| | [IdeaLens-ModernBERT-L-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-PerItem) | ModernBERT-large | single outline items, pooled | | |
| | [IdeaLens-LogisticClassifier-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem) | logistic regression over text-embedding-3-large | single outline items, pooled | | |
| Training data: [WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines). | |
| All IdeaLens models and datasets are in the [IdeaLens collection](https://huggingface.co/collections/rishanthrajendhran/idealens-6abee785ce6196fc0be9200f). | |
| ## License | |
| OpenMDW-1.1, the license of the base model (see `LICENSE`). | |
| ## Citation | |
| ```bibtex | |
| @article{idealens2026, | |
| title = {IdeaLens: Detecting AI Ideas in Long-form Writing}, | |
| author = {Rajendhran, Rishanth and Choi, Minjoon and Russell, Jenna and Namuduri, Ramya and B{\"o}l{\"o}ni-Turgut, Deniz and Karpinska, Marzena and Wieting, John and Iyyer, Mohit}, | |
| journal = {arXiv preprint arXiv:2610.06778}, | |
| year = {2026}, | |
| eprint = {2610.06778}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CL}, | |
| url = {https://arxiv.org/abs/2610.06778} | |
| } | |
| ``` | |