rishanthrajendhran commited on
Commit
9daf2d1
·
verified ·
1 Parent(s): 11f78dc

Rename the corpus IdeaLens-1M -> WildOutlines

Browse files
Files changed (2) hide show
  1. README.md +6 -6
  2. thresholds.json +1 -1
README.md CHANGED
@@ -10,7 +10,7 @@ tags:
10
  - ai-text-detection
11
  - idea-provenance
12
  datasets:
13
- - rishanthrajendhran/IdeaLens-1M
14
  extra_gated_prompt: "Access is granted individually. Please say who you are and what you intend to use the weights for."
15
  ---
16
 
@@ -21,7 +21,7 @@ role-labelled outline of the document (an ordered list of items, each giving one
21
  such as *Central Development* or *Open Question*) and returns P(human), the probability that the ideas are human.
22
 
23
  IdeaLens is `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` fine-tuned with LoRA (rank 64) on the outlines of 1M
24
- English web documents ([IdeaLens-1M](https://huggingface.co/datasets/rishanthrajendhran/IdeaLens-1M)). The training
25
  outlines were paraphrased to remove the documents' wording, so the model has to fit its labels through the ideas.
26
 
27
  ## Results
@@ -145,7 +145,7 @@ Scores agree with the training-time scores to about 0.001 in P(human) on average
145
 
146
  IdeaLens flags a document as having AI ideas when P(human) is below a cut. Each cut is set so
147
  that a given share of human documents is flagged (the false-positive rate, FPR), measured on the 80,000 human
148
- documents in IdeaLens-1M's `calibration` split (10,000 per format). The paper's operating point is the global
149
  cut at 1% FPR.
150
 
151
  | FPR | 0.1% | 0.5% | 1% | 2% | 5% | 10% | 20% |
@@ -154,7 +154,7 @@ cut at 1% FPR.
154
 
155
  Per-format cuts give each format its own operating point. They need the document's format, which the paper
156
  assigns with WebOrganizer's annotation prompt run on Gemini 3.7 Flash; the calibration documents use the formats
157
- recorded in IdeaLens-1M. Each is the
158
  format's own quantile, shrunk toward the global cut with weight n / (n + 2500); at 0.1% FPR 10,000 documents
159
  per format are too few, so there is no per-format cut. A document outside these eight formats has no
160
  per-format cut; do not fall back to the global cut for it.
@@ -198,7 +198,7 @@ input at a time:
198
  Inputs of 6,000 tokens do not fit on one 80 GB GPU and 12,000 do not fit on two; lowering the Mamba chunk size from
199
  128 to 64 did not change either limit.
200
 
201
- IdeaLens reads outlines, which are short. The outlines in IdeaLens-1M's calibration split average about 640
202
  tokens with the prompt, and the longest is under 3,800, so one 80 GB GPU (A100 80GB or H100 80GB) is enough. Outline
203
  extraction runs through an LLM API and needs no local GPU.
204
 
@@ -226,7 +226,7 @@ extraction runs through an LLM API and needs no local GPU.
226
  | [IdeaLens-ModernBERT-L-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-PerItem) | ModernBERT-large | single outline items, pooled |
227
  | [IdeaLens-LogisticClassifier-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem) | logistic regression over text-embedding-3-large | single outline items, pooled |
228
 
229
- Training data: [IdeaLens-1M](https://huggingface.co/datasets/rishanthrajendhran/IdeaLens-1M).
230
 
231
  ## License
232
 
 
10
  - ai-text-detection
11
  - idea-provenance
12
  datasets:
13
+ - rishanthrajendhran/WildOutlines
14
  extra_gated_prompt: "Access is granted individually. Please say who you are and what you intend to use the weights for."
15
  ---
16
 
 
21
  such as *Central Development* or *Open Question*) and returns P(human), the probability that the ideas are human.
22
 
23
  IdeaLens is `nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` fine-tuned with LoRA (rank 64) on the outlines of 1M
24
+ English web documents ([WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines)). The training
25
  outlines were paraphrased to remove the documents' wording, so the model has to fit its labels through the ideas.
26
 
27
  ## Results
 
145
 
146
  IdeaLens flags a document as having AI ideas when P(human) is below a cut. Each cut is set so
147
  that a given share of human documents is flagged (the false-positive rate, FPR), measured on the 80,000 human
148
+ documents in WildOutlines's `calibration` split (10,000 per format). The paper's operating point is the global
149
  cut at 1% FPR.
150
 
151
  | FPR | 0.1% | 0.5% | 1% | 2% | 5% | 10% | 20% |
 
154
 
155
  Per-format cuts give each format its own operating point. They need the document's format, which the paper
156
  assigns with WebOrganizer's annotation prompt run on Gemini 3.7 Flash; the calibration documents use the formats
157
+ recorded in WildOutlines. Each is the
158
  format's own quantile, shrunk toward the global cut with weight n / (n + 2500); at 0.1% FPR 10,000 documents
159
  per format are too few, so there is no per-format cut. A document outside these eight formats has no
160
  per-format cut; do not fall back to the global cut for it.
 
198
  Inputs of 6,000 tokens do not fit on one 80 GB GPU and 12,000 do not fit on two; lowering the Mamba chunk size from
199
  128 to 64 did not change either limit.
200
 
201
+ IdeaLens reads outlines, which are short. The outlines in WildOutlines's calibration split average about 640
202
  tokens with the prompt, and the longest is under 3,800, so one 80 GB GPU (A100 80GB or H100 80GB) is enough. Outline
203
  extraction runs through an LLM API and needs no local GPU.
204
 
 
226
  | [IdeaLens-ModernBERT-L-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-ModernBERT-L-PerItem) | ModernBERT-large | single outline items, pooled |
227
  | [IdeaLens-LogisticClassifier-PerItem](https://huggingface.co/rishanthrajendhran/IdeaLens-LogisticClassifier-PerItem) | logistic regression over text-embedding-3-large | single outline items, pooled |
228
 
229
+ Training data: [WildOutlines](https://huggingface.co/datasets/rishanthrajendhran/WildOutlines).
230
 
231
  ## License
232
 
thresholds.json CHANGED
@@ -3,7 +3,7 @@
3
  "input": "outline",
4
  "flag_rule": "flag the document as AI when P(human) < cut",
5
  "calibration": {
6
- "data": "IdeaLens-1M, calibration split (human documents only)",
7
  "scores_from": "the training checkpoint, scored on Tinker",
8
  "n_humans": 80000,
9
  "n_per_format": {
 
3
  "input": "outline",
4
  "flag_rule": "flag the document as AI when P(human) < cut",
5
  "calibration": {
6
+ "data": "WildOutlines, calibration split (human documents only)",
7
  "scores_from": "the training checkpoint, scored on Tinker",
8
  "n_humans": 80000,
9
  "n_per_format": {