Research
IdeaLens: Detecting AI Ideas in Long-form Writing
IdeaLens: Detecting AI Ideas in Long-form Writing Overview Research area: Natural Language Processing; AI text detection, authorship attribution, and computational writing research. Technical level: I

- arXiv
- 2610.06778
- Published
- 2026-10-05
- Authors
- Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Ramya Namuduri, Deniz Bölöni-Turgut, Marzena Karpinska, John Wieting, Mohit Iyyer
AI summary
IdeaLens: Detecting AI Ideas in Long-form WritingOverview
Research area: Natural Language Processing; AI text detection, authorship attribution, and computational writing research.
Technical level: Intermediate. The core idea is intuitive and explained in plain terms, but the paper assumes familiarity with detectors, false positive rates, classifiers, and fine-tuning.
Scope (one sentence): The paper introduces IdeaLens, a detector trained on document outlines rather than raw text that identifies whether a document's ideas came from a human or an AI, regardless of who wrote the words.
What This Paper Is About
Existing AI detectors answer the question "who wrote these words?" but emerging publishing and academic policies increasingly care about "who came up with these ideas?" The two can differ: a person can brainstorm a document and have an AI write it, or an AI can supply the plan that a person then writes out by hand. The authors build a detector that targets idea provenance rather than prose provenance, and evaluate it against prose-focused detectors such as Pangram 4 and EditLens on both kinds of mixed authorship. The paper is motivated by a real incident in which Pangram flagged a Wall Street Journal op-ed by Stanley Druckenmiller as 100% AI-generated, while the editor argued that what mattered was whether the op-ed reflected the author's original argument.
Key Contributions
-
IdeaLens, an idea provenance detector. Instead of reading a document's words, IdeaLens reads an outline: an ordered list of discourse roles, each paired with a short paraphrased description of content. Outlines are largely stripped of surface-level wording, so the model must fit its training labels mainly through ideas.
-
ProseLens, a controlled comparison model. ProseLens shares the same backbone, the same 1M training documents, and the same Pangram silver labels as IdeaLens, but trains on raw text. Any behavioral difference between the two is therefore attributable to the outline representation.
-
Training data at detector scale. The authors build WildOutlines, a corpus of 1M outlines extracted from FineWeb documents labeled with Pangram 3.3.2 via the WildAI dataset, with eight WebOrganizer formats and a minimum length of 500 English words. The corpus cost about $66.4k to build.
-
Three new evaluation datasets and a public release. IdeaShift (AI writing from increasingly detailed human plans), IdeaShift-X (24 languages), and TwiceTold (50 stories written from scratch by humans following AI-generated outlines) are introduced, alongside an item-level analysis of 90K IdeaLens predictions. Models, training data, evaluation datasets, and the idea-level analysis are released.
Main Findings
IdeaLens tracks ideas, not prose. Across 51 evaluation splits from 22 benchmarks, IdeaLens is the only detector that is accurate under both shared provenance (95.3%) and mixed provenance (81.3%). ProseLens and Pangram 4 are extremely accurate under shared provenance (99.1% and 98.5%) but fail under mixed provenance (25.4% and 25.9%). On weighted average accuracy over mixed-provenance benchmarks, IdeaLens reaches 92.6% versus 46.8% for ProseLens and 41.1% for Pangram.
It detects human ideas inside AI-generated text. On IdeaShift, as the human plan in the prompt becomes more detailed, IdeaLens's flag rate falls from 94.9% at level 0 to 6.8% at level 5. ProseLens and Pangram 4 flag 86–100% of documents regardless of the level of human ideation. The largest drop occurs between levels 3 and 4, when prompts begin including outline content and not just structural role information. When plans come from AI rather than humans, IdeaLens flags more than 96% of documents at every level.
It detects AI ideas inside human-written text. On TwiceTold, IdeaLens flags 68% (95% CI 54–79%) of stories that humans wrote from scratch from AI outlines, compared to 0% for ProseLens and 8% for Pangram 4. All detectors reliably flag the AI-generated original stories (96% for IdeaLens, 100% for ProseLens and Pangram). Of the 16 stories IdeaLens marked as human, 12 contained speculative or fantasy elements. The authors note Pangram 4's flags on these stories are likely due to character names typical of AI writing, such as Mara or Bellweather.
It is insensitive to AI rewording of human ideas. Aggregated over 12 benchmarks where AI makes surface-level changes to human text, IdeaLens labels 94.7% of documents as human, versus 50.7% for ProseLens and 43.8% for Pangram 4. On the introduction's framing of the same comparison, IdeaLens classifies 94.7% of AI-generated documents based on human ideas as human while Pangram labels only 43.8% as human. Dataset-specific results: in GEDE, IdeaLens flags 0.5% of human essays as AI after minor LLM corrections (Pangram 4: 54.6%; ProseLens: 45.8%) and 2.6% after a content-preserving rewrite (Pangram 4: 94.3%; ProseLens: 77.1%). In HART, IdeaLens flags 4.9% of LLM paraphrases of human documents (Pangram 4: 42.8%; ProseLens: 44.5%). On the Academic Integrity dataset, it flags 8.2% of abstracts an LLM generated from a human-written reference abstract (Pangram 4: 96.4%; ProseLens: 83.9%). Across 10 external benchmarks of polishing, paraphrasing, and rewriting, IdeaLens flags 5.6% of documents on average versus 46.0% for ProseLens and 52.0% for Pangram. On 12 benchmarks where AI makes surface-level changes, IdeaLens labels 94.7% of documents as human.
It remains competitive on conventional detection. On the 19 evaluation splits with shared provenance, IdeaLens detects 91.1% of AI documents at 0.6% FPR, versus 81.8% at 0.2% FPR for EditLens-3B, the strongest existing open baseline (both calibrated to a 1% target). ProseLens does better at 98.9% at 0.7% FPR, rivaling Pangram 4 at its published threshold (97.3% at 0.3%). On 10K pre-ChatGPT documents from C4, IdeaLens flags only 0.01%, compared with 0% for ProseLens and Pangram 4. Against humanization and adversarial attacks, IdeaLens's detection rate drops 6.4 points on average, between Pangram's 4.9 and ProseLens's 10.6. Its one noted weakness is DIPPER paraphrasing in GEDE, where it detects 77.6% versus 97.1% for ProseLens and 95.1% for Pangram 4, possibly because that attack changes content as well as style.
It transfers across 24 languages without non-English training. On IdeaShift-X, IdeaLens flags 95.3% of level-0 documents but only 0.7% of level-5 documents, despite training only on English. ProseLens and Pangram 4 flag 82–85% of level-0 and 45–62% of level-5 documents, and the gap widens in low-resource languages such as Tamil and Amharic. On 10K pre-ChatGPT human-written documents in these 24 languages from FineWeb2, IdeaLens flags only 0.1% as AI. Aggregated across 24 languages, IdeaLens detects 95.3% of AI documents at 0.1% FPR, compared with 82.5% and 0.0% for ProseLens.
Idea-level errors are systematic. From 90K item-level predictions, IdeaLens scores ideas as more AI when they cite numbers or methods without enough context to check them, or claim more than their evidence supports. It scores ideas as more human when they include specific sources a reader could look up, or discuss downsides and open problems. The paper's Table 1 contrasts human and AI ideas; the first listed contrast is "Decisive vs. Universal Details" (human details could alter the author's point or conclusion; AI details could apply to any similar subject). The remaining rows of that table are not reproduced in the available content.
Methodology in Plain English
The core problem is a label mismatch: large-scale labels exist for who wrote a document's words, but not for who produced its ideas. The authors' solution is to keep the cheap prose labels and change what the classifier sees.
-
Define ideas as an outline. Drawing on cognitive writing models that separate planning from translating a plan into prose, an "idea" is what an author decided to say. Each document is reduced to an outline: an ordered list of discourse roles (for example, an Event Account followed by an Evaluation in the op-ed example, which has 17 items) each paired with a brief content description. Because roles are format-specific, separate role sets are induced for each of the eight formats (the paper mentions 42 roles for nonfiction).
-
Extract outlines automatically. Role induction adapts the TopicGPT framework, prompting GPT-5.6 Sol to propose base roles per format and then labeling 200 documents per format while inventing new roles as needed, followed by deduplication and pruning. Outline extraction prompts Gemini 3.7 Flash and Gemini 3.1 Pro with the format's role set and few-shot demonstrations from Claude Fable 5. Paraphrasing uses Gemini 3.1 Pro.
-
Paraphrase to remove wording. Extracted outlines can copy original phrasing, letting a model cheat via style. Paraphrasing pushes all outlines toward a homogeneous, AI-like style: for human-written documents, the fraction of outline text Pangram 4 labels as human drops from 8.5% to 0.3%. Meaning is preserved: paraphrased outlines have cosine similarity 0.93 with the originally extracted outlines (versus 0.39 for unrelated outlines) and 0.76 with the source document (versus 0.26). Paraphrasing is used only at training time; at inference IdeaLens sees the unparaphrased outline.
-
Train two matched models. IdeaLens and ProseLens both start from a Nemotron-3.5-Lightning-30B-A3B backbone, fine-tuned for one epoch with LoRA using Tinker. The 1M documents are split into roughly 842k training, 30k validation, 80k calibration (human documents only), and 50k test. Because the calibration split contains only human documents, its threshold can target a specific false positive rate; by default detectors are evaluated at 1% FPR on calibration.
-
Compare against baselines. Baselines include Pangram 4, EditLens-Llama-3.2-3B, Fast-DetectGPT, Binoculars, MELD, and Desklib, with Pangram's ternary labels binarized per its technical report. Evaluation covers 19 external benchmarks, plus IdeaShift (500 seed documents, prompts from level 0 containing only a topic up to level 5 containing the full extracted outline), IdeaShift-X (10K human-authored documents in 24 languages from FineWeb2, with AI versions at levels 0 and 5), and TwiceTold (50 stories written from scratch by 17 computer science graduate students or faculty given only the outline of a 500-word GPT-5.6 Sol story, with the 50 AI originals included).
Why This Matters
Impact on research. The paper reframes AI detection as an authorship-of-ideas problem and shows that the outline representation, not a larger model or better labels, is what produces the behavioral shift: IdeaLens and ProseLens share a backbone, training documents, and labels. The released models, labeled datasets, and idea-level analysis give the field a shared testbed for a question that prior work only approached through simulated collaborations, which the authors argue are expensive and may not transfer to real ones. The multilingual result is notable because it emerges without any non-English training.
Real-world applications.
- Editorial and publishing review, where an outlet may want to know whether a submission's argument is the author's own, as in the Wall Street Journal op-ed case the paper opens with.
- Academic integrity and peer review workflows, such as judging whether an abstract or manuscript's conceptual content originated with the listed authors, especially given findings like the Academic Integrity dataset result.
- Education, where an assignment might permit AI-assisted wording but not AI-supplied reasoning, a distinction prose detectors cannot express.
- Multilingual content platforms, since IdeaLens applies an English outline extractor to documents in 24 languages, including low-resource ones.
Industry relevance. The four-quadrant framing (fully human, fully AI, human ideas with AI prose, AI ideas with human prose) maps directly onto how AI writing tools are actually used in workplaces, and onto AI-use policies that are being written around the notion of original argument. The paper notes that IdeaLens's backbone is a mixture-of-experts model with only 3B active parameters, and reports a 0.4B-parameter IdeaLens-ModernBERT variant that detects 79.1% of AI documents at 1.5% FPR, which is relevant to deployability.
Future Directions
- Understanding the level 3–4 transition. IdeaLens's flag rate changes most sharply between IdeaShift levels 3 and 4, when prompts begin carrying outline content rather than only structural role information. The paper reports the pattern; what specifically about content-bearing plans shifts the signal is left open.
- Explaining the residual misses on human writing. IdeaLens flags 68% of TwiceTold stories, and 12 of its 16 errors involved speculative or fantasy elements. Whether this reflects a genuine property of AI ideation in those genres or a limitation of the outline extractor is unresolved.
- Robustness to attacks that alter content. The reported weakness is DIPPER paraphrasing in GEDE (77.6% detection versus 97.1% for ProseLens and 95.1% for Pangram 4), which the authors suspect changes content as well as style. Defending against content-altering attacks is an open problem.
- What the systematic differences mean. The item-level analysis of 90K predictions produces patterns such as unverifiable citations or overclaiming correlating with AI ideation, but the paper treats these as detector behavior to characterize rather than as a settled account of how humans and AI differ when generating ideas.
Target Audience
Researchers in AI text detection, authorship attribution, and computational linguistics; NLP practitioners building provenance or content-moderation systems; journal editors, peer-review bodies, and academic integrity officers drafting or enforcing AI-use policies; educators evaluating the distinction between AI-assisted wording and AI-supplied reasoning; and multilingual platform teams who need detection that works beyond English. Readers without an NLP background can follow the Motivation and Findings but will need familiarity with false positive rates and classifier evaluation to interpret the quantitative comparisons.
Authors’ abstract
While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a brief, paraphrased description of the content, minimizing word-level overlap with the raw text. We train IdeaLens on 1M FineWeb documents with silver labels from Pangram, a prose provenance detector. Since the outlines are largely stripped of surface-level information, the labels must be fit mainly through the ideas. In a controlled study, IdeaLens's AI flag rate drops from 95% to 7% as models write from increasingly detailed human plans, while Pangram 4 still flags 92%; from AI-derived plans, IdeaLens stays above 96%. Conversely, on a new dataset of 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% of the stories as AI, compared to 8% for Pangram 4. On a comprehensive suite of 19 existing detection benchmarks, we show that IdeaLens maintains strong detection rates at low false positive rates, suggesting that ideas themselves provide a powerful discriminative signal, and its performance holds across domains, formats, and languages. Finally, we examine 90K predictions from IdeaLens to characterize systematic differences between human and AI ideation. We release our models and labeled datasets to facilitate future research on idea provenance detection.