Skip to main content

idealens

idealens detects whether a document's ideas came from a person or an AI model, whoever wrote the words. It runs the IdeaLens pipeline end to end:

  1. Classify the document into one of eight long-form formats (an LLM with WebOrganizer's annotation prompt, or WebOrganizer's encoder).
  2. Extract an outline: an LLM writes a role-labelled list of the document's ideas, using the same prompt and six worked examples the detector was trained with.
  3. Score the outline with a detector, which returns P(human) and a verdict at a calibrated false-positive rate.

The same package also scores the prose-level detector ProseLens and the baselines released with the paper.

Models: IdeaLens and the other detectors listed under Models. Data: WildOutlines (training corpus), IdeaShift, IdeaShift-X and TwiceTold (evaluation sets).

Install

pip install "idealens[vllm]"          # the default scoring backend (Nemotron models, one 80 GB GPU)
pip install "idealens[hf]"            # transformers backend; needed for the ModernBERT and Qwen models
pip install "idealens[openai]"        # OpenAI, OpenRouter or any OpenAI-compatible endpoint (also the logistic models)
pip install "idealens[anthropic]"     # Claude
pip install "idealens[vertex]"        # Gemini on Vertex AI (batch mode also needs a GCS bucket)
pip install "idealens[weborganizer]"  # WebOrganizer's encoder for format classification
pip install "idealens[all]"

Python 3.10 or later. The vLLM extra pins vllm==0.21.*, transformers>=5.15 and xgrammar==0.2.1.

The model repos are gated: request access on the model's Hugging Face page, then huggingface-cli login.

Credentials

Provider (--provider) Variable
gemini (default) GEMINI_API_KEY or GOOGLE_API_KEY
vertex GOOGLE_CLOUD_PROJECT and application-default credentials; IDEALENS_GCS_BUCKET for batch mode
openai OPENAI_API_KEY
anthropic ANTHROPIC_API_KEY
openrouter OPENROUTER_API_KEY
compatible --base-url and, if the endpoint needs one, --api-key

The logistic models embed outlines with OpenAI's text-embedding-3-large, so they also need OPENAI_API_KEY.

Quick start

Input is JSONL with a text field and, optionally, id, url, format and topic.

idealens run docs.jsonl -o scores.jsonl --dry-run   # estimate tokens and cost; sends nothing
idealens run docs.jsonl -o scores.jsonl             # classify, extract and score with IdeaLens

Or one step at a time, which lets you extract once and score with several models:

idealens classify docs.jsonl     -o formats.jsonl
idealens extract  formats.jsonl  -o outlines.jsonl
idealens score    outlines.jsonl -o scores.jsonl --model IdeaLens
idealens score    outlines.jsonl -o scores_np.jsonl --model IdeaLens-NoParaphrase

Every command keeps the input record's fields and appends to its output. Re-running a command resumes it: records already in the output are skipped, and records whose provider call failed are retried (the failed attempts are moved to <output>.failed.jsonl).

In Python:

import idealens as il

texts = ["...", "..."]
formats = il.classify(texts)                   # gemini-3.7-flash by default
outlines = il.extract(texts, formats)
with il.Detector("IdeaLens") as det:           # vLLM; the repo's published thresholds
    records = det.score_outlines(outlines, format=formats)

for r in records:
    print(r["p_human"], r["verdict"]["ai"])

il.run(texts, det) does all three steps. Any provider can be passed as an object:

from idealens.providers import make
prov = make("openai", "gpt-6-sol")
outlines = il.extract(texts, formats, provider=prov, mode="batch")

Ways to use idealens

Which detector?

You want Use
The paper's main idea-level detector IdeaLens (one 80 GB GPU)
The idea-level detector trained on outlines exactly as this package extracts them IdeaLens-NoParaphrase
Idea-level detection on a small GPU or a CPU IdeaLens-ModernBERT-L (or -NoParaphrase)
Idea-level detection with no GPU at all IdeaLens-LogisticClassifier (needs OpenAI embeddings)
A score for every outline item, not just the document the -PerItem models (item_p_human in each record)
How much the outline's structure alone gives away IdeaLens-ModernBERT-L-RolesOnly (reads only the role sequence)
Who wrote the prose, for comparison ProseLens, ProseLens-ModernBERT-L (no outline, no LLM calls)

1. From documents, end to end

idealens run classifies each document's format, extracts its outline and scores it. This is the route the published thresholds assume when the extractor is gemini-3.7-flash (the default).

idealens run docs.jsonl -o scores.jsonl --model IdeaLens

2. Extract once, score with several detectors

Extraction is the expensive step. Keep its output and score it with as many outline models as you like:

idealens classify docs.jsonl     -o formats.jsonl
idealens extract  formats.jsonl  -o outlines.jsonl
idealens score outlines.jsonl -o idealens.jsonl  --model IdeaLens
idealens score outlines.jsonl -o modernbert.jsonl --model IdeaLens-ModernBERT-L
idealens score outlines.jsonl -o per_item.jsonl   --model IdeaLens-ModernBERT-L-PerItem

3. Score outlines you already have

The outline field (or the field named by --input-field) can hold the extractor's JSON object or plain text with one [Role] content line per item:

{"id": "doc1", "format": "News Article", "outline": "[Central Development] The council approved the new pipeline.\n[Background Context] The reservoir has been shrinking for three summers."}
with il.Detector("IdeaLens-NoParaphrase") as det:
    records = det.score_outlines(["[Central Development] ...\n[Background Context] ..."], format=["News Article"])

Roles must come from the format's role vocabulary for the scores to mean what they meant in training; the extractor in this package uses it.

4. Score documents directly

The prose detectors read the document itself, so no outline and no LLM call is needed:

idealens score docs.jsonl -o prose.jsonl --model ProseLens

5. Choose the LLM that classifies and extracts

Any provider in Credentials works for classify, extract, run and calibrate, online or as a batch job (--mode batch, half price where the provider has a batch API). That includes a model you serve yourself behind an OpenAI-compatible endpoint (vLLM, Ollama, ...):

idealens run docs.jsonl -o scores.jsonl --provider openai   --llm-model gpt-6-sol --mode batch
idealens run docs.jsonl -o scores.jsonl --provider compatible --base-url http://localhost:8000/v1 --llm-model <served model>

The published thresholds were fitted on gemini-3.7-flash outlines. With another extractor every record carries a warning, and calibrating on your own human documents (step 8) restores a known false-positive rate.

6. Give formats yourself, or classify locally

A record's format field, or --format for the whole file, skips classification. A format outside the eight is refused unless you pass --force-fit, which maps it to the closest one. --method weborganizer classifies with WebOrganizer's encoder on your own machine instead of an LLM (pip install "idealens[weborganizer]").

idealens run reviews.jsonl -o scores.jsonl --format "User Reviews"
idealens classify docs.jsonl -o formats.jsonl --method weborganizer --device cuda

7. Choose how the model runs

Backend When
vllm (default for the Nemotron models) fastest; one 80 GB GPU
hf transformers; for the Nemotron models, --hf-mode adapter downloads the base model plus a 3 GB adapter instead of the merged weights
logistic the two logistic models; Detector(..., embed=fn) takes your own text-embedding-3-large vectors (for example, cached ones) instead of calling OpenAI

--weights PATH scores with a local copy of a model's weights instead of downloading them (a directory, or the .npz file for the logistic models). In Python, Detector passes extra keyword arguments to the backend, e.g. il.Detector("IdeaLens", tensor_parallel_size=2).

8. Choose the operating point, or calibrate your own

Every record carries verdicts at all calibrated false-positive rates and schemes (see Verdicts and thresholds); --fpr and --scheme pick the default one. For a new domain, fit cuts on human documents from it and score against them:

idealens calibrate my_humans.jsonl --save-as my_domain --model IdeaLens --group-by source
idealens score outlines.jsonl -o scores.jsonl --thresholds my_domain --group-by source --scheme group:source

9. Large jobs

--dry-run prices a run before anything is sent; idealens cost out.jsonl totals a finished one. Every command appends to its output and can be re-run after an interruption: finished records are skipped and failed provider calls are retried. --chunk sets how many records are scored per write, --workers how many online requests run at once.

10. Without this package

Each model card on Hugging Face shows how to load the weights with transformers or vLLM and read P(human) directly, and each model repo holds its thresholds.json.

Models

idealens models lists them. All read the output of the same extraction step, except the two ProseLens models, which read the document itself.

Model Reads Backends
IdeaLens (default) outline vllm, hf
ProseLens document vllm, hf
IdeaLens-NoParaphrase outline vllm, hf
IdeaLens-Qwen3.5-9B outline hf
IdeaLens-Qwen3.5-9B-PerItem each outline item hf
IdeaLens-ModernBERT-L outline hf
IdeaLens-ModernBERT-L-NoParaphrase outline hf
IdeaLens-ModernBERT-L-RolesOnly the outline's role sequence hf
IdeaLens-ModernBERT-L-PerItem each outline item hf
ProseLens-ModernBERT-L document hf
IdeaLens-LogisticClassifier outline logistic
IdeaLens-LogisticClassifier-PerItem each outline item logistic

Per-item models score each item and pool the item scores by their mean log-odds. Each model uses its own backend unless you pass --backend.

The package does not paraphrase outlines. IdeaLens was trained and calibrated on paraphrased outlines; IdeaLens-NoParaphrase was trained on outlines as extracted, and is the closer match to what this package produces.

Verdicts and thresholds

A document is flagged as AI when its P(human) is strictly below the cut. The default verdict uses the 1% global cut: the score below which 1% of the human calibration documents fall. Each record carries:

Field Contents
p_human the detector's P(human)
verdict the default verdict: fpr, scheme, cut, ai
verdicts every calibrated false-positive rate (0.1%, 0.5%, 1%, 2%, 5%) under each scheme: global, per_format, per_topic, and group:<field> for cuts you fitted yourself
warnings departures from how the cuts were fitted, such as a different extractor model
item_p_human per-item scores (per-item models)

Choose another operating point with --fpr 0.005 --scheme per_format. When a document has no cut under a scheme (its format or topic was not calibrated, or its format was force-fitted), that verdict is None with a reason. It never falls back to the global cut.

The published cuts were fitted on 80,000 human documents with outlines extracted by gemini-3.7-flash with the six worked examples. Other extractor models work, but their outlines can shift the score distribution, so the calibrated false-positive rate is no longer guaranteed. Every record says so in warnings.

Calibrating on your own data

If your documents differ from web text, fit cuts on human documents from your own domain:

idealens calibrate my_humans.jsonl --save-as my_domain --model IdeaLens
idealens score outlines.jsonl -o scores.jsonl --thresholds my_domain

The input can be scored records, outlines or documents; the missing steps are run first. --group-by FIELD fits a cut per value of a field, and --label-field / --human-value let you pass a mixed file and calibrate on its human rows. A cut is fitted only where there are enough documents to estimate it (at least 25 expected below the cut and 200 in the group). Profiles are saved under ~/.config/idealens/thresholds (or $IDEALENS_HOME/thresholds).

Formats

The detectors were trained on eight formats: Academic Writing, Creative Writing, Knowledge Article, News Article, Nonfiction Writing, Personal About Page, Personal Blog and User Reviews. A document in another format is assigned the closest of the eight by a second LLM pass (or by the encoder's probabilities). Such a document has forced_format: true, keeps its original label in original_format, and gets no per-format verdict.

If you know a document's format, give it in the record's format field or with --format; classification is then skipped. idealens run --check-format classifies anyway and reports disagreements.

Cost

--dry-run estimates the tokens and cost of a run before anything is sent, and idealens cost output.jsonl totals a finished one. Extraction dominates: the prompt carries six worked examples, and reasoning tokens make up most of the output. Batch mode (--mode batch) costs half as much where the provider has a batch API (Gemini on Vertex, OpenAI, Anthropic); not every model is served by it.

Hardware

Model Needs
IdeaLens, ProseLens, IdeaLens-NoParaphrase one 80 GB GPU (A100 80GB or H100 80GB); the weights take 59 GiB; about 60 GiB of CPU RAM while loading
IdeaLens-Qwen3.5-9B models one GPU; the weights take about 16 GB in bf16 (tested on 80 GB GPUs)
ModernBERT models any GPU; a CPU works for small jobs
Logistic models CPU only

The vLLM backend checks that the weights fit before it starts, and stops with a message otherwise. Use tensor_parallel_size (Detector(..., tensor_parallel_size=2)) to spread a model over two smaller GPUs.

Development

pip install -e ".[dev]"
pytest tests

The tests need no GPU, network or API key.

License

The code is released under the Apache License 2.0. Each model has its own license, given on its model page.

Citation

A citation will be added here once the paper is on arXiv.

Metadata

Release files for idealens 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for idealens 0.1.1
File Size Uploaded
idealens-0.1.1.tar.gz 723.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for idealens 0.1.1
File Interpreter ABI Platform
idealens-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 1.4 MB

Release files / idealens-0.1.1.tar.gz

Download URL idealens-0.1.1.tar.gz
Size 723.6 kB
Tags Source
SHA-256 checksum
How to use checksums
ea8bc1f4839c12615108cf3356b784323e6f1617285c63a06a5a9f78732d4c96
BLAKE2b-256 checksum
How to use checksums
98359e714da520fdee65528cf7f879795bf976e8354c4d8d7d26b32155285704
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / idealens-0.1.1-py3-none-any.whl

Download URL idealens-0.1.1-py3-none-any.whl
Size 723.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6ce55e1bd4775795ccd9174bf187b3c4dd2ea384023585c4361afc4cca55f829
BLAKE2b-256 checksum
How to use checksums
494c8e638a4938fd2a126af54b468ce0cc785fc4a37081d0a700b39536ebfa19
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page