Skip to main content

paper-preflight

English · 简体中文

CI License: MIT Python 3.11–3.14

Check every reference of a LaTeX paper against real scholarly records before you submit. No LLM guessing, no false accusations.

paper-preflight checking the demo paper: errors for an undefined citation key, a DOI that belongs to another paper, a reference no source knows and a retracted paper; warnings for a published preprint, a wrong year and a LaTeX-escaped DOI

Language models invent references, and copy-pasted BibTeX carries wrong years, wrong authors and dead DOIs. paper-preflight reads your .tex and .bib files and asks Crossref, dblp, arXiv, DataCite, PubMed and OpenAlex (and Semantic Scholar, if you have a key) about every cited work:

  • Does it exist?
  • Does it match what you wrote?
  • Has it been retracted?
  • Has the preprint you cite been published since?

When it cannot tell, it says so instead of guessing.

Status: v0.3, an early release. False positives are the bugs we most want to hear about: please open an issue.

The repository's demo paper cites eleven works, several of them wrong on purpose. A real run, against the live sources:

$ paper-preflight check examples/demo-paper
paper-preflight 0.3.0 · main.tex · 12 entries, 12 cited keys

error   CIT001 main.tex:31
    Citation key 'nonexistent2023' is not defined in any bibliography file (1 use(s)).
error   REF001 refs.bib:43
    The doi of 'devlin2019bert' (10.1109/cvpr.2016.90) resolves to a different work in Crossref: "Deep Residual Learning for Image Recognition" (He et al., 2016).
error   REF003 refs.bib:66
    'lindqvist2024quantum' was not found in Crossref, dblp and Semantic Scholar, and every source responded. Check that the work exists and that its title is correct.
error   REF004 refs.bib:73
    'wakefield1998ileal' has been retracted (reported by Crossref, OpenAlex). Cite it only if the text discusses the retraction.
error   CIT002 refs.bib:127
    Entry key 'kingma2015adam' is already defined at line 47; BibTeX ignores this one.
warning REF015 refs.bib:31
    'he2015residual' cites a preprint that has been published in CVPR (2016), DOI 10.1109/cvpr.2016.90. Cite the published version and keep the eprint field.
warning CIT004 refs.bib:37
    Entries 'devlin2019bert' and 'he2016deep' look like the same work (same DOI).
warning REF013 refs.bib:51
    'kingma2015adam' gives the year 2016, but dblp records 2014, 2015.
warning REF017 refs.bib:111
    The doi of 'tacl2019example' contains LaTeX escapes: '10.1162/tacl\_a\_00276'. Write it as: 10.1162/tacl_a_00276
info    REF005 refs.bib:73
    'wakefield1998ileal' has a published correction (reported by Crossref).
info    REF090 refs.bib:86
    'goodfellow2016deep' could not be verified: grey literature without an identifier (book, report, software, web page).
info    REF090 refs.bib:94
    'zhou2016ml' could not be verified: non-Latin titles are not supported yet; grey literature without an identifier (book, report, software, web page).
info    CIT003 refs.bib:115
    Entry 'lecun1998gradient' is never cited.

References: 6 verified · 1 metadata mismatch · 1 identifier conflict · 1 not found · 2 cannot determine
5 error(s) · 4 warning(s) · 4 info

Each finding is backed by a record (or by every source answering "no"). The correct NeurIPS paper is verified through dblp even though Crossref only holds fake copies of it, and the two books without identifiers are reported as "cannot determine" instead of "not found".

What it catches

Rule Finding
REF001 The DOI or arXiv ID points to a different paper
REF002 The DOI or arXiv ID does not exist
REF003 The work was not found in any source, and every source answered
REF004 · REF005 The work was retracted, or has an expression of concern or a correction
REF010–REF014 Authors, title, year or venue differ from the real record
REF015 A cited preprint has been formally published
REF016 The registry has a DOI the entry lacks (offered as a safe fix)
REF017 An identifier is written so that links break (10.1162/tacl\_a\_00276, …v1)
CIT001–CIT008 Undefined, duplicate, unused or near-duplicate citation keys; broken .bib syntax
REF090 Cannot determine, always with the reason (source unavailable, grey literature, …)

paper-preflight explain REF003 describes any rule.

How accurate is it?

Three measurements, all against the live sources: the bibliographies of real papers, a head-to-head with published tools, and a public benchmark.

On real papers

The bibliographies of 20 arXiv papers from the turn of July and August 2026 (cs, stat, q-bio, quant-ph and astro-ph), chosen mechanically and collected only after every fix in this release, with every warning and error reviewed by hand:

References Flags Real problems False positives Unclear False positives per 100 references
1,005 88 66 19 3 1.9
  • Fewer than one false alarm per paper (50 references on average), against 66 real problems: 31 errors in the entries (invented co-authors and given names, wrong or malformed DOIs, wrong titles and years) and 35 cited preprints that have since been published.
  • The false alarms are mostly names written another way (initials without dots, a generational suffix, a nickname or an English name) and records the registries got wrong (a registry listing 3 of 10 authors, a garbled title). The previous release has 2.1 on these papers.
  • Four earlier batches of 20 papers were used to find false positives, each first measured as it came out (0.1.0: 4.5 per 100 references; 0.1.1: 2.3; 0.1.2 before its last fixes: 3.0; 0.1.2: 1.7). On all four, this release has 0.1 to 0.7. Details in evals/README.md.

Next to other tools

Badalova & Mayr (2026) checked 104 references by hand and published what five tools flagged. On the same references, with their labels:

Tool Precision [95% CI] Recall False flags per 100 correct references
CheckIfExist 47.7% [36.0%, 59.6%] 93.9% 47.9
HalluCiteChecker 47.4% [32.5%, 62.7%] 54.5% 28.2
Hallucinator 50.9% [38.3%, 63.4%] 87.9% 39.4
HalRef 31.2% [21.9%, 42.2%] 72.7% 74.6
RefChecker 47.1% [35.7%, 58.8%] 97.0% 50.7
paper-preflight 72.5% [57.2%, 83.9%] 87.9% 15.5

The sample is small, so the intervals are wide. Some flags count as false here because the study labels a reference correct when the work exists: five of paper-preflight's flags on such references point at real errors (a wrong author, a broken DOI). Two causes of false flags found in this data were fixed, and four names the study's CSV garbled were restored, before the run above; the first run measured 62.8%. See evals/results/badalova-mayr.md.

On a benchmark: HALLMARK

HALLMARK is a public benchmark of real and hallucinated BibTeX entries.

Split Mode Precision Recall False-positive rate Coverage
test_public: 831 entries, never used during development Any issue 98.1% 88.9% 2.2% 97.0%
Fabrication 99.0% 49.0% 0.6% 97.0%
dev_public: 1,119 entries, used during development Any issue 97.6% 90.7% 2.1% 98.4%
Fabrication 98.1% 52.7% 1.0% 98.4%

HALLMARK v1.2.3, every entry of both public splits, run on 2026-10-04. Fabrication counts a wrong identifier, a work not found and no author in common; any issue also counts wrong authors, title, year or venue.

  • The held-out split confirms the development numbers: the same precision and two points less recall on entries no rule was ever tuned on.
  • Every flag on a dev_public entry labelled VALID was checked by hand. The 11 that remain are not correct citations: DOIs that belong to other papers, author lists naming people who did not write the paper, a shifted year and a truncated title.
  • Without them, both modes reach 100% precision and 0% false positives. The list, each item with a reason one lookup confirms, is in evals/hallmark_disputed.toml.
  • What is still missed: invented venues on papers known only as preprints (an arXiv record cannot contradict a venue) and author lists that merely leave people out. See evals/results/ for every hallucination type.

Precision comes first: a reference is called fabricated only on positive evidence, and an unanswered or ambiguous lookup is reported as "cannot determine", never as "not found". The evaluation harness and every run's summary are in evals/.

Quick start

With uv nothing needs installing (or pip install paper-preflight):

uvx paper-preflight check path/to/paper

path/to/paper is the project directory, its main .tex file, or a single .bib file. A project that ships no .bib, as many arXiv sources do, is read from its compiled .bbl (checked, but never edited).

No LaTeX at all? A reference list as plain text works too, in the common styles (APA, IEEE, ACM, Nature, Vancouver, Springer, Elsevier, Chicago, MLA), one reference per line, per paragraph or numbered:

uvx paper-preflight check references.txt
pbpaste | uvx paper-preflight check -      # or from stdin

Any arXiv paper, by its ID: the source is downloaded to a temporary folder, checked, and deleted.

uvx paper-preflight check arxiv:2607.06922

Only the PDF? Its reference list is read too, with the pdf extra:

uvx --from 'paper-preflight[pdf]' paper-preflight check paper.pdf
Option Effect
--format json / --format sarif Machine-readable output (SARIF works with GitHub code scanning)
--offline Never touch the network; use only answers already in the local cache
--refresh Ask every source again instead of using cached answers (after a correction, say)
--fail-on warning Make warnings fail the run too (the default is errors)
--lang zh Chinese messages (also chosen automatically from your locale)

Exit codes:

Code Meaning
0 Nothing at or above --fail-on was found
1 Blocking findings
2 No blocking findings, but a source was unavailable, so the paper cannot be called clean yet
3 Usage error

Fetch verified BibTeX

Instead of writing an entry from memory, ask for it by DOI, arXiv ID or title. Every field comes from the registry record, which a comment above the entry names:

paper-preflight bib fetch 1810.04805
% Verified with paper-preflight against dblp (conf/naacl/DevlinCLT19), 2026-10-03
@inproceedings{devlin2019bert,
  title         = {{BERT:} Pre-training of Deep Bidirectional Transformers for Language Understanding},
  author        = {Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
  booktitle     = {NAACL-HLT (1)},
  year          = {2019},
  doi           = {10.18653/v1/n19-1423},
  eprint        = {1810.04805},
  archivePrefix = {arXiv},
}
  • Preprints: an arXiv preprint that has been published comes back as the published version, with its eprint kept (--prefer preprint for the preprint itself).
  • Titles: --title (with --author/--year if needed) lists the candidates instead of choosing when several works match.
  • Retractions: a retracted work comes with a warning.
  • Agents: --format json is for scripts and agents.

Fix the bibliography

bib fix turns findings into edits of your .bib files, taken from the verified records. It prints a diff and changes nothing until you add --apply:

paper-preflight bib fix path/to/paper --level unsafe
--- a/refs.bib
+++ b/refs.bib
@@ -48,7 +47,7 @@
   title     = {Adam: A Method for Stochastic Optimization},
   author    = {Kingma, Diederik P. and Ba, Jimmy},
   booktitle = {International Conference on Learning Representations (ICLR)},
-  year      = {2016},
+  year      = {2015},
 }
  • --level safe (the default) only fixes what cannot change which work is cited: identifiers written so that links break, and DOIs the registry has but the entry lacks.
  • --level unsafe also rewrites authors, title, year and venue from the record, and removes identifiers that point to another work. Review the diff first.
  • Only the affected fields change; comments, formatting, line endings and encoding are kept. A reference nobody could find is never "fixed": only you can say what was meant.

Silence a finding you have checked

A comment directly above an entry silences rules for that entry, with an optional reason:

% preflight: ignore[REF003] reason="internal technical report, not indexed anywhere"
@techreport{lab2024internal,
  ...
}

The verdict stays in the JSON report; only the finding is dropped. A suppression that silenced nothing is reported as CFG001 (info), so stale comments do not pile up. Reference rules are only judged after a complete online run, since offline answers and outages may leave them unrun.

Experimental: find the passage behind each citation

support looks in each cited work for a passage that says what the citing sentence claims. It reads the work's text: the arXiv source, an open-access full text or PDF, or else the abstract. A small local model (HHEM-2.1-open, 0.4 GB) then scores the passages ranked best for the claim.

pip install "paper-preflight[support]"
paper-preflight support path/to/paper --download-model --all

--download-model fetches the model's weights once. --all also lists the confirmed citations with their quotes, and arxiv:<id> works as a target, as it does for check.

  • What it says: "confirmed", with the passage quoted word for word, or "could not confirm". A citation it could not confirm comes with the reason: no text, only the abstract, or no passage close enough.
  • It never calls a citation wrong. On a gold set of 298 citations, 97% of its confirmations were right [95% CI 85%, 99%]. But it confirms only about one real citation in eight, and a low score pointed at a mis-citation less than half the time. The gold set's labels were made by AI models, not experts; see evals/results/support.md.
  • What leaves your machine: the claims are scored locally. Only the cited works' identifiers go out, to fetch their text, which is then kept in the local cache.

Use it from your coding agent

Claude Code — install the plugin. It bundles an MCP server and a skill that makes Claude check the references before calling a paper finished, fix only what is proven wrong, and never invent a reference.

claude plugin marketplace add amos689/paper-preflight
claude plugin install paper-preflight@paper-preflight

Codex, Cursor, VS Code and other MCP clients — run paper-preflight mcp. The tools are read-only and confined to your workspace; see docs/mcp.md.

pre-commit — check citation keys and cached verdicts on every commit in seconds; see docs/pre-commit.md.

GitHub Actions — uses: amos689/paper-preflight@main checks the paper on every push, with the report in the job summary and optional code-scanning alerts; see docs/github-action.md.

Better results with free credentials

paper-preflight works without any account. These optional environment variables make it faster and more complete; their values are never printed or logged.

Variable Effect
PAPER_PREFLIGHT_EMAIL Crossref's polite pool: faster, more reliable lookups
OPENALEX_API_KEY A larger OpenAlex budget for retraction checks
S2_API_KEY Semantic Scholar as a rescue source for references nobody else found

paper-preflight doctor shows which are set and whether each source answers right now.

How it works

  1. Source-first. It reads the LaTeX project as LaTeX sees it: comments, \iffalse blocks and \includeonly are respected, .aux files are used when they are fresh, and the first definition of a duplicated key wins, as in BibTeX.
  2. Identifier-first routing. DOIs go to their registration agency (doi.org tells which: Crossref, DataCite, …). arXiv IDs go to arXiv, with DataCite as a fallback, and PMIDs and PMCIDs to PubMed (which also marks retracted articles). Entries without identifiers are searched by title in dblp and Crossref.
  3. Field-by-field matching with guards. It compares titles (including earlier arXiv version titles), authors (tolerating transcriptions such as Reiß/Reis), year and venue. A search result is used only when enough of these agree and no other work fits as well; known fake DOI copies are skipped.
  4. One verdict per reference: verified, metadata mismatch, identifier conflict, not found, or cannot determine with a reason. "Not found" needs every required source to answer "no".
  5. No LLM anywhere in the verdict. Answers are cached locally (SQLite), so re-runs are fast and --offline works.

Design principles

  • Positive confirmation or abstain. Rate limits, outages and unindexed works lead to "cannot determine", never to "not found".
  • Neutral wording. Findings state observations ("not found in Crossref, dblp and Semantic Scholar, and every source responded"), never accusations.
  • Local-first, no telemetry. Only the metadata of the cited works (DOIs, titles, authors) is sent to the public scholarly APIs above. Your manuscript never leaves your machine.

What it will never do

Help evade plagiarism or AI-text detection, scrape paywalled or bot-protected sites, recommend or "complete" references from memory, or name and shame authors.

Roadmap

  • Done: releases on PyPI (v0.1); references from a .bbl, plain text, a PDF or an arXiv ID (v0.2); an experimental evidence finder for citations, support (v0.3)
  • Next: catch more of what is still missed (partial author lists, invented venues), each round measured on a new week of real papers
  • Later: Chinese-language references

Progress is tracked in docs/PROGRESS.md (in Chinese) and the changelog.

Contributing

Bug reports with a reproducible .bib entry are the most valuable contribution, especially false positives. See CONTRIBUTING.md.

License

MIT. See THIRD_PARTY_NOTICES.md for adapted code.

Metadata

Release files for paper-preflight 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for paper-preflight 0.3.0
File Size Uploaded
paper_preflight-0.3.0.tar.gz 1.3 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for paper-preflight 0.3.0
File Interpreter ABI Platform
paper_preflight-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.5 MB

Release files / paper_preflight-0.3.0.tar.gz

Download URL paper_preflight-0.3.0.tar.gz
Size 1.3 MB
Tags Source
SHA-256 checksum
How to use checksums
6ab4cb8a72e2f5465690c3a73758ca9d505f4343e357947e5b67f79ac9b40f73
BLAKE2b-256 checksum
How to use checksums
21efdc5fec1a9f34e45453d014d1cc8404f730749d739605180fb15103c8a7ab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / paper_preflight-0.3.0-py3-none-any.whl

Download URL paper_preflight-0.3.0-py3-none-any.whl
Size 196.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b2a1e43fe45e50423816c0914a4399c210dfdc5f840a52a95790505e15270b00
BLAKE2b-256 checksum
How to use checksums
26bcc9fbc90853e4222be65d0e685a33e8ddc27e1a8a84bc8256107c976f70c3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page