Skip to main content

citeguard

Retracted papers get cited for years after retraction. Andrew Wakefield's fraudulent 1998 paper linking the MMR vaccine to autism — retracted in 2010 — has been cited well over a thousand times since its retraction, by researchers who had no easy way to know. A 2026 JMIR study found that freely available AI tools "cannot reliably flag retracted literature," right as AI-assisted research and writing has exploded. citeguard closes that gap: check a citation against Crossref — free, no API key — before it goes in a paper, a summary, or a review.

Three surfaces, one verified detection algorithm:

Surface For Location
Python library + CLI academic writers, scripts src/citeguard/
MCP server AI agents doing research/writing mcp-server/
GitHub Action CI on a lab's or journal's repo action.yml

Why this is built the way it is

The detection logic isn't guessed at from documentation — it was built by querying Crossref's real API for known cases and reading the actual response shapes, then writing tests against the saved real responses (committed in tests/fixtures/, shared by both the Python and TypeScript implementations). Two real papers anchor the ground truth:

  • Wakefield et al., 1998, The Lancet (the MMR-autism paper): the publisher's own Crossref metadata has no structured retraction data — an update-to check alone would silently miss it. What actually carries the signal is a separate updated-by field, where Crossref has backfilled the retraction (and an earlier 2004 correction) from the Retraction Watch database itself. Missing this field would have meant missing the single most famous retracted paper in medicine.
  • Mehra et al., 2020, The Lancet (the Surgisphere-linked COVID/hydroxychloroquine paper): here the publisher did attach structured data directly, via update-to. A different field, a different provenance, same underlying fact.
  • Watson & Crick, 1953, Nature serves as the clean control in every test suite — a definitely-real, definitely-not-retracted paper that must never be flagged.

So the checker looks at three independent signals, in order: update-to (publisher-asserted), updated-by (often Retraction-Watch-sourced, catching what publishers miss), and a title-prefix fallback ("RETRACTED:", "WITHDRAWN:", etc.) for older or unlinked cases with no structured metadata on either field. Each signal is mapped to the right severity — an "Expression of Concern:" title is not the same thing as a retraction, and earlier versions of this code collapsing that distinction was itself a bug caught by testing against real titles, not just synthetic ones. See src/citeguard/analyze.py for the fully-commented implementation.

Quick start

CLI:

pip install git+https://github.com/wedo911/citeguard.git
citeguard doi 10.1016/S0140-6736(97)11096-0
citeguard file references.bib --fail-on retracted   # for CI, see below

(Installing from git rather than PyPI: the name citeguard on PyPI belongs to an unrelated project. Use pip install -e . if you've cloned this repo.)

MCP server (add to your client's config, e.g. .mcp.json):

{ "mcpServers": { "citeguard": { "command": "npx", "args": ["-y", "citeguard-mcp-server"] } } }

Published on npm as citeguard-mcp-server — nothing to clone or build. To run it from a local checkout instead, point command at node and args at mcp-server/dist/index.js.

GitHub Action — on the GitHub Marketplace. Add to any repository that holds a manuscript and its bibliography:

- uses: wedo911/citeguard@v0.1.1
  with:
    path: references.bib
    fail-on: concern   # never | retracted | concern | corrected

What this is not

  • Not proof a paper's content is correct. It only checks retraction status, not whether a non-retracted paper's findings hold up.
  • Not exhaustive. The title-prefix heuristic only catches the publisher conventions it's been tested against; a clean result means "no known signal found," not "guaranteed never retracted."
  • Not a bulk-scraping tool. It's built for the size of a real bibliography (tens of citations), with a small fixed delay between Crossref requests and an optional persistent cache (src/citeguard/cache.py) — good API citizenship for a free public service, not a tool for scanning millions of DOIs.

Running the tests

# Python (55 tests, including against the real fixtures above)
pip install -e ".[dev]" && pytest -v

# MCP server (17 tests against the same real fixtures)
cd mcp-server && npm install && npm run build && npm test

Both suites are network-free and deterministic — they run against the committed real API responses, not live calls, so they're fast and don't depend on Crossref being reachable. Live end-to-end behavior (the actual CLI, the actual MCP tool, hitting the real API) was separately verified by hand during development; the GitHub Action additionally has its own CI job (action-smoke-test) that runs the real composite action against a known-retracted and a known-clean bibliography on every push, so the Action itself — not just the underlying library — is continuously verified against the live API.

Contributing

New signal types, additional publisher title conventions, and false-positive reports are all welcome. If you add a case, prefer adding it as a real, cited Crossref fixture over a synthetic one where possible — that's what caught the two real bugs this project's own test suite found during development (a URL-encoding bug in the DOI request path, and a BibTeX parser that could swallow an adjacent entry when parsing a malformed @comment block).

License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

citeguard_cli-0.1.0.tar.gz (22.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

citeguard_cli-0.1.0-py3-none-any.whl (16.0 kB view details)

Uploaded Python 3

File details

Details for the file citeguard_cli-0.1.0.tar.gz.

File metadata

  • Download URL: citeguard_cli-0.1.0.tar.gz
  • Upload date:
  • Size: 22.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.10

File hashes

Hashes for citeguard_cli-0.1.0.tar.gz
Algorithm Hash digest
SHA256 ea91496de2ae01fe1295e1848bbb713887c949fe25fb8742101d9a921012e868
MD5 a65b586c2802433055062a9f7253ebd7
BLAKE2b-256 648c9fe6b6b3d359cb5776db09d6e61f93ef3d3f645a36a887f9f9d44a9f1b0f

See more details on using hashes here.

File details

Details for the file citeguard_cli-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: citeguard_cli-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 16.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.10

File hashes

Hashes for citeguard_cli-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 9b42e2be8d38eda4871797abda2420d8710fb2b1c91363595612aa9c7ec2084e
MD5 28c026707b8a7a08a8e95fd58987625a
BLAKE2b-256 a485a851cda3a02bbeafc3cf5a4d14644de5d384da40c7fee847785d6af5ef8c

See more details on using hashes here.

Release history Release notifications | RSS feed

0.1.1

2 files

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page