Skip to main content

pydfdoi

pydfdoi is a small Python package for DOI-aware PDF cleanup. It extracts DOI values from PDF metadata or PDF text, writes DOI-related PDF metadata, and can rename academic PDFs from Crossref citation data.

Author: Yihtsy yihtsy@outlook.com

Features

  • Extract DOI values from existing PDF metadata and the first pages of text.
  • Write /doi and /doiURL into PDF metadata.
  • Set /Title to the file name for metadata-only cleanup.
  • Configure PDFs to open on the first page.
  • Ask PDF viewers to show the file name in the title bar.
  • Query Crossref metadata for DOI records.
  • Rename journal article PDFs as Author et al. Year Title.pdf.
  • Align journal article PDF page labels with Crossref page metadata.
  • Capture journal DOI metadata, book DOI metadata, or general DOI metadata.
  • Look up common publisher abbreviations and DOI prefixes.
  • Install optional Windows context-menu and SendTo shortcuts.

Package Name

The project uses one consistent package name:

  • PyPI distribution: pydfdoi
  • Python import: pydfdoi
  • Console command: pydfdoi
  • Compatibility command: pdfdoi

This avoids the common mismatch where a package is installed with a hyphenated name but imported with an underscore.

Installation

After the package is published to PyPI:

python -m pip install pydfdoi

For local development from a cloned repository:

C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -U pip
python -m pip install -e ".[dev]"

Editable installation means changes under src\pydfdoi are immediately used by the command line and tests. Do not develop directly inside Python's site-packages; keep the source in a normal project folder and use editable installation.

Command Line

Update DOI metadata only:

pydfdoi --mode metadata-only "D:\path\paper.pdf"

Rename one or more PDFs from Crossref citation metadata:

pydfdoi --mode rename "D:\path\paper1.pdf" "D:\path\paper2.pdf"

Scan more pages for DOI text:

pydfdoi --mode metadata-only --max-pages 5 "D:\path\paper.pdf"

The command-line default is 20 pages, which is friendlier to ebooks whose DOI often appears on the title or copyright pages rather than the first page.

Preview actions without writing changes:

pydfdoi --mode rename --dry-run "D:\path\paper.pdf"

Print JSON output:

pydfdoi --mode metadata-only --json "D:\path\paper.pdf"

Align PDF page labels with Crossref journal article page metadata:

pydfdoi --mode page-labels "D:\path\article.pdf"

If the PDF has extra non-article pages, they are labeled with skip- by default. Extra pages are assumed to be at the front unless you choose another placement:

pydfdoi --mode page-labels --extra-placement front "D:\path\article.pdf"
pydfdoi --mode page-labels --extra-placement back "D:\path\article.pdf"
pydfdoi --mode page-labels --extra-placement split "D:\path\article.pdf"

You can customize the special label prefix:

pydfdoi --mode page-labels --skipped-label-prefix "nonarticle-" "D:\path\article.pdf"

If no file is passed, the command opens a file picker.

Python API

Metadata-only cleanup

Use capture_and_write_doi_metadata() when you only want to capture DOI information and update the PDF metadata.

from pydfdoi import capture_and_write_doi_metadata

doi = capture_and_write_doi_metadata("paper.pdf")

This function:

  • extracts the DOI from the PDF;
  • writes /doi;
  • writes /doiURL;
  • sets /Title to the file name without .pdf;
  • configures the PDF to open on the first page;
  • does not query Crossref;
  • does not rename the file.

Rename by citation

from pathlib import Path

from pydfdoi import extract_doi, rename_by_citation

path = Path("paper.pdf")
doi = extract_doi(path)
if doi:
    renamed_path = rename_by_citation(path, doi)

Journal, book, and general DOI capture

from pydfdoi import capture_book_doi, capture_doi_metadata, capture_journal_doi

journal = capture_journal_doi("article.pdf")
book = capture_book_doi("book.pdf")
any_work = capture_doi_metadata("unknown.pdf")
  • capture_journal_doi() accepts Crossref journal-like work types.
  • capture_book_doi() accepts Crossref book-like work types.
  • capture_doi_metadata() is the general DOI metadata entry point.

capture_book_doi() scans more pages by default than the journal helper because books often put DOI information on title, copyright, or series pages rather than the first page. Its default is 20 pages:

book = capture_book_doi("book.pdf")
book = capture_book_doi("book.pdf", max_pages=30)
book = capture_book_doi("book.pdf", strict_type=True)

When strict_type=False, which is the default, capture_book_doi() can still accept a DOI if Crossref returns an unknown work type but the page-count heuristic says the PDF is book-like. If Crossref clearly identifies the work as journal-like, it is still rejected by the book helper.

The general entry point can also infer a coarse work category:

result = capture_doi_metadata("unknown.pdf", article_page_limit=80)

print(result.inferred_kind)           # "journal", "book", or None
print(result.classification_source)   # "page-count", "crossref", or None
print(result.page_count)

By default, capture_doi_metadata() first makes a simple page-count guess: PDFs with 80 pages or fewer are treated as journal-like, and longer PDFs are treated as book-like. If Crossref returns a known work type, the Crossref type overrides the page-count guess.

Publisher helpers

from pydfdoi import PUBLISHERS, publisher_for_doi, publisher_for_name

publisher = publisher_for_doi("10.1038/s41586-020-2649-2")
publisher = publisher_for_name("John Wiley & Sons")

The publisher table is a curated starter list of common publishers, DOI prefixes, abbreviations, and aliases. It is intentionally not exhaustive: DOI prefixes are assigned to registrants and may remain in use after publisher mergers, platform changes, or imprint transfers.

Journal article page labels

Use align_journal_article_page_labels() to align PDF page labels with Crossref article page metadata.

from pydfdoi import align_journal_article_page_labels

plan = align_journal_article_page_labels(
    "article.pdf",
    extra_placement="front",
    skipped_label_prefix="skip-",
)

For a Crossref page range such as 10-12, a four-page PDF is labeled as:

skip-1, 10, 11, 12

The function can also use Crossref article-number when a conventional page range is absent. For article-number-only records, labels look like e12345-1, e12345-2, and so on.

Because Crossref usually does not say whether extra PDF pages are at the front or the back, extra_placement is explicit:

  • front: extra pages are before the article content.
  • back: extra pages are after the article content.
  • split: extra pages are split between front and back.

Windows Context Menu

The repository includes compatibility scripts for Windows Explorer workflows.

Run PowerShell in the project folder:

.\install_context_menu.ps1

This adds two PDF right-click entries:

  • PDF DOI: rename by citation
  • PDF DOI: metadata only

It also installs SendTo shortcuts, which are usually the most reliable route for processing multiple selected PDFs in Windows Explorer.

To uninstall:

.\install_context_menu.ps1 -Uninstall

pdf_doi_tool.py is kept as a compatibility wrapper for these scripts and for older local workflows.

Development

Project layout:

src/
  pydfdoi/
    __init__.py
    citation.py
    cli.py
    crossref.py
    doi.py
    metadata.py
    page_labels.py
    pdf.py
    processing.py
    publishers.py
tests/

Run tests:

C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m pytest

Build local distributions:

C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m build

Check built distributions before publishing:

C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m twine check dist/*

Publisher-specific observations and heuristic notes live in docs/development-notes.md.

Release Notes

Current version: 0.1.1

See CHANGELOG.md for the full version history.

Publishing

Before publishing a new release:

  1. Update version in pyproject.toml.
  2. Update __version__ in src/pydfdoi/__init__.py.
  3. Add a new entry to CHANGELOG.md.
  4. Run python -m pytest.
  5. Run python -m build.
  6. Run python -m twine check dist/*.
  7. Tag the release in Git.
  8. Upload to PyPI with python -m twine upload dist/*.

For the first PyPI upload, create an account-wide API token on PyPI and run the local helper:

.\scripts\publish_pypi.ps1 -SkipBuild

The helper prompts for the token without echoing it, sets Twine credentials only for that process, uploads the built 0.1.0 distributions, and clears the temporary environment variables afterward.

License

pydfdoi is released under the MIT License. See LICENSE.

Metadata

Release files for pydfdoi 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pydfdoi 0.1.1
File Size Uploaded
pydfdoi-0.1.1.tar.gz 20.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pydfdoi 0.1.1
File Interpreter ABI Platform
pydfdoi-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 39.8 kB

Release files / pydfdoi-0.1.1.tar.gz

Download URL pydfdoi-0.1.1.tar.gz
Size 20.8 kB
Tags Source
SHA-256 checksum
How to use checksums
629827ea1e4327d500a33165dfe58b102ccc8ba9b3eef295cf940b9991955e31
BLAKE2b-256 checksum
How to use checksums
0d7a66c36b00e0f29807c947f01516e460c2e9c2c5e98a4049d45ecbc070fa54
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / pydfdoi-0.1.1-py3-none-any.whl

Download URL pydfdoi-0.1.1-py3-none-any.whl
Size 19.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8650fffb02fe72d3c443bb17d26d07f9df3445cefd67d5df9ea59bc47077c28f
BLAKE2b-256 checksum
How to use checksums
338a34297a6bcb376eaa16fc8980bc7e2ea724405bb8cc7f1165f7abd451b343
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page