pydfdoi
pydfdoi is a small Python package for DOI-aware PDF cleanup. It extracts DOI
values from PDF metadata or PDF text, writes DOI-related PDF metadata, and can
rename academic PDFs from Crossref citation data.
Author: Yihtsy yihtsy@outlook.com
Features
- Extract DOI values from existing PDF metadata and the first pages of text.
- Write
/doiand/doiURLinto PDF metadata. - Set
/Titleto the file name for metadata-only cleanup. - Configure PDFs to open on the first page.
- Ask PDF viewers to show the file name in the title bar.
- Query Crossref metadata for DOI records.
- Rename journal article PDFs as
Author et al. Year Title.pdf. - Align journal article PDF page labels with Crossref page metadata.
- Capture journal DOI metadata, book DOI metadata, or general DOI metadata.
- Look up common publisher abbreviations and DOI prefixes.
- Install optional Windows context-menu and SendTo shortcuts.
Package Name
The project uses one consistent package name:
- PyPI distribution:
pydfdoi - Python import:
pydfdoi - Console command:
pydfdoi - Compatibility command:
pdfdoi
This avoids the common mismatch where a package is installed with a hyphenated name but imported with an underscore.
Installation
After the package is published to PyPI:
python -m pip install pydfdoi
For local development from a cloned repository:
C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -U pip
python -m pip install -e ".[dev]"
Editable installation means changes under src\pydfdoi are immediately used by
the command line and tests. Do not develop directly inside Python's
site-packages; keep the source in a normal project folder and use editable
installation.
Command Line
Update DOI metadata only:
pydfdoi --mode metadata-only "D:\path\paper.pdf"
Rename one or more PDFs from Crossref citation metadata:
pydfdoi --mode rename "D:\path\paper1.pdf" "D:\path\paper2.pdf"
Scan more pages for DOI text:
pydfdoi --mode metadata-only --max-pages 5 "D:\path\paper.pdf"
The command-line default is 20 pages, which is friendlier to ebooks whose DOI often appears on the title or copyright pages rather than the first page.
Preview actions without writing changes:
pydfdoi --mode rename --dry-run "D:\path\paper.pdf"
Print JSON output:
pydfdoi --mode metadata-only --json "D:\path\paper.pdf"
Align PDF page labels with Crossref journal article page metadata:
pydfdoi --mode page-labels "D:\path\article.pdf"
If the PDF has extra non-article pages, they are labeled with skip- by
default. Extra pages are assumed to be at the front unless you choose another
placement:
pydfdoi --mode page-labels --extra-placement front "D:\path\article.pdf"
pydfdoi --mode page-labels --extra-placement back "D:\path\article.pdf"
pydfdoi --mode page-labels --extra-placement split "D:\path\article.pdf"
You can customize the special label prefix:
pydfdoi --mode page-labels --skipped-label-prefix "nonarticle-" "D:\path\article.pdf"
If no file is passed, the command opens a file picker.
Python API
Metadata-only cleanup
Use capture_and_write_doi_metadata() when you only want to capture DOI
information and update the PDF metadata.
from pydfdoi import capture_and_write_doi_metadata
doi = capture_and_write_doi_metadata("paper.pdf")
This function:
- extracts the DOI from the PDF;
- writes
/doi; - writes
/doiURL; - sets
/Titleto the file name without.pdf; - configures the PDF to open on the first page;
- does not query Crossref;
- does not rename the file.
Rename by citation
from pathlib import Path
from pydfdoi import extract_doi, rename_by_citation
path = Path("paper.pdf")
doi = extract_doi(path)
if doi:
renamed_path = rename_by_citation(path, doi)
Journal, book, and general DOI capture
from pydfdoi import capture_book_doi, capture_doi_metadata, capture_journal_doi
journal = capture_journal_doi("article.pdf")
book = capture_book_doi("book.pdf")
any_work = capture_doi_metadata("unknown.pdf")
capture_journal_doi()accepts Crossref journal-like work types.capture_book_doi()accepts Crossref book-like work types.capture_doi_metadata()is the general DOI metadata entry point.
capture_book_doi() scans more pages by default than the journal helper because
books often put DOI information on title, copyright, or series pages rather than
the first page. Its default is 20 pages:
book = capture_book_doi("book.pdf")
book = capture_book_doi("book.pdf", max_pages=30)
book = capture_book_doi("book.pdf", strict_type=True)
When strict_type=False, which is the default, capture_book_doi() can still
accept a DOI if Crossref returns an unknown work type but the page-count
heuristic says the PDF is book-like. If Crossref clearly identifies the work as
journal-like, it is still rejected by the book helper.
The general entry point can also infer a coarse work category:
result = capture_doi_metadata("unknown.pdf", article_page_limit=80)
print(result.inferred_kind) # "journal", "book", or None
print(result.classification_source) # "page-count", "crossref", or None
print(result.page_count)
By default, capture_doi_metadata() first makes a simple page-count guess:
PDFs with 80 pages or fewer are treated as journal-like, and longer PDFs are
treated as book-like. If Crossref returns a known work type, the Crossref type
overrides the page-count guess.
Publisher helpers
from pydfdoi import PUBLISHERS, publisher_for_doi, publisher_for_name
publisher = publisher_for_doi("10.1038/s41586-020-2649-2")
publisher = publisher_for_name("John Wiley & Sons")
The publisher table is a curated starter list of common publishers, DOI prefixes, abbreviations, and aliases. It is intentionally not exhaustive: DOI prefixes are assigned to registrants and may remain in use after publisher mergers, platform changes, or imprint transfers.
Journal article page labels
Use align_journal_article_page_labels() to align PDF page labels with Crossref
article page metadata.
from pydfdoi import align_journal_article_page_labels
plan = align_journal_article_page_labels(
"article.pdf",
extra_placement="front",
skipped_label_prefix="skip-",
)
For a Crossref page range such as 10-12, a four-page PDF is labeled as:
skip-1, 10, 11, 12
The function can also use Crossref article-number when a conventional page
range is absent. For article-number-only records, labels look like e12345-1,
e12345-2, and so on.
Because Crossref usually does not say whether extra PDF pages are at the front
or the back, extra_placement is explicit:
front: extra pages are before the article content.back: extra pages are after the article content.split: extra pages are split between front and back.
Windows Context Menu
The repository includes compatibility scripts for Windows Explorer workflows.
Run PowerShell in the project folder:
.\install_context_menu.ps1
This adds two PDF right-click entries:
PDF DOI: rename by citationPDF DOI: metadata only
It also installs SendTo shortcuts, which are usually the most reliable route for processing multiple selected PDFs in Windows Explorer.
To uninstall:
.\install_context_menu.ps1 -Uninstall
pdf_doi_tool.py is kept as a compatibility wrapper for these scripts and for
older local workflows.
Development
Project layout:
src/
pydfdoi/
__init__.py
citation.py
cli.py
crossref.py
doi.py
metadata.py
page_labels.py
pdf.py
processing.py
publishers.py
tests/
Run tests:
C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m pytest
Build local distributions:
C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m build
Check built distributions before publishing:
C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m twine check dist/*
Publisher-specific observations and heuristic notes live in docs/development-notes.md.
Release Notes
Current version: 0.1.1
See CHANGELOG.md for the full version history.
Publishing
Before publishing a new release:
- Update
versioninpyproject.toml. - Update
__version__insrc/pydfdoi/__init__.py. - Add a new entry to
CHANGELOG.md. - Run
python -m pytest. - Run
python -m build. - Run
python -m twine check dist/*. - Tag the release in Git.
- Upload to PyPI with
python -m twine upload dist/*.
For the first PyPI upload, create an account-wide API token on PyPI and run the local helper:
.\scripts\publish_pypi.ps1 -SkipBuild
The helper prompts for the token without echoing it, sets Twine credentials only
for that process, uploads the built 0.1.0 distributions, and clears the
temporary environment variables afterward.
License
pydfdoi is released under the MIT License. See LICENSE.
Metadata
Release files for pydfdoi 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pydfdoi-0.1.1.tar.gz | 20.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pydfdoi-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 39.8 kB
Release files / pydfdoi-0.1.1.tar.gz
| Download URL | pydfdoi-0.1.1.tar.gz |
|---|---|
| Size | 20.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
629827ea1e4327d500a33165dfe58b102ccc8ba9b3eef295cf940b9991955e31
|
|
BLAKE2b-256 checksum How to use checksums |
0d7a66c36b00e0f29807c947f01516e460c2e9c2c5e98a4049d45ecbc070fa54
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|
Release files / pydfdoi-0.1.1-py3-none-any.whl
| Download URL | pydfdoi-0.1.1-py3-none-any.whl |
|---|---|
| Size | 19.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8650fffb02fe72d3c443bb17d26d07f9df3445cefd67d5df9ea59bc47077c28f
|
|
BLAKE2b-256 checksum How to use checksums |
338a34297a6bcb376eaa16fc8980bc7e2ea724405bb8cc7f1165f7abd451b343
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|