Skip to main content

Extract DOIs or titles from PDF papers and generate a BibTeX bibliography

Project description

Bib Extractor

PyPI version License: MIT

A tiny, pip‑installable Python utility that scans a folder of PDF papers, extracts their DOI (or a title fallback), and produces a JSON file that can be turned into a BibTeX bibliography.


✨ Features

  • Works on any folder of PDFs.
  • Uses pdftotext (Poppler) to read text from PDFs.
  • Detects DOI strings with a robust regular expression.
  • Multiple API Support: Queries doi.org and falls back to Crossref for metadata.
  • Auto-Rename: Automatically renames PDFs to Year - Author - Title.pdf.
  • Formatted Citations: Generates APA/MLA style reference lists in a separate text file.
  • Visual Progress: Includes a terminal progress bar for high‑volume processing.
  • Zero external Python dependencies (standard library only).

📦 Installation

From PyPI (recommended)

pip install bib-extractor

From source

  1. Install Popplerpdftotext is required.
  2. Clone the repository
git clone https://github.com/msrtarit/bib_extractor.git
cd bib_extractor
  1. (Optional) Create a virtual environment and install the package in editable mode:
python -m venv .venv
.venv\\Scripts\\activate   # Windows
# or source .venv/bin/activate on Unix
pip install -e .

🚀 Usage

Extract DOIs / titles

# Using the installed command (if installed via pip)
bib-extractor --dir path/to/papers --output paper_info.json

# Or run the script directly from the source checkout
python extract_bib_info.py --dir path/to/papers --output paper_info.json
  • --dir defaults to the current working directory.
  • --output defaults to paper_info.json.

Fetch BibTeX entries & Auto-Rename

# Fetch entries and automatically rename your PDFs
bib-fetch --input paper_info.json --output papers.bib --rename --dir path/to/papers
  • --input: The JSON file from the extractor.
  • --output: The destination .bib file.
  • --citations: (Optional) Output file for a formatted reference list (e.g., refs.txt).
  • --style: (Optional) Citation style for the list (apa or mla, default is apa).
  • --rename: (Optional) Automatically renames the files in --dir to a standard format: Year - Author - Title.pdf.
  • --dir: (Required if renaming) The folder where your original PDFs are located.

The extractor prints progress and writes a JSON array like:

[
  {"file": "1.pdf", "doi": "10.1109/XYZ.2023.123456"},
  {"file": "2.pdf", "title": "An Interesting Study on …"}
]

Next steps

  • Convert the generated .bib file to the citation style you need.
  • Extend the workflow with additional scripts or integrate into your bibliography manager.

🤝 Contributing

Please see the CONTRIBUTING.md for guidelines on how to fork the repo, set up a development environment, and submit pull requests.


📜 License

This project is licensed under the MIT License – see the LICENSE file for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

bib_extractor-0.1.2.tar.gz (7.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

bib_extractor-0.1.2-py3-none-any.whl (8.3 kB view details)

Uploaded Python 3

File details

Details for the file bib_extractor-0.1.2.tar.gz.

File metadata

  • Download URL: bib_extractor-0.1.2.tar.gz
  • Upload date:
  • Size: 7.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for bib_extractor-0.1.2.tar.gz
Algorithm Hash digest
SHA256 82ad4c1af5232bac3c9f80205501918cb1387beb8e6bac7db39ab721637c2004
MD5 ecd664717ce6c88b4f7b8b156099e710
BLAKE2b-256 e5de15188522f4234f2433d3e283b1ae224dffa93e9a3c443fd432fde81f41dc

See more details on using hashes here.

Provenance

The following attestation bundles were made for bib_extractor-0.1.2.tar.gz:

Publisher: publish.yml on msrtarit/bib-extractor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file bib_extractor-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: bib_extractor-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 8.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for bib_extractor-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 4fe99b9b98b9309a3a9baf3aa87726e99b9a4ccab285c008f67eb582387e94e3
MD5 b8441e3bff8d529bcffb6fca9f4a0855
BLAKE2b-256 1529dabc128a40192ec1cd58be339936df2b65fb22bdf1cbd91245588f97314b

See more details on using hashes here.

Provenance

The following attestation bundles were made for bib_extractor-0.1.2-py3-none-any.whl:

Publisher: publish.yml on msrtarit/bib-extractor

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page