Skip to main content

sourcelens sourcelens logo

sourcelens collects the latest research on a topic and keeps it in one chronological catalogue. Give it a topic and it gathers five kinds of material:

  • papers and preprints from PubMed, Europe PMC, arXiv and OpenAlex;
  • the open-access full texts, as PDF, Markdown and plain text;
  • the code repositories and software packages those papers link to;
  • the websites and databases of the field;
  • an exact reference for every paper, ready to export to Zotero or LaTeX.

Run it again later and it adds only what is new. Nothing you have edited in the catalogue is ever overwritten.

pip install sourcelens
sourcelens "CRISPR base editing"

Contents

Install

sourcelens runs on Linux and macOS. There are two implementations of the same command. Install one of them.

Python (3.10 or newer):

pip install sourcelens            # or: pipx install sourcelens
pip install "sourcelens[pdf]"     # adds PyMuPDF for better PDF to Markdown conversion

Go (a single static binary, about 8 MB):

go install github.com/bhagesh-h/sourcelens/cmd/sourcelens@latest

Prebuilt binaries for Linux and macOS are attached to each release.

Docker:

docker build -t sourcelens https://github.com/bhagesh-h/sourcelens.git
docker run --rm -v "$HOME/sourcelens:/data" sourcelens "CRISPR base editing"

Optional:

  • poppler-utils (apt install poppler-utils, brew install poppler) converts PDFs to text and Markdown. The Go binary needs it for that; the Python package uses it when PyMuPDF is not installed.
  • gh, the GitHub CLI: when it is logged in, sourcelens uses its token for faster GitHub searches.

Quick start

  1. Tell sourcelens your email address. Crossref, OpenAlex and NCBI ask for one from automated clients, and Unpaywall needs it to find open-access PDFs.

    sourcelens config set contact_email you@example.org
    
  2. Choose where catalogues are stored. The default is ~/sourcelens.

    sourcelens config set output ~/Documents/research
    
  3. Start a catalogue:

    sourcelens "CRISPR base editing"
    

    This searches all sources for work published in the last 12 months. It then downloads the open-access full texts, finds linked code and websites, and writes the catalogue to ~/Documents/research/crispr-base-editing/. At the end it prints the newest additions and the path of progress.csv.

    A first run takes a few minutes for a narrow topic. Large topics, or runs that download thousands of PDFs, take longer.

  4. Later, fetch what has been published since:

    sourcelens "CRISPR base editing"
    
  5. Look at the results, or export them:

    sourcelens query --topic "CRISPR base editing" --range 1m
    sourcelens export --topic "CRISPR base editing" --format BIB --out crispr.bib
    

Faster, smaller first runs:

sourcelens "CRISPR base editing" --types papers           # metadata only, no downloads
sourcelens "CRISPR base editing" --range 1m               # only the last month
sourcelens "CRISPR base editing" --dry-run                # show what would run
you type sourcelens searches for
sourcelens "CRISPR base editing" titles or abstracts containing CRISPR, base and editing
sourcelens '"base editing"' the exact phrase "base editing"
sourcelens '"base editing", "prime editing"' either phrase
sourcelens "graph neural networks OR GNN" all three words, or GNN

Rules:

  • Commas, semicolons and OR separate alternatives.
  • Within an alternative, every word is required.
  • Double quotes keep a phrase together.

The terms are written to the catalogue's configuration file. Edit that file to refine the search, for example to add synonyms or a second required term (see Configuration).

The time window:

  • A new topic starts 12 months back. --range 5y or --from 2015 start it earlier.
  • Asking later for older work extends the window.
  • Each later update searches the whole window up to today, and adds only what the catalogue does not already contain.

Commands

command what it does
sourcelens "TOPIC" same as sourcelens update --topic "TOPIC"
update search, download and rebuild a catalogue (the default topic when no --topic is given)
retry download again the full texts that were rate-limited, incomplete or failed
query filter a catalogue; print the rows or save them as CSV
export write catalogue entries as references
status size, full-text coverage, recent runs and failures of a catalogue
list the catalogues in the output folder
config show or change the settings of this machine
sources the values that --sources, --types, --paper-types and --format accept
test check, on a throwaway catalogue, that updates only ever add
version print the version

sourcelens help COMMAND or sourcelens COMMAND --help lists every option of a command.

Choosing a catalogue

Every command that works on a catalogue takes:

option meaning
--topic TEXT the catalogue of that topic (default: the default topic)
--dir DIR the catalogue in DIR, anywhere on disk

update

option values default
--range span back from --to: 1d 7d 2w 1m 6m 1y 5y the catalogue's window
--from, --to 2024, 2024-03, 2024-03-15, today, or a span (--from 2y) window start, today
--types papers (metadata), pdf, md, txt, xml, repo, website, all all
--sources pubmed europepmc arxiv openalex local pmc biorxiv unpaywall github cran bioconductor pypi zenodo websites, or groups literature, fulltext, repos, packages all
--paper-types article review preprint report thesis "conference paper" "book chapter" (full texts only) all
--tiers full texts for landmark,core,related rows all three
--scope search groups focused, broad, all all
--workers parallel full-text downloads 12
--parallel steps run at the same time 5
--min-stars GitHub search hits need this many stars 3
--heartbeat seconds between progress lines (0: none) 120
--stop-on-error stop after the first failed stage
--dry-run print the plan and a full-text estimate

The steps of an update run in stages:

stage steps (run in parallel)
1 searches (PubMed, Europe PMC, arXiv, OpenAlex), local reference folders, GitHub search
2 metadata for DOIs found in reference folders
3 catalogue build
4 citation counts and exact dates, preprint to journal links, full texts
5 catalogue refresh
6 links to code, data and websites found in the papers
7 repositories and packages, websites
8 final catalogue build
9 summary of findings

query

Filters combine with AND. A comma list within one filter combines with OR.

sourcelens query --range 1m                                     # published in the last month
sourcelens query --type review --sort cited_by --limit 20       # most-cited reviews
sourcelens query --title "base editor" --from 2024
sourcelens query --fulltext "off-target" --type article         # search inside the downloaded full texts
sourcelens query --added-since 2026-10-01 --out new.csv         # what the last runs added
filter matches
--text RE, --title RE regular expression (any case) over title, entities, venue, matched groups and details, or the title only
--doi LIST, --uid LIST a comma list, or @file with one per line
--from, --to, --range publication date
--tier, --type, --category exact values
--modality, --entity, --species parts of the column value
--added-since DATE rows that entered the catalogue on or after DATE
--has-fulltext rows with a downloaded Markdown full text
--fulltext RE regular expression searched in the downloaded full texts; prints the matching passage
--sort date, cited_by (highest first), title, author

query prints 50 rows (--limit N changes that). --out FILE saves every matching row as CSV.

export

Takes the same filters as query, plus --format with one or more three-letter codes:

code format code format
APA APA 7th IEE IEEE
AMA AMA 11th NAT Nature
MLA MLA 9th BIB BibTeX
CHI Chicago author-date RIS RIS (Zotero, Mendeley, EndNote)
HAR Harvard ENW EndNote tagged
VAN Vancouver CSL CSL-JSON (Zotero, pandoc)
sourcelens export --topic "CRISPR base editing" --type review --format APA
sourcelens export --doi 10.1038/s41586-019-1711-4 --format BIB,RIS --out anzalone
sourcelens export --from 2025 --tier core --format CSL --out core2025.json

--out name.ext writes one format to that file. --out name writes one file per format (name.apa.txt, name.bib, ...). Relative paths go to the catalogue's exports/ folder. Without --out, the references are printed.

Where the results go

Every catalogue is a folder:

<catalogue>/
  progress.csv               the catalogue, oldest first
  config/sourcelens.yaml      its configuration (edit freely)
  fulltext/<year>/<id>/      paper.pdf, paper.md, paper.txt, paper.jats.xml, metadata.json
  fulltext/fulltext_index.csv
  corpus/                    every harvested record with abstract (records.jsonl.gz),
                             references.csv (full author lists, for export), search hits,
                             citation counts, preprint links, broad-search hits
  repos/  websites/  seeds/  intermediate tables of those steps
  changelog/                 added_<date>.csv for every day with new rows, runs.csv
  summary/findings.md        growth per year, most-cited work, tools, websites, coverage
  exports/                   files written by query and export
  logs/cli_<stamp>/          plan, one log per step, summary, the configuration used

The output folder holds the default topic's catalogue directly, and one subfolder per other topic.

progress.csv

One row per resource, sorted by date.

column meaning
added_on the day the row entered the catalogue; never changes
date, year publication date; for repositories the creation date; for packages the first release; for websites the first Internet Archive snapshot
resource_type article, review, preprint, report, thesis, conference paper, book chapter, dataset, repository, software package, archive, database, web calculator, website
tier landmark (key papers named in your reference folders), core (topic in the title, or a method, benchmark, review or software paper), related (topic in the abstract only)
category, modality, entities, species from the classification rules of the configuration
title, authors, venue, doi, pmid, pmcid, url bibliographic data (first three authors; all of them in corpus/references.csv)
code_links, related code and data links found in the paper; preprint and journal versions of the same work
open_access, license, fulltext_status, fulltext_pdf, fulltext_md, fulltext_txt, metadata_file what was downloaded, and where (paths relative to the catalogue)
cited_by highest citation count from OpenAlex, Europe PMC or Crossref
details, matched_groups, found_by, local_refs repository stars and languages; which searches found the row; which reference folder mentions it
status empty, undated, or no longer matched by pipeline (kept)
uid stable key: doi:..., pmid:..., pmcid:..., epmc:... or url:...
notes, user_tags yours; sourcelens never changes them

Updates only add

  • A row whose uid is already in the catalogue keeps its added_on, notes and user_tags. Derived columns are refreshed.
  • A row the searches no longer find is kept and marked in status.
  • Every update writes the new rows of the day to changelog/added_<date>.csv.
  • Full texts already on disk are skipped. Rate-limited, incomplete and failed downloads are retried on the next run; papers without an open copy are retried after 30 days.
  • Only one update runs on a catalogue at a time (.pipeline.lock).
  • sourcelens test checks these rules on a throwaway catalogue.

Settings

Settings belong to the machine, not to a catalogue:

sourcelens config                                   # show them (keys are masked)
sourcelens config set output ~/research             # where catalogues are stored
sourcelens config set contact_email you@example.org
sourcelens config set ncbi_api_key KEY              # optional: faster PubMed
sourcelens config set github_token TOKEN            # optional: or log in with the gh CLI
sourcelens config unset github_token
setting environment variable used for
output SOURCELENS_OUTPUT folder that holds the catalogues (default ~/sourcelens)
contact_email SOURCELENS_EMAIL polite access to Crossref, OpenAlex and NCBI; required for Unpaywall
github_token GITHUB_TOKEN GitHub search rate limit
openalex_api_key OPENALEX_API_KEY OpenAlex premium access
ncbi_api_key NCBI_API_KEY PubMed rate limit

The settings file is ~/.config/sourcelens/settings.yaml on Linux and ~/Library/Application Support/sourcelens/settings.yaml on macOS. It is readable by you only. Environment variables override it.

Configuration of a catalogue

Each catalogue has its own config/sourcelens.yaml, created on its first update. It has four sections:

section sets
search time window, sources, search groups and their terms
classify rules that fill the category, modality, entities, species and tier columns
repos GitHub searches, a relevance pattern for repositories, package searches
websites curated websites and a relevance pattern for sites cited by papers

Top-level keys name the topic and list local reference folders. sourcelens mines those folders for DOIs and links; the papers they cite become seeds of the catalogue.

A configuration created from a topic searches all four literature sources. It has general classification rules: review, commentary, correction, method development, benchmark, software, trial.

Add your own:

  • entities: names of methods, models or measures to tag;
  • modality: data types to track;
  • websites.sites: websites to include.

Every key is described in docs/configuration.md.

Sources

source used for
PubMed (NCBI E-utilities) search, biomedical literature
Europe PMC search (preprints and records outside MEDLINE), full-text XML
arXiv search, PDFs
OpenAlex search across all fields, citation counts, exact dates, open-access PDF links
Crossref, DataCite metadata for DOIs found in reference folders
PMC open-access bucket (AWS) full texts: JATS XML, text, PDF, licence
bioRxiv / medRxiv full texts of preprints, links from preprints to journal versions
Unpaywall open-access PDFs (needs a contact email)
GitHub, CRAN, Bioconductor, PyPI, Zenodo repositories and packages
Internet Archive first-seen dates of websites

OpenAlex matches word stems, so sourcelens keeps an OpenAlex result only when the topic words appear as whole words in its title or abstract. It also skips further versions of a work already in the catalogue: the same title, venue and year under another DOI, as with Zenodo and figshare versions.

Only open-access copies are downloaded. Requests are spaced per host and retried with back-off. bioRxiv and medRxiv rate limits mark downloads as deferred instead of waiting; sourcelens retry picks them up later.

Python and Go

The two implementations take the same commands and options, write the same files and print the same text. Use whichever is easier to install. Both always run their own steps; they share no code.

Python Go
install pip install sourcelens go install or a release binary
source src/sourcelens/ cmd/sourcelens/
PDF conversion PyMuPDF if installed, else poppler poppler

Known differences:

  • Markdown made from PDFs differs in detail between PyMuPDF and poppler.
  • Some publisher sites (for example nature.com) answer Go's HTTP client with a bot check instead of the PDF. sourcelens does not try to get past such checks.
  • The keys inside corpus/records.jsonl.gz may be in a different order. Both read either file.

Both implementations can work on the same catalogue, one after the other.

Development

git clone https://github.com/bhagesh-h/sourcelens.git && cd sourcelens
make python        # pip install -e ".[dev]"
make go            # bin/sourcelens
make test          # go vet, go test, pytest, and sourcelens test for both
make parity        # run parity/cases.txt through both and compare every output

parity/check.sh runs about 70 commands through both implementations:

  • help texts, plans and errors;
  • queries and exports of a real catalogue;
  • a new topic created and built offline.

It then compares stdout, stderr, exit codes and every file written. A change to one implementation is finished when the parity check passes.

Releasing: set the version in src/sourcelens/__init__.py and cmd/sourcelens/main.go, then push a tag such as v0.0.2. Two workflows run:

  • publish uploads the Python package to PyPI;
  • release attaches the Go binaries for Linux and macOS to a GitHub release.

publish.md has the one-time PyPI setup and the release steps, including a trial upload to TestPyPI.

Limitations

  • Coverage. PubMed, Europe PMC, arXiv and OpenAlex are searched. Google Scholar, Scopus and Web of Science have no open API and are not.
  • Matching. Searches match title and abstract. A topic word that appears only in the full text is not found.
  • Classification. Categories, tiers and tags come from regular expressions. They are useful for sorting and filtering, but counts derived from them are approximate.
  • Full texts. Only open-access copies. Paywalled papers are listed without a full text.
  • Websites. A website's date is its first Internet Archive snapshot, which can be older than its current purpose.
  • Platforms. Linux and macOS (file locking uses flock).

Default topic: aging clocks

sourcelens was built to follow research on measuring biological aging. That catalogue is the default topic: sourcelens update without --topic builds it.

The default topic covers:

  • epigenetic and DNA methylation clocks, and named clocks such as Horvath, Hannum, PhenoAge, GrimAge and DunedinPACE;
  • clocks built from omics, imaging and clinical data;
  • biological age and aging scores, frailty indices and methylation risk scores;
  • biomarkers of aging, and forensic age estimation.

It starts in 2011. Broad searches for aging and frailty are kept as metadata only. The reference folders of the author's projects provide landmark papers: a clock registry with the origin paper of each clock, and curated reading lists.

State of that catalogue on 2026-10-04:

count
rows in progress.csv 27,926
papers 19,901 articles, 3,724 reviews, 3,415 preprints, 46 reports
tiers 157 landmark, 13,716 core, 14,053 related
code and data 654 repositories, 27 packages, 46 archives, 43 datasets
websites, databases, calculators 69
full texts 15,999 complete, 377 partial, 481 deferred

Findings, from its summary/findings.md:

  1. Growth. Catalogued papers rose from 255 in 2011 to 4,495 in 2025. Preprints have made up 13 to 17% of each year since 2020.
  2. Epigenetic clocks came in waves. These dates are from the origin papers of the 178 clocks in the registry:
    • 2011 to 2014: chronological-age predictors (Hannum, Horvath);
    • 2016 to 2019: tissue- and species-specific clocks;
    • 2017 to 2022: clocks trained on mortality and morbidity (PhenoAge, GrimAge, DunedinPACE);
    • 2022 onward: reliability-focused and deep-learning clocks;
    • 2025 onward: system-level clocks.
  3. Other data layers. Clocks built from proteomic, multi-omic and clinical-biomarker data grew fastest after 2023.
  4. Frailty indices form the largest single block of the catalogue, mostly as clinical outcome measures.
  5. Openness. 59% of catalogued papers have an open full text.

License

GPL-3.0. See LICENSE.

Metadata

Release files for sourcelens 0.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sourcelens 0.0.1
File Size Uploaded
sourcelens-0.0.1.tar.gz 125.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sourcelens 0.0.1
File Interpreter ABI Platform
sourcelens-0.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 264.8 kB

Release files / sourcelens-0.0.1.tar.gz

Download URL sourcelens-0.0.1.tar.gz
Size 125.5 kB
Tags Source
SHA-256 checksum
How to use checksums
059843d2ec7cf64ceab15d7755e15442889ec2c59b7d2730531b0b8125f2b45a
BLAKE2b-256 checksum
How to use checksums
e642d8a13308f1b1b94f57a3dd7c51d3091647e5a78636c013a96553f54bcaed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / sourcelens-0.0.1-py3-none-any.whl

Download URL sourcelens-0.0.1-py3-none-any.whl
Size 139.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
13b7757202098b327e0478a7b1c7cfad96efe27e9fe6fff237fb846b388588c8
BLAKE2b-256 checksum
How to use checksums
d641d69c61bbba1282a15e6ab997c1bcb01164d3fe56dcb45a69fb95d685147c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

1.0.1

2 release files

1.0.0

2 release files

This release

0.0.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page