Skip to main content

pagespring

PyPI

Acquire and normalize online documentation into clean, convertible source files — the acquisition front-end to pagespeak.

Point it at a manual's URL. A pattern recognizes the source type, acquires the raw pages (stdlib urllib), and normalizes them into ONE clean HTML/markdown file with absolute asset URLs under incoming/<slug>/. That clean file is the deliverable; converting it into the finished RAG corpus is a separate step (pagespeak) that consumes incoming/ on its own — pagespring never runs it.

Lean by design: pf-core[cli] (PyPI) + beautifulsoup4, pyyaml, and pypdfium2 for PDF page counts. Stdlib fetch, no ML stack.

Intended use

pagespring is for publicly available documentation — vendor manuals, help centers, open textbooks, API specs. It fetches only what the source serves to any reader: there is no login/session handling, no paywall traversal, and no bot-detection evasion. It is a polite client: it identifies itself with a pagespring/<version> User-Agent (PAGESPRING_UA overrides it), honors 429 Retry-After, backs off on server errors, paces crawl requests, and caps crawl sizes.

It is a user-invoked, one-manual-at-a-time archiver — closer to "Save Page As" than to an autonomous crawler — so it does not consult robots.txt (which governs bots that discover URLs on their own; you supply the URL). Before mirroring a site, check its terms of use. What you may do with the acquired copy (personal RAG corpus, internal search, redistribution) is governed by the source's license — the deliverable under incoming/ stays on your machine, and nothing is re-published by this tool.

Install

pip install pagespring

Quick start

pagespring ingest https://docs.tableplus.com   # acquire + normalize → incoming/tableplus/
pagespring renormalize <slug>                   # replay normalize from kept raw/ — no re-crawl
pagespring refresh --all                        # re-check every manual against its source
pagespring audit --all                          # $0 sanity checks on everything staged
pagespring localize <slug>                      # pull a deliverable's images later (resumable; --all)
pagespring patterns                             # list the source patterns
pagespring classify <url>                       # which pattern handles a URL (no fetch)
pagespring status                               # what's been acquired

Deliverables land in ./incoming/<slug>/ under the directory you run from.

Dev

bin/setup   # clone → venv + editable install with dev extras
bin/test    # pytest
bin/lint    # ruff check + ruff format --check + mypy (strict) + structural gate + framework-first

See docs/usage.md for the full command set and docs/architecture.md for the acquire → normalize flow and how to add a new source pattern.

License

Apache-2.0 — see LICENSE. Releases through 0.11.0 were published under the MIT license and stay MIT.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pagespring-0.12.0.tar.gz (228.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pagespring-0.12.0-py3-none-any.whl (128.5 kB view details)

Uploaded Python 3

File details

Details for the file pagespring-0.12.0.tar.gz.

File metadata

  • Download URL: pagespring-0.12.0.tar.gz
  • Upload date:
  • Size: 228.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pagespring-0.12.0.tar.gz
Algorithm Hash digest
SHA256 b4e173eabea432816b112d6f6d3ab09e0a5b8896e261aa8116b385c4509beb52
MD5 dae6ce18b5b3d24cb7ab909d46bfeee3
BLAKE2b-256 b3a97277f91a39da92f3d71fac16ec06d6ff6ae8a6089269baecc289c4f568ec

See more details on using hashes here.

Provenance

The following attestation bundles were made for pagespring-0.12.0.tar.gz:

Publisher: publish.yml on phierceweb/pagespring

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pagespring-0.12.0-py3-none-any.whl.

File metadata

  • Download URL: pagespring-0.12.0-py3-none-any.whl
  • Upload date:
  • Size: 128.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pagespring-0.12.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a0d6f05cf69c1256a46f9f2d9997b23147a428a78de02b4e9687c4848febb7a2
MD5 371ea0a88c29a423534af674dd7408a0
BLAKE2b-256 c68f73264cc074458760118b29d0f42709b7550b2a78b81ac48c00ef74fd9a4d

See more details on using hashes here.

Provenance

The following attestation bundles were made for pagespring-0.12.0-py3-none-any.whl:

Publisher: publish.yml on phierceweb/pagespring

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.12.0 This release

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page