Skip to main content

Metawarc

Metawarc indexes WARC collections into a versioned DuckDB catalog with Parquet sidecars, then provides bounded tools for querying, exporting payloads, extracting metadata, and analyzing collections. Source archives are always treated as immutable.

The 2.0 implementation replaces filename-derived tables and whole-archive buffers with stable archive IDs, explicit workspace metadata, batched writes, atomic publication, and typed queries shared by the CLI, REST API, and MCP server.

Install

Python 3.10 or newer is required.

pip install metawarc              # core CLI
pip install 'metawarc[api]'       # REST API (includes website replay routes)
pip install 'metawarc[replay]'    # same as api; documents replay intent
pip install 'metawarc[mcp]'       # MCP server
pip install 'metawarc[all]'       # all runtime interfaces
pip install -e '.[all,dev]'       # contributor checkout

Quick start

metawarc index 'archives/**/*.warc*' --dbfile collection.db --resume
metawarc catalog --dbfile collection.db
metawarc stats --dbfile collection.db --mode mimes
metawarc list-files --dbfile collection.db --mimes application/pdf
metawarc dump --dbfile collection.db --exts pdf --limit 100 --output exported

The default sidecar directory for collection.db is collection.data. It may be relocated with --data-dir. Catalog paths are stored relative to that workspace whenever possible, so moving the database and its data directory together remains supported.

Website replay

Local Wayback-style replay is available through the REST server:

pip install 'metawarc[replay]'   # or metawarc[api]
metawarc serve --dbfile collection.db
# Home page:  http://127.0.0.1:8000/
# Replay URL: http://127.0.0.1:8000/replay/<YYYYMMDDHHMMSS>mp_/https://example.com/
metawarc replay --dbfile collection.db   # alias of serve
metawarc export-cdxj --dbfile collection.db -o collection.cdxj --path-index paths.tsv
Path Purpose
/ or /replay HTML index of archived hosts with Open links
/replay/sites JSON list of hosts, entry URLs, and replay paths
/replay/<stamp>/<url> Closest exact-URL capture; HTML/CSS rewritten by default
/replay/<stamp>mp_/<url> Explicit rewritten mode
/replay/<stamp>id_/<url> Raw identity mode (no rewrite or banner)

Capture selection uses the DuckDB catalog (exact URL, closest or exact timestamp). Rewritten HTML/CSS keeps the original charset (UTF-8, windows-1251, KOI8-R, and related declarations) and re-emits UTF-8 so Cyrillic and other non-ASCII text render correctly. The home page prefers an https://host/ 200 response over an http:// redirect when both exist.

JavaScript is not rewritten. For full Wombat/JS fidelity, export CDXJ and point pywb at the original WARCs. Archived scripts may be hostile; keep the default loopback bind unless you configure a token or --allow-insecure.

Incremental operation and recovery

metawarc ingest 'archives/**/*.warc*' --dbfile collection.db --dry-run
metawarc ingest 'archives/**/*.warc*' --dbfile collection.db --resume
metawarc doctor --dbfile collection.db
metawarc doctor --dbfile collection.db --repair       # dry-run repair plan
metawarc doctor --dbfile collection.db --repair --apply

add, update, rescan, and force are explicit index --mode values. Changed archives retain their stable catalog ID. A moved source is reported as a candidate and requires metawarc rebind ARCHIVE_ID NEW_PATH; Metawarc never silently guesses identity.

Metadata and analysis

metawarc index-content --dbfile collection.db --type links --type pdfs
metawarc analyze summary --dbfile collection.db --output summary.json
metawarc analyze metadata --dbfile collection.db --type all --top 20
metawarc analyze hashes --dbfile collection.db --resume
metawarc analyze duplicates --dbfile collection.db --output duplicates.csv --output-format csv
metawarc analyze links --dbfile collection.db --output links.parquet --output-format parquet
metawarc analyze integrity --dbfile collection.db --deep --max-records 1000

Extraction uses MIME, extension, and bounded signature signals. Results include a versioned envelope, normalized metadata, raw parser output, warnings, stable error codes, inspected bytes, and duration. ZIP/XML, payload-size, and time limits are applied before a derived sidecar is published.

Supported content families include Office Open XML documents, templates, macro-enabled files, binary workbooks, and presentations (including PPSX); PDF; GIF, SVG, WebP, icons, bitmap, camera, and editing images; common MP4/QuickTime, AVI, WebM/Matroska, Ogg, MPEG, ASF/WMV, and FLV video; MP3, WAV, AIFF, FLAC, Ogg/Opus, M4A, WMA, MIDI, and RealAudio; and TTF/OTF, font collections, WOFF, WOFF2, and EOT fonts.

analyze metadata reads the stored extraction envelopes for PDF, image, OOXML, OLE, video, audio, and font sidecars. --type all covers those seven types; link metadata continues to use the dedicated analyze links report.

Query and export safety

Normal filtering uses allowlisted fields and bound parameters. Raw SQL is available only through the CLI's visibly named --unsafe-where option; it is not exposed by REST or MCP. Payload exports sanitize record IDs, avoid overwrites, stream in source order, enforce record/byte limits, and write a JSONL manifest with SHA-256 checksums.

REST API and MCP

METAWARC_API_TOKEN='replace-me' metawarc serve --dbfile collection.db
metawarc mcp --dbfile collection.db                    # stdio
metawarc mcp --dbfile collection.db --transport http  # loopback only by default

serve exposes the typed record API and the website replay routes described above. Both network services bind to loopback by default. Non-loopback REST binding requires a bearer token or an explicit --allow-insecure acknowledgement. The MCP surface is read-only and contains no raw SQL, filesystem-path, payload, or mutation tool. Non-loopback MCP transport requires explicit acknowledgement.

See architecture, CLI reference, security, changelog, and the release checklist for operational detail.

Compatibility

  • A 2.0 workspace uses schema version 2 and is opened only by compatible code.
  • Legacy 1.2/1.3 catalogs with files/tables are detected. A backup is made before their paths are migrated into the versioned catalog.
  • If a legacy layout cannot be migrated unambiguously, doctor reports rebuild guidance instead of rewriting source archives.

The canonical repository is https://github.com/datacoon/metawarc; master is the release branch and feature work is integrated through reviewed pull requests. Releases use signed vMAJOR.MINOR.PATCH tags.

License

MIT. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

metawarc-2.0.1.tar.gz (92.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

metawarc-2.0.1-py3-none-any.whl (81.6 kB view details)

Uploaded Python 3

File details

Details for the file metawarc-2.0.1.tar.gz.

File metadata

  • Download URL: metawarc-2.0.1.tar.gz
  • Upload date:
  • Size: 92.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.7

File hashes

Hashes for metawarc-2.0.1.tar.gz
Algorithm Hash digest
SHA256 b7439362e68a422d1cfea5d5f9840e9677803901a1c60c57c60120f32369529a
MD5 df5f8b4388fb8c06959569af3890c262
BLAKE2b-256 8d926045727efe25de7243f4f22a0fcfe9f41ea15b7594f44e428cb5db01c77d

See more details on using hashes here.

File details

Details for the file metawarc-2.0.1-py3-none-any.whl.

File metadata

  • Download URL: metawarc-2.0.1-py3-none-any.whl
  • Upload date:
  • Size: 81.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.7

File hashes

Hashes for metawarc-2.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 e664e65c76441f0d3a779392ff39c308b7b0bb95aa2a608afce46b0096070b1d
MD5 b646bfc1d4d79c39c570af9e3f7b3a80
BLAKE2b-256 94f4618e75021432c251b84fc092cb0487874746cf19b6001560e1a57624688b

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

2.0.1 This release

2 files

2.0.0

2 files

1.3.1

2 files

1.1.1

1 file

1.0.2

1 file

1.0.1

1 file

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page