Skip to main content

filegetter

PyPI version Python Versions License: MIT Tests

filegetter is a command-line tool and Python library for automated bulk file collection from public data sources.


Table of Contents


Overview

Filegetter automates downloading large numbers of files from URLs listed in configuration-driven project files. It was built to streamline file collection from datasets produced by other tools (API scrapers, data extracts), with support for several input formats and storage backends, resumable runs and a detailed per-file report.


Features

✅ Multiple Input Formats - CSV, JSON Lines (JSONL) and plain text lists ✅ Flexible URL Handling - URL prefixes, {}-patterns and absolute URLs ✅ Storage Options - ZIP archive or filesystem directory tree ✅ Resume Capability - Already-downloaded files are skipped; failed files are retried ✅ Integrity & Reporting - CSV report with HTTP status, MIME type, size and SHA-256 checksum ✅ WARC Output - Optional archival WARC/1.0 records (pip install filegetter[warc]) ✅ Robust Downloads - Retries with backoff, timeouts, per-file error isolation ✅ Politeness Controls - Configurable delay, User-Agent, worker count and file size limit ✅ Zip Compression Toggle - compression option for smaller archives or faster writes


Installation

Requirements

  • Python 3.9 or higher

Install from PyPI

pip install --upgrade pip
pip install --upgrade filegetter

# with WARC output support (see [storage] write_warc below)
pip install --upgrade filegetter[warc]

Install from Source

git clone https://github.com/ruarxive/filegetter.git
cd filegetter
pip install -e .

Quick Start

This example demonstrates archiving files from the Russian federal draft budget law 2023-2025.

1. Create a Project Directory

mkdir budget2023
cd budget2023

2. Create Configuration File

Create a file named filegetter.cfg:

[project]
name = budget2023
source = dataset.csv
source_type = csv
delimiter = ,

[data]
data_key = href

[files]
fetch_mode = prefix
root_url = https://sozd.duma.gov.ru
keys = href
storage_mode = filepath
transfer_ext = True

[storage]
storage_type = zip
compression = True

3. Run the Collection

filegetter run

Downloaded files are stored in storage/files.zip, with the per-file report in storage/processed.csv and the cached source id list in storage/allfiles.csv.


Usage

Command Syntax

filegetter [OPTIONS] COMMAND [ARGS]

Commands

  • run - Execute the file collection project

Options

Option Description
--projectpath PATH, -p PATH Project directory (default: current directory). All relative paths in the config are resolved against this directory, so you can run projects from anywhere.
--verbose, -v Debug-level logging (default: INFO)
--dry-run List pending downloads without fetching or writing anything
--limit N Download at most N pending files in this run
--refresh Re-read the source file instead of the cached allfiles.csv
--version Print the version and exit

Examples

# Run in current directory
filegetter run

# Run a project located elsewhere
filegetter run --projectpath /path/to/project

# See what would be downloaded, without touching the network
filegetter run --dry-run

# Download at most 100 pending files with debug logging
filegetter run --limit 100 --verbose

# Source file changed since the last run - rebuild the id cache
filegetter run --refresh

Exit Codes

  • 0 - all requested files were processed successfully (or skipped)
  • 1 - configuration error, or some files failed to download

Failures are always recorded in storage/processed.csv (with a non-200 status) and retried on the next run.

Proxies

Downloads go through a standard requests session, so the usual proxy environment variables (HTTP_PROXY, HTTPS_PROXY, NO_PROXY) are honoured automatically.


Python Library

Filegetter can also be used programmatically. FilegetterBuilder reads the project config and run() performs the download cycle, returning a stats dict:

from filegetter.cmds.project import ConfigError, FilegetterBuilder

try:
    builder = FilegetterBuilder("/path/to/project")
except ConfigError as e:
    print("Invalid configuration:", e)
    raise SystemExit(1)

stats = builder.run()          # same options as the CLI: dry_run=, limit=, refresh=
print(stats)                   # {'total': 42, 'skipped': 40, 'downloaded': 2, 'failed': 0}

if stats["failed"]:
    ...                        # failed files are retried on the next run()

Configuration Reference

All configuration is stored in filegetter.cfg using INI format. Invalid values produce a single error listing every problem found.

[project] Section

Option Required Description
name Yes Short name for the project
source Yes Source data file path, resolved relative to the project directory
source_type Yes One of: csv, jsonl, list
delimiter No Column delimiter for CSV (tab means tab; default ,)

[data] Section

Option Required Description
data_key Yes for csv/jsonl Column name (CSV) or dot-path into JSON records (JSONL) containing URLs or URL parts

[files] Section

Option Required Description
fetch_mode Yes prefix (prepend root_url) or pattern (root_url with a {} placeholder)
root_url Yes Base URL. May be empty when source ids are absolute URLs. Ids starting with http:///https:// always override it.
keys Yes Comma-separated list of keys containing file URLs/ids (used for JSONL sources)
storage_mode No filepath (mirror the URL path; default) or id (use the raw id as filename)
transfer_ext No If True, add an extension (from Content-Disposition, or default_ext as fallback) to files that have none
default_ext No Extension to add to extension-less files
file_storage_type No zip (default) or filesystem
delay No Delay in seconds between requests (default 0)
retries No Retry attempts for transient HTTP errors (429/5xx) and connection errors (default 2, exponential backoff)
timeout No Request timeout in seconds (default 30)
workers No Parallel download threads (default 1)
max_filesize No Skip files larger than this many bytes (default 270000000)
user_agent No User-Agent header (default: a Firefox browser string)
verify_ssl No Verify TLS certificates (default True)

[storage] Section

Option Required Description
storage_path No Directory for storage files, relative to the project (default storage)
compression No True (default) compresses the ZIP archive; False stores entries uncompressed
storage_type Legacy Alias for file_storage_type used by configs from 1.0.x
write_warc No If True, additionally write every successful response (full HTTP headers + body) to storage/files.warc.gz in WARC/1.0 format. Requires the warc extra: pip install filegetter[warc]

Report Format

storage/processed.csv has one row per attempted file:

Column Meaning
url Requested URL
filename Name under which the file was stored
mime Content-Type header
ext Extension detected from Content-Disposition
disp_name Filename from Content-Disposition
filesize Size in bytes
status HTTP status code (200, 404, ...) or error for connection failures / size limit rejections
sha256 SHA-256 checksum of the stored content

How It Works

  1. The source file (csv, jsonl or list) is read and reduced to a de-duplicated list of file identifiers, cached in allfiles.csv (pass --refresh to rebuild it after the source changes).
  2. For every id a URL is built (prefix or pattern mode); ids already recorded with status 200 in processed.csv are skipped.
  3. Each pending file is downloaded (streamed, with retries, timeout and the size limit enforced), stored via the configured backend and recorded in processed.csv with its checksum - the row is flushed immediately, so an interrupted run can simply be re-run.
  4. Files that failed (HTTP errors, connection problems) are recorded with their status and retried automatically on the next run; the command exits with code 1 whenever anything failed.
  5. With [storage] write_warc = True every successful response is also appended to storage/files.warc.gz as a WARC record with the complete HTTP headers, preserving full provenance for archival use.

Examples

See the examples/ directory for complete working examples:

  • budget2023 - Russian federal budget documents
  • goskatalog - Government catalog images
  • rupolitparties - Russian political parties data

Each example includes a complete filegetter.cfg and source data file.


Development

Setting Up Development Environment

git clone https://github.com/ruarxive/filegetter.git
cd filegetter
python -m venv .venv && source .venv/bin/activate
pip install -e .
pip install -r requirements-dev.txt
pre-commit install

Running Tests

# Run all tests (with coverage)
pytest

# Run a specific test file
pytest tests/test_storage.py

Code Quality

black filegetter/ tests/
isort filegetter/ tests/
flake8 filegetter/ tests/
mypy filegetter/

Contributing

Contributions are welcome! Please see CONTRIBUTING.md for guidelines.


License

This project is licensed under the MIT License - see the LICENSE file for details.

Copyright (c) 2022-2026 Russian national digital archive


Author

Ivan Begtin - ivan@begtin.tech


Metadata

Release files for filegetter 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for filegetter 1.1.0
File Size Uploaded
filegetter-1.1.0.tar.gz 28.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for filegetter 1.1.0
File Interpreter ABI Platform
filegetter-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 46.7 kB

Release files / filegetter-1.1.0.tar.gz

Download URL filegetter-1.1.0.tar.gz
Size 28.4 kB
Tags Source
SHA-256 checksum
How to use checksums
26f148c683e39aa828990c13fe0a3803d424d9300fa83011ca7e00a74c0732a2
BLAKE2b-256 checksum
How to use checksums
0e2ee78963468a5f49a4a88b95b3b3e4308a727dead9852365de7490e23ebff5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.7

Release files / filegetter-1.1.0-py3-none-any.whl

Download URL filegetter-1.1.0-py3-none-any.whl
Size 18.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8847141330a5ff2d73f795d5f8b5a8dc2f94d3c6ccec4d658e19a0a7801278c3
BLAKE2b-256 checksum
How to use checksums
1dacafc5062c9d099cea808813903295b38d6774bbdb64bbdeb15b03725977c8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.7

Release history Release notifications | RSS feed

This release

1.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page