filegetter
filegetter is a command-line tool and Python library for automated bulk file collection from public data sources.
Table of Contents
- Overview
- Features
- Installation
- Quick Start
- Usage
- Python Library
- Configuration Reference
- How It Works
- Examples
- Development
- License
Overview
Filegetter automates downloading large numbers of files from URLs listed in configuration-driven project files. It was built to streamline file collection from datasets produced by other tools (API scrapers, data extracts), with support for several input formats and storage backends, resumable runs and a detailed per-file report.
Features
✅ Multiple Input Formats - CSV, JSON Lines (JSONL) and plain text lists
✅ Flexible URL Handling - URL prefixes, {}-patterns and absolute URLs
✅ Storage Options - ZIP archive or filesystem directory tree
✅ Resume Capability - Already-downloaded files are skipped; failed files are retried
✅ Integrity & Reporting - CSV report with HTTP status, MIME type, size and SHA-256 checksum
✅ WARC Output - Optional archival WARC/1.0 records (pip install filegetter[warc])
✅ Robust Downloads - Retries with backoff, timeouts, per-file error isolation
✅ Politeness Controls - Configurable delay, User-Agent, worker count and file size limit
✅ Zip Compression Toggle - compression option for smaller archives or faster writes
Installation
Requirements
- Python 3.9 or higher
Install from PyPI
pip install --upgrade pip
pip install --upgrade filegetter
# with WARC output support (see [storage] write_warc below)
pip install --upgrade filegetter[warc]
Install from Source
git clone https://github.com/ruarxive/filegetter.git
cd filegetter
pip install -e .
Quick Start
This example demonstrates archiving files from the Russian federal draft budget law 2023-2025.
1. Create a Project Directory
mkdir budget2023
cd budget2023
2. Create Configuration File
Create a file named filegetter.cfg:
[project]
name = budget2023
source = dataset.csv
source_type = csv
delimiter = ,
[data]
data_key = href
[files]
fetch_mode = prefix
root_url = https://sozd.duma.gov.ru
keys = href
storage_mode = filepath
transfer_ext = True
[storage]
storage_type = zip
compression = True
3. Run the Collection
filegetter run
Downloaded files are stored in storage/files.zip, with the per-file report
in storage/processed.csv and the cached source id list in
storage/allfiles.csv.
Usage
Command Syntax
filegetter [OPTIONS] COMMAND [ARGS]
Commands
run- Execute the file collection project
Options
| Option | Description |
|---|---|
--projectpath PATH, -p PATH |
Project directory (default: current directory). All relative paths in the config are resolved against this directory, so you can run projects from anywhere. |
--verbose, -v |
Debug-level logging (default: INFO) |
--dry-run |
List pending downloads without fetching or writing anything |
--limit N |
Download at most N pending files in this run |
--refresh |
Re-read the source file instead of the cached allfiles.csv |
--version |
Print the version and exit |
Examples
# Run in current directory
filegetter run
# Run a project located elsewhere
filegetter run --projectpath /path/to/project
# See what would be downloaded, without touching the network
filegetter run --dry-run
# Download at most 100 pending files with debug logging
filegetter run --limit 100 --verbose
# Source file changed since the last run - rebuild the id cache
filegetter run --refresh
Exit Codes
0- all requested files were processed successfully (or skipped)1- configuration error, or some files failed to download
Failures are always recorded in storage/processed.csv (with a non-200
status) and retried on the next run.
Proxies
Downloads go through a standard requests session, so the usual proxy
environment variables (HTTP_PROXY, HTTPS_PROXY, NO_PROXY) are honoured
automatically.
Python Library
Filegetter can also be used programmatically. FilegetterBuilder reads the
project config and run() performs the download cycle, returning a stats
dict:
from filegetter.cmds.project import ConfigError, FilegetterBuilder
try:
builder = FilegetterBuilder("/path/to/project")
except ConfigError as e:
print("Invalid configuration:", e)
raise SystemExit(1)
stats = builder.run() # same options as the CLI: dry_run=, limit=, refresh=
print(stats) # {'total': 42, 'skipped': 40, 'downloaded': 2, 'failed': 0}
if stats["failed"]:
... # failed files are retried on the next run()
Configuration Reference
All configuration is stored in filegetter.cfg using INI format. Invalid
values produce a single error listing every problem found.
[project] Section
| Option | Required | Description |
|---|---|---|
name |
Yes | Short name for the project |
source |
Yes | Source data file path, resolved relative to the project directory |
source_type |
Yes | One of: csv, jsonl, list |
delimiter |
No | Column delimiter for CSV (tab means tab; default ,) |
[data] Section
| Option | Required | Description |
|---|---|---|
data_key |
Yes for csv/jsonl |
Column name (CSV) or dot-path into JSON records (JSONL) containing URLs or URL parts |
[files] Section
| Option | Required | Description |
|---|---|---|
fetch_mode |
Yes | prefix (prepend root_url) or pattern (root_url with a {} placeholder) |
root_url |
Yes | Base URL. May be empty when source ids are absolute URLs. Ids starting with http:///https:// always override it. |
keys |
Yes | Comma-separated list of keys containing file URLs/ids (used for JSONL sources) |
storage_mode |
No | filepath (mirror the URL path; default) or id (use the raw id as filename) |
transfer_ext |
No | If True, add an extension (from Content-Disposition, or default_ext as fallback) to files that have none |
default_ext |
No | Extension to add to extension-less files |
file_storage_type |
No | zip (default) or filesystem |
delay |
No | Delay in seconds between requests (default 0) |
retries |
No | Retry attempts for transient HTTP errors (429/5xx) and connection errors (default 2, exponential backoff) |
timeout |
No | Request timeout in seconds (default 30) |
workers |
No | Parallel download threads (default 1) |
max_filesize |
No | Skip files larger than this many bytes (default 270000000) |
user_agent |
No | User-Agent header (default: a Firefox browser string) |
verify_ssl |
No | Verify TLS certificates (default True) |
[storage] Section
| Option | Required | Description |
|---|---|---|
storage_path |
No | Directory for storage files, relative to the project (default storage) |
compression |
No | True (default) compresses the ZIP archive; False stores entries uncompressed |
storage_type |
Legacy | Alias for file_storage_type used by configs from 1.0.x |
write_warc |
No | If True, additionally write every successful response (full HTTP headers + body) to storage/files.warc.gz in WARC/1.0 format. Requires the warc extra: pip install filegetter[warc] |
Report Format
storage/processed.csv has one row per attempted file:
| Column | Meaning |
|---|---|
url |
Requested URL |
filename |
Name under which the file was stored |
mime |
Content-Type header |
ext |
Extension detected from Content-Disposition |
disp_name |
Filename from Content-Disposition |
filesize |
Size in bytes |
status |
HTTP status code (200, 404, ...) or error for connection failures / size limit rejections |
sha256 |
SHA-256 checksum of the stored content |
How It Works
- The source file (
csv,jsonlorlist) is read and reduced to a de-duplicated list of file identifiers, cached inallfiles.csv(pass--refreshto rebuild it after the source changes). - For every id a URL is built (
prefixorpatternmode); ids already recorded with status200inprocessed.csvare skipped. - Each pending file is downloaded (streamed, with retries, timeout and the
size limit enforced), stored via the configured backend and recorded in
processed.csvwith its checksum - the row is flushed immediately, so an interrupted run can simply be re-run. - Files that failed (HTTP errors, connection problems) are recorded with their status and retried automatically on the next run; the command exits with code 1 whenever anything failed.
- With
[storage] write_warc = Trueevery successful response is also appended tostorage/files.warc.gzas a WARC record with the complete HTTP headers, preserving full provenance for archival use.
Examples
See the examples/ directory for complete working examples:
- budget2023 - Russian federal budget documents
- goskatalog - Government catalog images
- rupolitparties - Russian political parties data
Each example includes a complete filegetter.cfg and source data file.
Development
Setting Up Development Environment
git clone https://github.com/ruarxive/filegetter.git
cd filegetter
python -m venv .venv && source .venv/bin/activate
pip install -e .
pip install -r requirements-dev.txt
pre-commit install
Running Tests
# Run all tests (with coverage)
pytest
# Run a specific test file
pytest tests/test_storage.py
Code Quality
black filegetter/ tests/
isort filegetter/ tests/
flake8 filegetter/ tests/
mypy filegetter/
Contributing
Contributions are welcome! Please see CONTRIBUTING.md for guidelines.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Copyright (c) 2022-2026 Russian national digital archive
Author
Ivan Begtin - ivan@begtin.tech
Links
- GitHub Repository: https://github.com/ruarxive/filegetter
- PyPI Package: https://pypi.org/project/filegetter/
- Issue Tracker: https://github.com/ruarxive/filegetter/issues
- Changelog: CHANGELOG.md
Metadata
Release files for filegetter 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| filegetter-1.1.0.tar.gz | 28.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| filegetter-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 46.7 kB
Release files / filegetter-1.1.0.tar.gz
| Download URL | filegetter-1.1.0.tar.gz |
|---|---|
| Size | 28.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
26f148c683e39aa828990c13fe0a3803d424d9300fa83011ca7e00a74c0732a2
|
|
BLAKE2b-256 checksum How to use checksums |
0e2ee78963468a5f49a4a88b95b3b3e4308a727dead9852365de7490e23ebff5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.7
|
Release files / filegetter-1.1.0-py3-none-any.whl
| Download URL | filegetter-1.1.0-py3-none-any.whl |
|---|---|
| Size | 18.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8847141330a5ff2d73f795d5f8b5a8dc2f94d3c6ccec4d658e19a0a7801278c3
|
|
BLAKE2b-256 checksum How to use checksums |
1dacafc5062c9d099cea808813903295b38d6774bbdb64bbdeb15b03725977c8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.7
|