Skip to main content

Harvesting for SCIGMA

License: AGPL-3.0

Installation

  1. Install the package, for example with uv:
uv pip install scigma_harvesting
  1. Copy .env.example to .env.
cp .env.example .env
  1. Fill in your local API keys and IDs.

  2. Optional you may adapt config.yaml to your needs. It contains all configuration values.

Usage

For working examples, see:

  • Minimal workflow – Basic harvest pipeline with Zotero/JSON/RIS export.
  • Advanced raw queries – Using raw query parameters (base_raw, ddb_raw, stcv_raw) for source-specific control.

Configuration

Most functions require a config parameter that is loaded from a YAML configuration file:

from scigma_harvesting import load_config, harvest

config = load_config("./config.yml")
records = harvest(config=config, must=["Aristoteles"])

See Configuration below for details on the config file format.

Top-Level Functions

The scigma_harvesting package exposes the following functions:

  • harvest(config, ...) – Main search function. Queries all sources (STCV, BASE, DDB), deduplicates results, returns list[HarvestRecord].
  • harvest_all(config, ...) – Like harvest, but without deduplication.
  • deduplicate(records, custom_filter=None) – Deduplicate HarvestRecords directly. Accepts optional custom_filter function for additional post-processing.
  • to_zotero(config, records, ...) – Exports HarvestRecords to Zotero (via API).
  • to_ris(records, path) – Exports HarvestRecords to a RIS file.
  • to_json(records, path) – Exports HarvestRecords to a JSON file.
  • activate_logging(debug=False) – Configure logging: INFO level by default, DEBUG level when debug=True.
  • load_config(path) – Load configuration from a YAML file.

Source-specific search functions (all require their respective config):

  • base_search(query, config, ...) – Search BASE API directly.
  • ddb_search(query, config, ...) – Search Deutsche Digitale Bibliothek directly.
  • stcv_search(config, ...) – Search STCV database directly.

Data Model

The normalized data model HarvestRecord of the output is importable from the package and documented in src/scigma_harvesting/pipeline/models.py.

Source-Specific Functions and Raw Dumps

Each search function (base_search, ddb_search, stcv_search) accepts an optional debug_dump_path parameter to save raw API/database responses for debugging and reproducibility. This is especially useful for investigating query issues or preserving data snapshots.

Additionally, stcv_search supports a sql_query parameter for direct SQL queries against the STCV SQLite database. When provided, it bypasses all other filter parameters and executes the raw SQL, which must return rows containing a cloi column.

BASE (scigma_harvesting.base)

from scigma_harvesting import load_config, base_search

config = load_config("./config.yml")

# Single page search
results = base_search(
    query="dccreator: lossau",
    config=config.base,
    debug_dump_path=Path("dumps/base_response.json")
)

Dump contents: JSON file containing source, query, params, and raw_xml_file (path to separate .xml file with pretty-printed raw response)

Special behavior: Creates TWO files:

  • base_response.json – Metadata and parsed results
  • base_response.xml – Raw XML response from BASE API (pretty-printed)

DDB (scigma_harvesting.ddb)

from scigma_harvesting import load_config, ddb_search

config = load_config("./config.yml")

results = ddb_search(
    query="lossau",
    config=config.ddb,
    debug_dump_path=Path("dumps/ddb_response.json")
)

Dump contents: JSON file containing source, query, params, and raw_response (complete DDB API JSON response)

STCV (scigma_harvesting.stcv)

from scigma_harvesting import load_config, stcv_search

config = load_config("./config.yml")

# Standard search with filters
results = stcv_search(
    config=config.stcv,
    must_terms=["Aristoteles"],
    debug_dump_path=Path("dumps/stcv_response.json")
)

# Direct SQL query (must return rows with cloi column)
results = stcv_search(
    config=config.stcv,
    sql_query="SELECT cloi FROM title WHERE title_ti LIKE '%Aristoteles%'"
)

Dump contents: JSON file containing source, search parameters (standalone, must_terms, must_not_terms, year_from, year_to, sql_query), and results (parsed records)

Note: STCV uses a local SQLite database, so dumps contain the already-parsed results rather than raw API responses.

Configuration

The package uses a centralized YAML-based configuration system. All API endpoints, paths, and other settings are managed through a configuration file.

Default config file: config.yml (can be specified when loading)

# BASE API
base:
  url: "https://api.base-search.net/cgi-bin/BaseHttpSearchInterface.fcgi"  # BASE API endpoint
  default_hits: 10  # Default results per page
  max_hits: 120  # Maximum results per request
  max_offset: 999  # Maximum offset for pagination

# DDB (Deutsche Digitale Bibliothek)
ddb:
  url: "https://api.deutsche-digitale-bibliothek.de/2/search/index/search/select"  # DDB API endpoint
  # ddb_time_parser_cache: "./data/internal/ddb/ddb_timeparser_cache.tsv"  # Legacy, now hardcoded in ddb_cache.py

# STCV (Short Title Catalogue Vlaanderen)
stcv:
  url: "https://anet.be/opendata/stcv/stcv.sqlite.gz"  # STCV API endpoint
  data_path: "./data/internal/stcv/latest/"  # Local database path
  timeout: 30  # Download/processing timeout in seconds

# Zotero API (environment variables are resolved from .env)
zotero:
  chunk_size: 50  # Items per batch request
  chunk_delay: 2.0  # Delay between batches in seconds
  http:
    read_timeout: 60.0  # HTTP read timeout in seconds
    connect_timeout: 30.0  # HTTP connect timeout in seconds

API keys for BASE and Zotero are stored in your .env file and loaded separately (not resolved automatically by load_config()). You can also build the config programmatically:

from pathlib import Path
from scigma_harvesting import Config, BaseConfig, DdbConfig, StcvConfig, ZoteroConfig, ZotHttpConfig

config = Config(
    base=BaseConfig(url="...", default_hits=10, max_hits=120, max_offset=999),
    ddb=DdbConfig(url="...", ddb_time_parser_cache=Path("./data/internal/ddb/ddb_timeparser_cache.tsv")),
    stcv=StcvConfig(url="...", data_path=Path("./data/internal/stcv/latest/"), timeout=30),
    zotero=ZoteroConfig(chunk_size=50, chunk_delay=2.0, http=ZotHttpConfig(read_timeout=60.0, connect_timeout=30.0))
)

Logging

The package uses Python's logging module. All modules (base, ddb, stcv, pipeline) log at DEBUG and INFO levels:

[!WARNING] When debug=True HTTP request URLs and headers are logged, which may include API keys. Never use debug=True in production or in environments where logs are shared or persisted.

  • DEBUG: Raw API requests/responses (query parameters, full XML/JSON responses), individual record details, Zotero batch creation summaries.
  • INFO: High-level progress (search start/end, result counts, harvest completion, Zotero chunk processing).

To enable detailed logging in your own code, use the provided helper:

from scigma_harvesting import activate_logging

# Activate logging (INFO level by default)
activate_logging()

# For DEBUG level (very verbose, including API requests/responses):
activate_logging(debug=True)

Troubleshooting

Common issues:

  • BASE API returns no results: Verify API_KEY_BASE in .env. Check query syntax (Lucene/SOLR). BASE has a hard limit of ~1080 results per query.
  • Zotero export fails: Verify API_KEY_ZOTERO and USER_ID_ZOTERO. Check collection key exists.
  • DDB IIIF/PDF links not accessible: Many DDB resources require institutional authentication.

Debug mode: Enable detailed logging or run development tests for debugging:

from scigma_harvesting import activate_logging
activate_logging(debug=True)  # Enable DEBUG logging for all HTTP requests/responses

Or run dev tests to inspect individual components:

uv run pytest -m dev -v

Debug logging in integration tests: To enable DEBUG logging for integration tests (e.g., to inspect HTTP requests), use the --log-debug flag:

# Without debug (default: no secrets in logs)
pytest tests/integration/

# With debug (WARNING: may log API keys and other secrets!)
pytest --log-debug tests/integration/

Note: The --log-debug flag is disabled by default. When enabled, API keys may appear in log files (tests/artifacts/logs/test_e2e_debug.log). Never use --log-debug in CI or shared environments.

Project Structure

src/scigma_harvesting/
├── base/       # BASE API integration
├── ddb/        # Deutsche Digitale Bibliothek
├── stcv/       # Short Title Catalogue Vlaanderen
└── pipeline/   # Core harvest pipeline
    ├── core/   # Main pipeline module (harvest, harvest_all)
    ├── models  # HarvestRecord and Zotero export models
    └── export  # Export functions (to_zotero, to_ris, to_json)

IO Quality

The tests/integration/test_io_quality.py module contains integration tests that verify data integrity during I/O operations (searching, processing, saving, loading). For manual comparision dumps can be found in tests/artifacts/io-quality after running this test.

DDB

  • API Access: The DDB search API is public and does not require an API key.
  • MODS (MARC21) URLs: Generated via /items/{id}/source/record — publicly accessible.
  • IIIF endpoints: Some IIIF manifests linked in DDB records may require institutional authentication tokens (not provided by this library).
  • METS endpoints: Typically require authentication and are not publicly accessible.

Updating the Cache

The DDB uses a custom day-based timestamp system for dates. The DDB_TIMEPARSER_CACHE dictionary in scigma_harvesting.ddb.ddb_cache contains precomputed day timestamps for January 1st of each year.

If you need to update the DDB time parser cache (e.g., to extend it to cover more years):

  1. Place your updated ddb_timeparser_cache.tsv file in data/internal/ddb/
  2. Run the conversion script:
uv run python scripts/convert_ddb_cache.py data/internal/ddb/ddb_timeparser_cache.tsv src/scigma_harvesting/ddb/ddb_cache.py

This will generate a new ddb_cache.py file with the updated cache as a Python dictionary.

The TSV file format is simple: each line contains <year>\t<ddb_day_timestamp> where <ddb_day_timestamp is the time_stamp for the first day of the year.

Development

Git and Versioning

GitFlow

Merge non-squashing from dev to main. Fast-forwarding is not allowed in order to see the different releases on main as merge commit.

Versioning

Versioning is handled primarily via the version in pyproject.toml. To bump run

uv version --bump [beta|patch|minor|major]

Then commit and push.

In order to make a release from that, a git tag will trigger the CI:

# git tag
git tag -a $(uv version | awk '{print $2}') -m "Release $(uv version | awk '{print $2}')"
# git push tag
git push origin $(uv version | awk '{print $2}')

Styling

The code is formated by black formatter.

Tests

Run all tests (including dev tests, not recommended for production!) with:

uv run pytest -v

Test markers:

  • -m prod – Production tests (safe for CI/CD, includes STCV database initialization).
  • -m dev – Development/debug tests only (not suitable for production).

If errors occur, you can run individual tests or test modules for debugging:

# Run a specific test file
uv run pytest tests/unit/pipeline/test_pipeline.py -v

# Run a single test function
uv run pytest tests/unit/pipeline/test_pipeline.py::test_base_raw_passed_through_directly -v

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scigma_harvesting-0.3.1.tar.gz (50.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scigma_harvesting-0.3.1-py3-none-any.whl (57.0 kB view details)

Uploaded Python 3

File details

Details for the file scigma_harvesting-0.3.1.tar.gz.

File metadata

  • Download URL: scigma_harvesting-0.3.1.tar.gz
  • Upload date:
  • Size: 50.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.15 {"installer":{"name":"uv","version":"0.11.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for scigma_harvesting-0.3.1.tar.gz
Algorithm Hash digest
SHA256 325fffc6bd5bf28ba8b38358b19c2ced2ac26ab016341d4dcfa061fc0d820d05
MD5 df9ecfdf1f775f7a6309131b0a6bb7ec
BLAKE2b-256 989955e796dde3a526f73818cd25a7745876588119234e0520cb6706632662fb

See more details on using hashes here.

File details

Details for the file scigma_harvesting-0.3.1-py3-none-any.whl.

File metadata

  • Download URL: scigma_harvesting-0.3.1-py3-none-any.whl
  • Upload date:
  • Size: 57.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.15 {"installer":{"name":"uv","version":"0.11.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for scigma_harvesting-0.3.1-py3-none-any.whl
Algorithm Hash digest
SHA256 aa88e4346b764c70b29d30dd696cbbdf2070d81deb26f36e208ab6f94f2406e2
MD5 ff8a973cc17bb0bf2599fe5f230789ce
BLAKE2b-256 376cb1e1248ec440c446b13d17ef7aab63e3cb64966b0ada96b754dae74b1a59

See more details on using hashes here.

Release history Release notifications | RSS feed

0.3.2

2 files

This release

0.3.1 This release

2 files

0.3.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page