Skip to main content

bnetza_bk6_scraper

License: MIT Python Versions (officially) supported PyPI Status Badge Unittests status badge Coverage status badge Linting status badge Formatting status badge

bnetza_bk6_scraper mirrors the documents published by the German Bundesnetzagentur (BNetzA) Beschlusskammer 6 (BK6) into a structured, git-diffable directory tree. BK6 regulates electricity network access and is a constant source of consultations, rulings (Festlegungen) and their attachments. Because the agency publishes these as loose PDFs on HTML pages with no changelog, tracking what changed and when is painful. This tool discovers every BK6 proceeding, downloads its PDFs and a normalized HTML snapshot of each phase page, and records structured metadata. Committing the output to git turns every regulatory update into a reviewable diff.

Installation

pip install bnetza_bk6_scraper

Usage

The package installs a single console command, bnetza-bk6-scraper, with a mirror subcommand:

bnetza-bk6-scraper mirror --target <dir> [--concurrency N] [--year YYYY] [-v]
Option Default Description
--target (required) Output directory (the mirror repository root).
--concurrency 4 Number of parallel HTTP fetches.
--year (all) Restrict the run to a single year, e.g. 2023.
-v, --verbose off Enable debug logging.

Example — mirror only the 2023 proceedings into ./mirror:

bnetza-bk6-scraper mirror --target ./mirror --year 2023 -v

Each run logs a summary such as run summary: 7 proceedings, 16 documents written, 0 failures.

Scoped mirroring (Python API)

The mirror CLI mirrors all proceedings. To mirror only a curated subset (e.g. the electricity GPKE/WiM/MaBiS Prozessdokumente), use the Python API: give the scraper a list of seed pages to crawl and a list of predicates (OR semantics — a document is downloaded if any predicate returns True). Predicates receive a CandidateDocument.

import asyncio
from bnetza_bk6_scraper import BnetzaBk6Scraper, CandidateDocument

GPKE = "https://www.bundesnetzagentur.de/DE/Beschlusskammern/BK06/BK6_83_Zug_Mess/831_gpke/gpke_node.html"

def is_prozessdokument_lesefassung(c: CandidateDocument) -> bool:
    name = c.filename.lower()
    return any(fw in name for fw in ("_gpke_", "_wim_", "_mabis_")) and "lesefassung" in name

asyncio.run(
    BnetzaBk6Scraper().mirror_seeds(
        target_dir=".", seeds=[GPKE], keep=[is_prozessdokument_lesefassung]
    )
)

Documents that carry an Aktenzeichen are written under <year>/<aktenzeichen>/; documents without one (e.g. the PID-Liste in the Datenformate tree) go under _other/…. A root manifest.json lists the kept documents.

Link detection is not limited to PDFs: .pdf, .xlsx and .xls are all mirrored. BNetzA publishes the Liste der Prüfidentifikatoren (PID) as both a PDF and a machine-friendly Excel file — the Excel is captured too, so a predicate like the one above picks up PID_..._info_....xlsx right next to its PDF.

Output layout

Proceedings are written under /{year}/{aktenzeichen}/, with a top-level index.json listing every mirrored proceeding:

<target>/
├── index.json                          # summary of all proceedings
└── 2023/
    └── BK6-23-241/
        ├── metadata.json               # structured proceeding metadata
        ├── BK6-23-241_beschluss.html   # normalized HTML snapshot of a phase page
        ├── BK6-23-241_beschluss_vom_07.05.26.pdf
        ├── BK6-23-241_bilarem.pdf
        └── BK6-23-241_anlage_bilarem.pdf
  • metadata.json captures the Aktenzeichen, year, title, status, Stand (last-modified date), any submission deadline (Frist), the phase pages, and one entry per document (title, type, source URL, filename).
  • The normalized *.html files are trimmed, stable snapshots of the source phase pages so that content changes surface as small diffs.
  • The PDFs are the proceeding's documents, downloaded verbatim.

Change detection is intentionally "dumb": the tool always writes the current state, and git diff in the mirror repository reveals what changed.

Mirror repository

The scraper is designed to feed a separate mirror repository, Hochfrequenz/bnetza_bk6_mirror. A scheduled GitHub Action there will periodically:

pip install bnetza_bk6_scraper
bnetza-bk6-scraper mirror --target .
git add -A && git commit -m "update BK6 mirror"

so that regulatory changes at BK6 become visible as reviewable git diffs and commit history. That Action is future work and does not live in this repository.

WAF / browser User-Agent

The BNetzA website sits behind a Web Application Firewall that rejects non-browser clients by serving a 200 OK "The requested URL was rejected" page instead of the real content. To get through, the scraper sends browser-like User-Agent and Accept headers and treats the rejection page as a retryable error. No credentials or API keys are required.

Contribute

This project uses tox for all quality gates. Create a one-shot development environment with everything installed:

tox -e dev

Individual gates: tox -e tests, tox -e linting, tox -e type_check, tox -e coverage, and tox -e spell_check. Run the full suite with tox.

Metadata

Release files for bnetza-bk6-scraper 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bnetza-bk6-scraper 0.1.1
File Size Uploaded
bnetza_bk6_scraper-0.1.1.tar.gz 52.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for bnetza-bk6-scraper 0.1.1
File Interpreter ABI Platform
bnetza_bk6_scraper-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 70.1 kB

Release files / bnetza_bk6_scraper-0.1.1.tar.gz

Download URL bnetza_bk6_scraper-0.1.1.tar.gz
Size 52.0 kB
Tags Source
SHA-256 checksum
How to use checksums
206e4e506a150b242afdb986bf78125e56560b20e2afdcb17606f94dfe721be0
BLAKE2b-256 checksum
How to use checksums
14150052ed8d631decd3aa57ed1e66d113901829a808a6f54e2650afe8e15ef2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 7, 2026.

Transparency log

Release files / bnetza_bk6_scraper-0.1.1-py3-none-any.whl

Download URL bnetza_bk6_scraper-0.1.1-py3-none-any.whl
Size 18.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
42d10b40f7019ed6080d72e231dea3e69653b007a7aa09586afa8a54548f6569
BLAKE2b-256 checksum
How to use checksums
26175d04cd69f8fbb2b26da32cb5767dc026a107a9665acbd531f46715b22c5a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 7, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page