Skip to main content

makeprov: Pythonic Provenance Tracking

makeprov is a small library for recording W3C PROV/JSON-LD provenance around Python functions that read and write files: which inputs produced which outputs, when, with what code and environment. A decorator wraps a function, tracks the files it declares as inputs/outputs, and writes a provenance record after each call. A minimal make-style dependency resolver and an optional Snakemake bridge are included, but the core contract of the library is the provenance record — not workflow orchestration, which tools like Snakemake already do well.

Features

  • Decorator-based rules that infer dependencies from InPath/OutPath parameters and write a PROV/JSON-LD record after every call.
  • A clean Plan → Run → Artifact model: the script at a commit is a prov:Plan, the runtime and the user are agents, and the two are tied together by prov:qualifiedAssociation/prov:hadPlan.
  • ArtifactRef lets a run cite external entities — a dataset IRI, an object-store key, a model checkpoint — without makeprov copying their metadata.
  • Provenance write failures are fatal by default (ProvenanceConfig(strict=True)), so a rule can't silently "succeed" with no record of what it did.
  • Resolve templated targets (results/{sample}.txt) via parse-style patterns, and a small dependency resolver (build/build_all) for chaining rules.
  • Serialize provenance as JSON-LD, or as RDF/TriG when rdflib is installed (pip install "makeprov[rdf]").
  • Optional Snakemake bridge that turns --d3dag and --detailed-summary output into PROV JSON-LD artifacts ready for inclusion in Snakemake HTML reports.

The provenance model

makeprov keeps PROV's distinction between the plan (the recipe) and the agent (whoever carried it out):

run.py @ git SHA          a prov:Plan, schema:SoftwareSourceCode
CPython 3.11              a prov:Agent, prov:SoftwareAgent
you (opt-in)              a prov:Agent, schema:Person

train-20260910T…-c2f6dc7b a prov:Activity
    prov:used                  dataset-X, the Python environment
    prov:wasAssociatedWith     runtime, person
    prov:qualifiedAssociation  [ prov:agent person ; prov:hadPlan run.py ]

results/model.txt         a prov:Entity
    prov:wasGeneratedBy        train-20260910T…-c2f6dc7b
    dct:identifier             sha256:…

Keeping the plan and the agent apart is what makes the graph mappable onto Workflow Run RO-Crate, whose instrument (the software that was run) and agent (a Person or Organization) are separate slots:

makeprov / PROV-O Process Run Crate
prov:Plan instrument
prov:Activity CreateAction
prov:used object
prov:wasGeneratedBy result
schema:Person agent agent
startedAtTime/endedAtTime startTime/endTime

The schema:Person agent is off by default: provenance documents are routinely committed and published, and a name and email address are personal data you should choose to publish rather than emit by accident. Turn it on with ProvenanceConfig(record_user=True), or --record-user on the Snakemake bridge. Without it, the qualified association names the runtime as the responsible agent.

Note that WRROC is Schema.org-native and defines no normative PROV-O mapping; the table above is a practical alignment, not an OWL equivalence. makeprov's own vocabulary stays prov:/schema: — RO-Crate and OpenLineage are intended as adapters over this model rather than changes to it.

Planned structure vs. observed execution

By default a document is purely retrospective: it records the activities that ran. A rule that was already up to date contributes nothing, because asserting an execution that did not happen would be worse than saying nothing.

Set emit_plan_graph = true (or pass --plan-graph) to additionally emit the prospective structure — each rule as a prov:Plan in its own right, linked to the plans it depends on by dct:requires:

run.py#rule-transform  a prov:Plan
    dct:requires   run.py#rule-extract
    dct:source     run.py

The activity's prov:hadPlan then points at the specific rule rather than the whole script. The Snakemake bridge does the same, collapsing the job DAG's edges to rule-level dct:requires edges.

Forge profiles

When no base_iri is set, makeprov derives one from the git remote. The supported hosts are declared in forges.toml — GitHub, GitLab, Bitbucket, Forgejo/Gitea/Codeberg and SourceHut — each giving the permalink layout for that host:

[[forge]]
name = "gitlab"
hosts = ["gitlab.com"]
blob = "{repo}/-/blob/{revision}/"

Point forge_profiles at your own TOML file to add self-hosted instances; entries there are matched first, so they can also override a built-in host. SSH and scp-style remotes (git@host:owner/repo.git) are understood, and any credentials embedded in a remote URL are stripped before it reaches a document.

Referencing things that aren't local files

ArtifactRef describes an entity a run consumed or produced. It is either local (makeprov stats and hashes it) or external (makeprov records the IRI and never touches the filesystem):

from makeprov import ArtifactRef, OutPath, rule

@rule()
def train(
    dataset: ArtifactRef = ArtifactRef.external(
        "https://example.org/datasets/train-v17",
        types=("prov:Entity", "schema:Dataset"),
        digest="sha256:...",
    ),
    model: OutPath = OutPath("models/m.pkl"),
):
    ...

The external object keeps its own detailed metadata; makeprov only records that this run used its stable IRI. External refs take no part in staleness checks, since they have no local mtime to compare.

Installation

You can install the module directly from PyPI:

pip install makeprov

Optional extras add RDF/TriG export, CLI subcommand support, or the Snakemake bridge:

pip install "makeprov[rdf]"        # rdflib + pyshacl for RDF/TriG export
pip install "makeprov[cli]"        # defopt, needed for makeprov.main()
pip install "makeprov[snakemake]"  # the makeprov-snakemake bridge

Usage

Here’s an example of how to use this package in your Python scripts:

from makeprov import rule, InPath, OutPath, build

@rule()
def process_data(
    sample: int | None = None,
    input_file: InPath = InPath('data/{sample:d}.txt'),
    output_file: OutPath = OutPath('results/{sample:d}.txt')
):
    with input_file.open('r') as infile, output_file.open('w') as outfile:
        data = infile.read()
        outfile.write(data.upper())

if __name__ == '__main__':
    # Build a specific templated target and its prerequisites
    from makeprov import build
    build('results/1.txt')

    # Or expose rules via a command line interface
    import defopt
    defopt.run(process_data)

You can execute examples/example.py via the CLI like so:

python examples/example.py build-all

# Or set configuration through the CLI
python examples/example.py build-all --conf='{"base_iri": "http://mybaseiri.org/", "prov_dir": "my_prov_directory"}' --force --input_file input.txt --output_file final_output.txt

# Or set configuration through a TOML file
python examples/example.py build-all -c @my_config.toml

# Inspect dependency resolution without executing rules
python examples/example.py --explain results/1.txt
python examples/example.py --to-dot results/1.txt

Complex CSV-to-RDF Workflow

For a more involved scenario, see examples/complex_example.py. It creates multiple CSV files, aggregates their contents, and emits an RDF graph that is both serialized to disk and embedded into the provenance dataset because the function returns an rdflib.Graph.

@rule()
def export_totals_graph(
    totals_csv: InPath = InPath("data/region_totals.csv"),
    graph_ttl: OutPath = OutPath("data/region_totals.ttl"),
) -> Graph:
    graph = Graph()
    graph.bind("sales", SALES)

    with totals_csv.open("r", newline="") as handle:
        for row in csv.DictReader(handle):
            region_key = row["region"].lower().replace(" ", "-")
            subject = SALES[f"region/{region_key}"]

            graph.add((subject, RDF.type, SALES.RegionTotal))
            graph.add((subject, SALES.regionName, Literal(row["region"])))
            graph.add((subject, SALES.totalUnits, Literal(row["total_units"], datatype=XSD.integer)))
            graph.add((subject, SALES.totalRevenue, Literal(row["total_revenue"], datatype=XSD.decimal)))

    with graph_ttl.open("w") as handle:
        handle.write(graph.serialize(format="turtle"))

    return graph

Run the entire workflow, including CSV generation and RDF export, with:

python examples/complex_example.py build-sales-report

Bundling nested provenance and directory outputs

Rules can merge the provenance from any rules they invoke by passing merge=True to makeprov.rule. Pair this with makeprov.OutDir to declare a directory and then materialize multiple outputs beneath it while keeping them linked to a single provenance record. Use makeprov.InDir for the same tracked-directory semantics on inputs. For nested structures, call subdir() on an OutDir/InDir to auto-wrap subfolders without manually constructing new instances. See examples/merge_outdir_example.py for an example.

Merging is enabled by default: top-level runs start a provenance buffer and flush it once the CLI finishes, so downstream rules end up in one document unless you explicitly turn buffering off with merge=False on a rule or in the global config. Nested merges append to their parent buffer rather than writing multiple files.

Configured context and isolated sessions

examples/context_demo_example.py demonstrates pinning a base IRI, writing provenance to a dedicated directory, and running rules inside an isolated session so registries and buffers do not leak across runs:

python examples/context_demo_example.py build-all

Snakemake workflows

Install the snakemake extra (pip install "makeprov[snakemake]") to get the makeprov-snakemake command, which shells out to Snakemake and converts the job DAG together with --detailed-summary metadata into a PROV document. It mirrors the familiar configuration flags from makeprov.config and writes JSON-LD by default. Note this is a best-effort bridge: it parses Snakemake's human-oriented text output, so treat it as a convenience for reports rather than an authoritative source of truth — it will raise rather than guess when it can't unambiguously parse a filename (e.g. one containing whitespace).

makeprov-snakemake --prov-path prov/snakemake -- --snakefile Snakefile --nolock

Wire the resulting file into a report by marking it with Snakemake’s report() helper:

rule provenance:
    input:
        "results/word_count.txt"
    output:
        "prov/snakemake.json"
    shell:
        (
            "makeprov-snakemake "
            "--prov-path prov/snakemake "
            "--out-fmt json --context --frame provenance "
            "-- "
            "--snakefile {workflow.snakefile} --nolock {input}"
        )

Using the optional --forceall-dag flag ensures that the job-level dependency edges in the provenance graph remain complete even when Snakemake skips nodes that are already up to date.

Configuration

You can customize the provenance tracking with the following options:

  • base_iri (str): Base IRI for new resources
  • prov_dir (str): Directory for writing PROV .json-ld or .trig files
  • force (bool): Force running of dependencies
  • dry_run (bool): Only check workflow, don't run anything
  • strict (bool, default True): Raise makeprov.ProvenanceWriteError if a rule's provenance record fails to write, instead of only logging a warning. A rule that produces a result but no provenance record is treated as a failure by default; set strict=False to opt out per-rule or globally.
  • run_id (str | None): Adopt an externally supplied run identity, such as a CI job id. When unset, each run gets a fresh unique id.
  • record_user (bool, default False): Record the invoking user, taken from git config user.name/user.email, as a schema:Person agent. Off by default so personal data isn't published by accident. CLI: --record-user.
  • emit_plan_graph (bool, default False): Also emit prospective structure — the rule dependency graph as prov:Plan nodes linked by dct:requires. Off by default, so a document describes only what actually ran. CLI: --plan-graph.
  • forge_profiles (str | None): TOML file of extra forge profiles, for self-hosted git hosts. CLI: --forge-profiles.

Upgrading to 0.7

0.7 changes the provenance model. The decorator API is unchanged — existing @rule functions using InPath/OutPath keep working — but the emitted graph differs:

  • The script is no longer a prov:SoftwareAgent. It is a prov:Plan, reached from the activity via prov:qualifiedAssociation/prov:hadPlan. Consumers that looked for prov:wasAssociatedWith to find the script should follow prov:hadPlan instead.
  • Outputs now carry a sha256 content digest. Previously the digest was computed and then discarded for outputs.
  • Entity IRIs derived from a GitHub remote are pinned to the commit rather than the branch, so an IRI no longer denotes different bytes after each push.
  • Run identifiers include seconds and a random suffix. Minute-resolution ids meant two runs of the same rule in one minute shared an activity IRI.
  • A declared output that is missing after a successful run now raises UnresolvedArtifactError instead of being dropped from the graph. Missing inputs are recorded without content metadata and logged, rather than disappearing.
  • Prov.create() takes list[ArtifactRef] instead of list[Path].
  • The Snakemake bridge follows the same model: each rule is now a prov:Plan at <base>rule/<name>, and each job activity carries a prov:qualifiedAssociation. Its agent node gained schema:SoftwareApplication.
  • The bridge no longer defaults to a urn:snakemake: namespace. That NID was never IANA-registered, so it named no real namespace and collided across unrelated workflows sharing rule names. It now shares the decorator API's identifier policy (makeprov.prov.resolve_iris): an explicit base_iri, else a commit-pinned base derived from a GitHub remote, else relative IRIs.
  • blob: identifiers are only minted for files inside the repository. An absolute path within the checkout is rewritten to its repo-relative form, and a path outside it gets a file: URI instead of a blob: IRI that would expand to a nonexistent location.
  • Entities carry schema:sha256 (bare hex) alongside the algorithm-qualified dct:identifier, so consumers no longer have to parse a prefix.
  • Activities state prov:generated as well as each entity's inverse prov:wasGeneratedBy, mirroring prov:used and mapping onto RO-Crate's result.
  • A dirty working tree is recorded as <sha>-dirty (the git describe --dirty convention) with a warning, instead of asserting a clean revision that does not describe what ran.
  • prov:wasDerivedFrom and rdfs:seeAlso on cached downloads are emitted as node references rather than string literals, so the links are traversable. CachedDownload(..., sha256=...) pins and verifies the cached copy.
  • The base heuristic covers all hosts in forges.toml, not just GitHub, and understands SSH remotes. Credentials embedded in a remote URL are stripped — previously a remote like https://user:token@github.com/o/r.git would have put the token into @base in every document.

Scoped spans and cached downloads

Use makeprov.span(label, prov_path=None, frame=None, context=None) as a context manager or decorator to bracket a chunk of work in its own provenance buffer. A span returns the merged Prov via span.prov, so nested spans can emit labeled artifacts without manual slicing/merging:

from makeprov import span

with span("model-run", prov_path="prov/models/model1"):
    run_model()

For remote resources that are cached locally, wrap the path with CachedDownload. It will lazily fetch on first access and record the source URL (and optional headers) in the provenance:

from makeprov import CachedDownload, rule

@rule()
def fetch_data(meta_json=CachedDownload("https://example.org/meta.json", "cache/meta.json")):
    with meta_json.open() as handle:
        return handle.read()

Documentation

Build the Sphinx docs (including autosummary API stubs) with the docs extra so that the CLI dependencies needed for imports are available:

pip install -e ".[docs]"
python docs/build.py

Contributing

Contributions are welcome! Please open an issue or submit a pull request.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Metadata

Release files for makeprov 0.7.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for makeprov 0.7.1
File Size Uploaded
makeprov-0.7.1.tar.gz 66.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for makeprov 0.7.1
File Interpreter ABI Platform
makeprov-0.7.1-py3-none-any.whl Python 3 none any Details

Total release size: 115.3 kB

Release files / makeprov-0.7.1.tar.gz

Download URL makeprov-0.7.1.tar.gz
Size 66.3 kB
Tags Source
SHA-256 checksum
How to use checksums
d928127492c059b9284757676f1a3456131121cebb90b74fc2a3270494a14e8b
BLAKE2b-256 checksum
How to use checksums
23172b9819987daaf1863e270078fb91077b4c9585e5623bde5316672e88017e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release files / makeprov-0.7.1-py3-none-any.whl

Download URL makeprov-0.7.1-py3-none-any.whl
Size 48.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
967897804baff8fe03387edc86e7f98bf0a919bcdbe8493671952c3b8255dfc9
BLAKE2b-256 checksum
How to use checksums
104d56a2d98d61b72f3122f41825fa84a0bfb99ecdc2fe33595d6712d849bc22
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release history Release notifications | RSS feed

0.8.1

2 release files

0.8.0

2 release files

This release

0.7.1 This release

2 release files

0.7.0

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page