Skip to main content

PurRDF for Python

PyPI License: MIT OR Apache-2.0 Python versions

PurRDF is a from-scratch, dependency-light RDF 1.2 engine — parsers and serializers, SPARQL, SHACL, ShEx, RDFC-1.0 canonicalization, and the GTS graph-transport container — written in Rust and carried verbatim into Python, JavaScript, and C. The purrdf package is the Python surface of that one engine: the same byte-identical semantics in every language, including triple terms, reifiers, and base-direction literals that most incumbent libraries do not carry.

Install

pip install purrdf

Requires Python 3.13+. Wheels bundle the native extension; no Rust toolchain needed.

Parse RDF

import purrdf

quads = purrdf.parse(
    '<https://example.org/alice> <http://xmlns.com/foaf/0.1/name> "Alice" .',
    purrdf.RdfFormat.TURTLE,
)

purrdf.parse accepts Turtle, TriG, N-Triples, N-Quads, TriX, and HexTuples (purrdf.RdfFormat); JSON-LD and RDF/XML travel through the dedicated purrdf.from_json_ld / purrdf.to_json_ld and purrdf.from_rdf_xml / purrdf.to_rdf_xml converters. All codecs are first-party with byte-deterministic output.

Configured JSON-LD and YAML-LD use one strict versioned options document. Compile a reusable context when serializing several datasets:

import json
import purrdf

options = json.dumps({
    "version": 1,
    "mode": "context",
    "prefixes": {"ex": "https://example.org/", "schema": "https://schema.org/"},
})
context = purrdf.CompiledJsonLdContext(options)
jsonld = purrdf.serialize_jsonld(
    nquads,
    format=purrdf.RdfFormat.N_QUADS,
    output_format="jsonld",
    context=context,
)

expanded, context, and deterministic dataset-IRI derived modes are explicit. PurRDF never infers a caller vocabulary or fetches a remote context.

Project graph, tabular, and research-object carriers

purrdf.project and purrdf.lift are thin calls into the same Rust projection engine used by every other surface. Configuration is mandatory, strict JSON: PurRDF supplies no vocabulary, identity IRI, or resource-limit default.

import json
import purrdf

config = json.dumps({
    "profile": "lpg-csv",
    "config": {
        "rdf_type": "https://example.org/type",
        "scope": {"mode": "all"},
        "limits": {
            "max_artifacts": 16,
            "max_artifact_bytes": 1_000_000,
            "max_total_bytes": 4_000_000,
            "max_archive_bytes": 5_000_000,
            "max_term_depth": 16,
        },
        "execution_limits": {
            "max_input_records": 1_000,
            "max_model_records": 1_000,
            "max_nodes": 1_000,
            "max_edges": 1_000,
        },
    },
})
package = purrdf.project(
    "@prefix ex: <https://example.org/> . ex:alice ex:knows ex:bob .",
    format=purrdf.RdfFormat.TURTLE,
    profile="lpg-csv",
    config=config,
)
lifted = purrdf.lift(package.archive, profile="lpg-csv", config=config)
assert lifted.dataset.quad_count() == 1
print([(loss.code, loss.location) for loss in package.losses])

Project profiles are lpg-csv, neo4j-csv, open-cypher, graphml, csvw-exact, csvw-terms, okf-terms, obo-graphs, skos, croissant-1.1, ro-crate-1.3, datacite-4.6, dcat-3, dcat-rdf, void, and frictionless-data-package-1. Curated CSVW/OKF terms, OBO Graphs, SKOS, native DCAT RDF, and VoID are deliberately write-only, ledgered views. Returned archives are canonical deterministic USTAR bytes and every result carries its always-computed structured loss records. Research-object contexts, vocabularies, identities, and profiles are all mandatory caller configuration.

Native RDF dataset descriptions use the same call. The complete JSON names the output syntax and either a mapped/CONSTRUCT DCAT source or the VoID source graphs, role vocabulary, dataset-prefix registries, and resource bounds:

from pathlib import Path
import purrdf

source = Path("void-source.trig").read_text()
void_config = Path("void.json").read_text()
description = purrdf.project(
    source,
    format=purrdf.RdfFormat.TRIG,
    profile="void",
    config=void_config,
)
Path("void.tar").write_bytes(description.archive)

Portable void-source.trig, void.json, and dcat-rdf.json examples are in crates/rdf/tests/fixtures/dataset-description/.

Attached RO-Crate packaging uses the same call with assets= set to a canonical payload-only USTAR archive and configuration packaging: "attached". The result contains the exact payloads, deterministic metadata, and self-contained preview; missing, unowned, reserved, or size-inconsistent members raise ValueError. See the runnable projection_roundtrip.py file-producing example.

For large LPG carriers, purrdf.project_artifacts(...) invokes a transactional artifact callback with package/artifact begin, bounded chunk, artifact finish, commit, and abort events. An optional progress callback receives immutable ProjectionProgress snapshots; callback exceptions abort the package and are returned unchanged. This path retains the selected canonical LPG model but not complete artifact bodies or USTAR bytes. See the runnable atomic-directory projection_stream.py example.

Validate with SHACL

The SHACL engine lives at purrdf.shapes (mirroring the Rust crate; purrdf.shacl is a back-compat alias):

from purrdf import shapes

report = shapes.validate(shapes_ttl=my_shapes, data_nt=my_data)
print(report["conforms"])

Complete SHACL Core, SHACL-SPARQL constraints/targets, and SHACL-AF sh:rule entailment via shapes.entail(...). Reusable parsed shapes are available as shapes.Shapes(shapes_ttl).validate_nt(data_nt).

Validate with ShEx

from purrdf import shex

results = shex.validate(
    my_schema_shexc,
    my_data_ttl,
    [("https://example.org/alice", "https://example.org/PersonShape")],
)
print(all(entry["conformant"] for entry in results))

The ShEx 2.1 validator passes 1,051/1,051 attempted validation tests of the official shexTest suite (see the repo's docs/CONFORMANCE.md).

rdflib compatibility layer

The package ships an rdflib-shaped API over the native engine:

from purrdf.compat.rdflib import Graph, URIRef

g = Graph()
g.parse(data=my_ntriples, format="nt")
print(len(g), g.serialize(format="turtle"))

For a literal, zero-change import rdflib, install the opt-in extra:

pip install purrdf[rdflib]

This pulls in the separate purrdf-rdflib distribution, whose top-level rdflib package re-exports the compat surface, so existing third-party code doing import rdflib / from rdflib.namespace import RDF transparently runs on purrdf. Caveat: that shadow claims the rdflib import name and must never be installed alongside the genuine rdflib — the two cannot co-inhabit one environment. It is a separate distribution (never bundled into the main purrdf wheel) precisely so environments that need the real rdflib simply omit it.

GTS graph transport and relational exports

GTS is PurRDF's single-file, content-addressed, append-only container for RDF 1.2 graphs. Build one from quads and export it straight to relational stores:

import purrdf

gts_bytes = purrdf.gts_from_quads(my_nquads_bytes, format=purrdf.RdfFormat.N_QUADS)

purrdf.gts_to_sqlite(gts_bytes, "graph.db")
purrdf.gts_to_duckdb(gts_bytes, "graph.duckdb")
files = purrdf.gts_to_parquet(gts_bytes, "out/")

The same entry points are grouped under purrdf.gts for discoverability.

Learn more

Licensed under MIT OR Apache-2.0, at your option.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

purrdf-0.8.5.tar.gz (3.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

purrdf-0.8.5-cp313-abi3-manylinux_2_34_x86_64.whl (7.3 MB view details)

Uploaded CPython 3.13+manylinux: glibc 2.34+ x86-64

File details

Details for the file purrdf-0.8.5.tar.gz.

File metadata

  • Download URL: purrdf-0.8.5.tar.gz
  • Upload date:
  • Size: 3.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for purrdf-0.8.5.tar.gz
Algorithm Hash digest
SHA256 ba68f79bb4b90717fcec9e621a0ee1e81711f507118e5f99e48a17a6c26db745
MD5 941881b03643552582c0b328026018ab
BLAKE2b-256 5cdbf922e7fdc8f680a02c6391c4037f8407ffed802992e86c50ece3490b7f36

See more details on using hashes here.

Provenance

The following attestation bundles were made for purrdf-0.8.5.tar.gz:

Publisher: release-pypi.yaml on Blackcat-Informatics/purrdf

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file purrdf-0.8.5-cp313-abi3-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for purrdf-0.8.5-cp313-abi3-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 24828f1f46ad3861ffa1895947dec90d40ba00735903337617bc8ae5bb56fa5d
MD5 07a4a6390450ffa8054385bc7e855130
BLAKE2b-256 bcf848819d3195967b856378ca0c4e0545c0a8e573fe294e1fb1d1e9d1723d66

See more details on using hashes here.

Provenance

The following attestation bundles were made for purrdf-0.8.5-cp313-abi3-manylinux_2_34_x86_64.whl:

Publisher: release-pypi.yaml on Blackcat-Informatics/purrdf

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page