Skip to main content

PurRDF for Python

PyPI License: MIT OR Apache-2.0 Python versions

PurRDF is a from-scratch, dependency-light RDF 1.2 engine — parsers and serializers, SPARQL, SHACL, ShEx, RDFC-1.0 canonicalization, and the GTS graph-transport container — written in Rust and carried verbatim into Python, JavaScript, and C. The purrdf package is the Python surface of that one engine: the same byte-identical semantics in every language, including triple terms, reifiers, and base-direction literals that most incumbent libraries do not carry.

Install

pip install purrdf

Requires Python 3.13+. Wheels bundle the native extension; no Rust toolchain needed.

Parse RDF

import purrdf

quads = purrdf.parse(
    '<https://example.org/alice> <http://xmlns.com/foaf/0.1/name> "Alice" .',
    purrdf.RdfFormat.TURTLE,
)

purrdf.parse accepts Turtle, TriG, N-Triples, N-Quads, TriX, and HexTuples (purrdf.RdfFormat); JSON-LD and RDF/XML travel through the dedicated purrdf.from_json_ld / purrdf.to_json_ld and purrdf.from_rdf_xml / purrdf.to_rdf_xml converters. All codecs are first-party with byte-deterministic output.

Configured JSON-LD and YAML-LD use one strict versioned options document. Compile a reusable context when serializing several datasets:

import json
import purrdf

options = json.dumps({
    "version": 1,
    "mode": "context",
    "prefixes": {"ex": "https://example.org/", "schema": "https://schema.org/"},
})
context = purrdf.CompiledJsonLdContext(options)
jsonld = purrdf.serialize_jsonld(
    nquads,
    format=purrdf.RdfFormat.N_QUADS,
    output_format="jsonld",
    context=context,
)

expanded, context, and deterministic dataset-IRI derived modes are explicit. PurRDF never infers a caller vocabulary or fetches a remote context.

Project graph, tabular, and research-object carriers

purrdf.project and purrdf.lift are thin calls into the same Rust projection engine used by every other surface. Configuration is mandatory, strict JSON: PurRDF supplies no vocabulary, identity IRI, or resource-limit default.

import json
import purrdf

config = json.dumps({
    "profile": "lpg-csv",
    "config": {
        "rdf_type": "https://example.org/type",
        "scope": {"mode": "all"},
        "limits": {
            "max_artifacts": 16,
            "max_artifact_bytes": 1_000_000,
            "max_total_bytes": 4_000_000,
            "max_archive_bytes": 5_000_000,
            "max_term_depth": 16,
        },
        "execution_limits": {
            "max_input_records": 1_000,
            "max_model_records": 1_000,
            "max_nodes": 1_000,
            "max_edges": 1_000,
        },
    },
})
package = purrdf.project(
    "@prefix ex: <https://example.org/> . ex:alice ex:knows ex:bob .",
    format=purrdf.RdfFormat.TURTLE,
    profile="lpg-csv",
    config=config,
)
lifted = purrdf.lift(package.archive, profile="lpg-csv", config=config)
assert lifted.dataset.quad_count() == 1
print([(loss.code, loss.location) for loss in package.losses])

Project profiles are lpg-csv, neo4j-csv, open-cypher, graphml, csvw-exact, csvw-terms, okf-terms, obo-graphs, skos, croissant-1.1, ro-crate-1.3, datacite-4.6, dcat-3, dcat-rdf, void, and frictionless-data-package-1. Curated CSVW/OKF terms, OBO Graphs, SKOS, native DCAT RDF, and VoID are deliberately write-only, ledgered views. Returned archives are canonical deterministic USTAR bytes and every result carries its always-computed structured loss records. Research-object contexts, vocabularies, identities, and profiles are all mandatory caller configuration.

Native RDF dataset descriptions use the same call. The complete JSON names the output syntax and either a mapped/CONSTRUCT DCAT source or the VoID source graphs, role vocabulary, dataset-prefix registries, and resource bounds:

from pathlib import Path
import purrdf

source = Path("void-source.trig").read_text()
void_config = Path("void.json").read_text()
description = purrdf.project(
    source,
    format=purrdf.RdfFormat.TRIG,
    profile="void",
    config=void_config,
)
Path("void.tar").write_bytes(description.archive)

Portable void-source.trig, void.json, and dcat-rdf.json examples are in crates/rdf/tests/fixtures/dataset-description/.

Attached RO-Crate packaging uses the same call with assets= set to a canonical payload-only USTAR archive and configuration packaging: "attached". The result contains the exact payloads, deterministic metadata, and self-contained preview; missing, unowned, reserved, or size-inconsistent members raise ValueError. See the runnable projection_roundtrip.py file-producing example.

For large LPG carriers, purrdf.project_artifacts(...) invokes a transactional artifact callback with package/artifact begin, bounded chunk, artifact finish, commit, and abort events. An optional progress callback receives immutable ProjectionProgress snapshots; callback exceptions abort the package and are returned unchanged. This path retains the selected canonical LPG model but not complete artifact bodies or USTAR bytes. See the runnable atomic-directory projection_stream.py example.

Validate with SHACL

The SHACL engine lives at purrdf.shapes (mirroring the Rust crate; purrdf.shacl is a back-compat alias):

from purrdf import shapes

report = shapes.validate(shapes_ttl=my_shapes, data_nt=my_data)
print(report["conforms"])

Complete SHACL Core, SHACL-SPARQL constraints/targets, and SHACL-AF sh:rule entailment via shapes.entail(...). Reusable parsed shapes are available as shapes.Shapes(shapes_ttl).validate_nt(data_nt).

Validate with ShEx

from purrdf import shex

results = shex.validate(
    my_schema_shexc,
    my_data_ttl,
    [("https://example.org/alice", "https://example.org/PersonShape")],
)
print(all(entry["conformant"] for entry in results))

The ShEx 2.1 validator passes 1,051/1,051 attempted validation tests of the official shexTest suite (see the repo's docs/CONFORMANCE.md).

rdflib compatibility layer

The package ships an rdflib-shaped API over the native engine:

from purrdf.compat.rdflib import Graph, URIRef

g = Graph()
g.parse(data=my_ntriples, format="nt")
print(len(g), g.serialize(format="turtle"))

For a literal, zero-change import rdflib, install the opt-in extra:

pip install purrdf[rdflib]

This pulls in the separate purrdf-rdflib distribution, whose top-level rdflib package re-exports the compat surface, so existing third-party code doing import rdflib / from rdflib.namespace import RDF transparently runs on purrdf. Caveat: that shadow claims the rdflib import name and must never be installed alongside the genuine rdflib — the two cannot co-inhabit one environment. It is a separate distribution (never bundled into the main purrdf wheel) precisely so environments that need the real rdflib simply omit it.

GTS graph transport and relational exports

GTS is PurRDF's single-file, content-addressed, append-only container for RDF 1.2 graphs. Build one from quads and export it straight to relational stores:

import purrdf

gts_bytes = purrdf.gts_from_quads(my_nquads_bytes, format=purrdf.RdfFormat.N_QUADS)

purrdf.gts_to_sqlite(gts_bytes, "graph.db")
purrdf.gts_to_duckdb(gts_bytes, "graph.duckdb")
files = purrdf.gts_to_parquet(gts_bytes, "out/")

The same entry points are grouped under purrdf.gts for discoverability.

Learn more

Licensed under MIT OR Apache-2.0, at your option.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

purrdf-0.8.2.tar.gz (3.1 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

purrdf-0.8.2-cp313-abi3-manylinux_2_34_x86_64.whl (7.3 MB view details)

Uploaded CPython 3.13+manylinux: glibc 2.34+ x86-64

File details

Details for the file purrdf-0.8.2.tar.gz.

File metadata

  • Download URL: purrdf-0.8.2.tar.gz
  • Upload date:
  • Size: 3.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for purrdf-0.8.2.tar.gz
Algorithm Hash digest
SHA256 6a277ade3ca58d86e16ae592f2ec33c3eb37953f2c08722bb4cba05ce17eb59a
MD5 3542bab9185aec27c4e1e4ab86bfc086
BLAKE2b-256 0b630f94de646a2844e30b982800f655f73f58113db8a6b3f81d191aaebf8db0

See more details on using hashes here.

Provenance

The following attestation bundles were made for purrdf-0.8.2.tar.gz:

Publisher: release-pypi.yaml on Blackcat-Informatics/purrdf

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file purrdf-0.8.2-cp313-abi3-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for purrdf-0.8.2-cp313-abi3-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 7788f9acfae4ce926b1292bde338f6863a13c8636f56842eb280f12e8a71a054
MD5 fae520912d6a0aadc779b798d1513d4d
BLAKE2b-256 639ab4292795d1ead32442036468894e16b2ce923ce9b9528df308085ba96575

See more details on using hashes here.

Provenance

The following attestation bundles were made for purrdf-0.8.2-cp313-abi3-manylinux_2_34_x86_64.whl:

Publisher: release-pypi.yaml on Blackcat-Informatics/purrdf

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page