Skip to main content

PurRDF for Python

PyPI License: MIT OR Apache-2.0 Python versions

PurRDF is a from-scratch, dependency-light RDF 1.2 engine — parsers and serializers, SPARQL, SHACL, ShEx, RDFC-1.0 canonicalization, and the GTS graph-transport container — written in Rust and carried verbatim into Python, JavaScript, and C. The purrdf package is the Python surface of that one engine: the same byte-identical semantics in every language, including triple terms, reifiers, and base-direction literals that most incumbent libraries do not carry.

Install

pip install purrdf

Requires Python 3.13+. Wheels bundle the native extension; no Rust toolchain needed.

Parse RDF

import purrdf

quads = purrdf.parse(
    '<https://example.org/alice> <http://xmlns.com/foaf/0.1/name> "Alice" .',
    purrdf.RdfFormat.TURTLE,
)

purrdf.parse accepts Turtle, TriG, N-Triples, N-Quads, TriX, and HexTuples (purrdf.RdfFormat); JSON-LD and RDF/XML travel through the dedicated purrdf.from_json_ld / purrdf.to_json_ld and purrdf.from_rdf_xml / purrdf.to_rdf_xml converters. All codecs are first-party with byte-deterministic output.

Configured JSON-LD and YAML-LD use one strict versioned options document. Compile a reusable context when serializing several datasets:

import json
import purrdf

options = json.dumps({
    "version": 1,
    "mode": "context",
    "prefixes": {"ex": "https://example.org/", "schema": "https://schema.org/"},
})
context = purrdf.CompiledJsonLdContext(options)
jsonld = purrdf.serialize_jsonld(
    nquads,
    format=purrdf.RdfFormat.N_QUADS,
    output_format="jsonld",
    context=context,
)

expanded, context, and deterministic dataset-IRI derived modes are explicit. PurRDF never infers a caller vocabulary or fetches a remote context.

Project graph, tabular, and research-object carriers

purrdf.project and purrdf.lift are thin calls into the same Rust projection engine used by every other surface. Configuration is mandatory, strict JSON: PurRDF supplies no vocabulary, identity IRI, or resource-limit default.

import json
import purrdf

config = json.dumps({
    "profile": "lpg-csv",
    "config": {
        "rdf_type": "https://example.org/type",
        "scope": {"mode": "all"},
        "limits": {
            "max_artifacts": 16,
            "max_artifact_bytes": 1_000_000,
            "max_total_bytes": 4_000_000,
            "max_archive_bytes": 5_000_000,
            "max_term_depth": 16,
        },
        "execution_limits": {
            "max_input_records": 1_000,
            "max_model_records": 1_000,
            "max_nodes": 1_000,
            "max_edges": 1_000,
        },
    },
})
package = purrdf.project(
    "@prefix ex: <https://example.org/> . ex:alice ex:knows ex:bob .",
    format=purrdf.RdfFormat.TURTLE,
    profile="lpg-csv",
    config=config,
)
lifted = purrdf.lift(package.archive, profile="lpg-csv", config=config)
assert lifted.dataset.quad_count() == 1
print([(loss.code, loss.location) for loss in package.losses])

Project profiles are lpg-csv, neo4j-csv, open-cypher, graphml, csvw-exact, csvw-terms, okf-terms, obo-graphs, skos, croissant-1.1, ro-crate-1.3, datacite-4.6, dcat-3, dcat-rdf, void, and frictionless-data-package-1. Curated CSVW/OKF terms, OBO Graphs, SKOS, native DCAT RDF, and VoID are deliberately write-only, ledgered views. Returned archives are canonical deterministic USTAR bytes and every result carries its always-computed structured loss records. Research-object contexts, vocabularies, identities, and profiles are all mandatory caller configuration.

Native RDF dataset descriptions use the same call. The complete JSON names the output syntax and either a mapped/CONSTRUCT DCAT source or the VoID source graphs, role vocabulary, dataset-prefix registries, and resource bounds:

from pathlib import Path
import purrdf

source = Path("void-source.trig").read_text()
void_config = Path("void.json").read_text()
description = purrdf.project(
    source,
    format=purrdf.RdfFormat.TRIG,
    profile="void",
    config=void_config,
)
Path("void.tar").write_bytes(description.archive)

Portable void-source.trig, void.json, and dcat-rdf.json examples are in crates/rdf/tests/fixtures/dataset-description/.

Attached RO-Crate packaging uses the same call with assets= set to a canonical payload-only USTAR archive and configuration packaging: "attached". The result contains the exact payloads, deterministic metadata, and self-contained preview; missing, unowned, reserved, or size-inconsistent members raise ValueError. See the runnable projection_roundtrip.py file-producing example.

For large LPG carriers, purrdf.project_artifacts(...) invokes a transactional artifact callback with package/artifact begin, bounded chunk, artifact finish, commit, and abort events. An optional progress callback receives immutable ProjectionProgress snapshots; callback exceptions abort the package and are returned unchanged. This path retains the selected canonical LPG model but not complete artifact bodies or USTAR bytes. See the runnable atomic-directory projection_stream.py example.

Validate with SHACL

The SHACL engine lives at purrdf.shapes (mirroring the Rust crate; purrdf.shacl is a back-compat alias):

from purrdf import shapes

report = shapes.validate(shapes_ttl=my_shapes, data_nt=my_data)
print(report["conforms"])

Complete SHACL Core, SHACL-SPARQL constraints/targets, and SHACL-AF sh:rule entailment via shapes.entail(...). Reusable parsed shapes are available as shapes.Shapes(shapes_ttl).validate_nt(data_nt).

Validate with ShEx

from purrdf import shex

results = shex.validate(
    my_schema_shexc,
    my_data_ttl,
    [("https://example.org/alice", "https://example.org/PersonShape")],
)
print(all(entry["conformant"] for entry in results))

The ShEx 2.1 validator passes 1,105/1,105 attempted validation tests of the official shexTest suite (see the repo's docs/CONFORMANCE.md).

Entailment regimes

The SPARQL entailment regimes live at purrdf.entail (mirroring the purrdf-entail Rust crate). It closes a dataset under a regime's own specification rule table and takes no shapes at all — not to be confused with purrdf.shapes.entail(...), which applies the SHACL-AF sh:rules a shapes graph declares.

import purrdf
from purrdf import entail

dataset = purrdf.RdfDataset(my_turtle, purrdf.RdfFormat.TURTLE)
closure, report = entail.materialize(dataset, "rdfs", "")
print(closure.to_nquads())
print(report)

For callers holding a document rather than a parsed dataset, entail.materialize_nt(text, regime, program) takes N-Triples/N-Quads and returns (canonical_nquads, report). Both accept the regime as a plain string ("simple", "rdf", "rdfs", "owl-rl", "owl-direct", "rif", "d") or as entail.Regime.RDFS.

All seven regimes close; none is refused. The third argument is the regime's own rule document. Six regimes take none, so theirs is "" — and a non-empty one raises rather than being silently discarded. "rif" is the exception: it entails under the caller's rules, which PurRDF does not declare, so its program is a normative RIF-in-XML document:

closure, report = entail.materialize(dataset, "rif", my_rif_xml)

"owl-direct" takes no program either, and that is a statement rather than an omission: its extra input is a query's class expressions, and this surface closes a dataset rather than answering a query — so what it runs is the query-independent tableau augmentation (the classification, the realization, the entailed role assertions and the owl:sameAs identifications the tableau decides about the ontology's own named terms).

The report is the second return value and is never optional. It is a byte-stable rendering naming which rules fired and how often, which specification rules did not fire, which constructs the run left at a boundary, what it consumed of the evaluator's fixed ceilings, and the contract hash of the calculus that ran — so a cached closure minted under a different rule set can be refused rather than trusted.

The rule tables are readable directly, so coverage is something you measure rather than something you take on faith:

defined = entail.rules("owl-rl")             # 78 — OWL 2 Profiles §4.3 Tables 4–9
fired = entail.implemented_rules("owl-rl")   # 78
missing = [rule for rule in entail.rules("rdfs") if rule not in entail.implemented_rules("rdfs")]
# [] — RDFS fires 18 of its 18 rules; the gap is legitimately empty
added = entail.extensions("owl-rl")          # ['ext-eq-diff-sym']

extensions(regime) is a third, disjoint inventory: the rules this build fires that no specification table states. owl-rl has one — ext-eq-diff-sym, symmetry of owl:differentFrom, sound and shaped exactly like prp-symp — and every other regime has none. It appears in neither rules() nor implemented_rules() for any regime, so the 78 above is unaffected by it: those two are statements about the specification, and firing a sound rule the table omits does not change what the table says. Asking is a question in its own right rather than something you learn by materializing a dataset and reading the report's extension line — though the report says the same thing, and the two cannot drift apart.

rdfD1, rdfD1a, rdfs14 and rdfs14a are in that fired set and each concludes about a fresh blank node. The restricted chase mints one as a frontier-addressed Skolem witness and closes under it, so the rules genuinely run — but every conclusion mentioning a witness is withheld when the closure is materialized back, because a SPARQL entailment regime draws its answers from the scoping graph and a minted blank node is not in it. The report says so with a boundary surrogate line rather than with a missing rule, and completeness reads exact-within-boundaries rather than exact.

78 / 78 is rule-table coverage, and rule-table coverage is not entailment conformance. The two are measured separately and stating only the first is the overclaim the reasoning report exists to prevent: on W3C's own OWL 2 RL entailment tests this chase reaches 11 of 27 published positive entailments and correctly withholds on 23 of 23 negative ones — the latter meaning no unsoundness was found. Both numbers are true. The full scoreboard, the typed divergence ledger, and every other suite are in docs/CONFORMANCE.md.

ValueError is raised for an unknown regime spelling (the message names the accepted set), for a program that is wrong for the regime — a non-empty one for any regime but "rif", or one "rif" cannot parse as a normative RIF-in-XML document — and for an exhausted evaluation ceiling. An exhausted ceiling is a refusal, never a truncated closure handed back as a complete one. Being "owl-direct" or "rif" is not itself a refusal: both materialize.

Description-Logic reasoning services

Materialization is the chase. The OWL 2 Direct-Semantics reasoner is a second lane on the same module — a SHOIQ(D) hypertableau — and every one of its services is on purrdf.entail. Each takes an N-Triples (or N-Quads) document and returns (answer, certificate) as a tuple, so a caller unpacks the evidence rather than being able to not ask for it:

Service Call Answer
Consistency entail.consistency(data) consistency true / false / unknownunknown means the tableau reached its step cap, and is never collapsed to false
Classification entail.classify(data) equivalent, subclass (transitively closed), direct (its reduction) and unsatisfiable lines
Realization entail.realize(data) type lines for the named individuals, then the most specific direct-type lines
Instance retrieval entail.instances(data, class_) instance <term> lines; class_ is ONE N-Triples term, angle brackets included
Axiom entailment entail.entails(data, axiom) entails true / false / unknown, then the axiom as it was read, so you can see which kind its predicate selected
Profile certification entail.profile(data) certified <profile> lines, most restrictive first (EL, QL, RL, DL, Full)
Module extraction entail.extract_module(data, signature, method) the locality module as canonical N-Quads; method is "bot", "top" or "star"
Justification entail.justify(data, axiom) a minimal subset of the ontology that still entails axiom, as canonical N-Quads
Proof entail.explain_conclusion(data, regime, conclusion) asserted, steps, and one rule line per rule the derivation cited
from purrdf import entail

ontology = (
    "<https://example.org/Cat>"
    " <http://www.w3.org/2000/01/rdf-schema#subClassOf>"
    " <https://example.org/Animal> .\n"
    "<https://example.org/felix>"
    " <http://www.w3.org/1999/02/22-rdf-syntax-ns#type>"
    " <https://example.org/Cat> .\n"
)

answer, certificate = entail.consistency(ontology)
assert answer.strip() == "consistency true"
assert certificate.startswith("purrdf-dl-certificate 1")
assert "completeness decided" in certificate

answer, _ = entail.instances(ontology, "<https://example.org/Animal>")
assert answer.strip() == "instance <https://example.org/felix>"

answer, _ = entail.profile(ontology)
assert answer.splitlines()[0] == "certified EL"

The certificate is the point, and there is a different one per kind of evidence. consistency, classify, realize, instances and entails render a purrdf-dl-certificate 1 block carrying the DL lane's own completeness — decided, decided-within-boundaries (an axiom that never became a DL clause, with each such construct named) or budget-exhausted. That is a different notion from the chase report's, which subtracts two rule tables, and it is rendered under a different banner so neither can be parsed as the other. profile reports no search at all — it is purely syntactic — so it renders a purrdf-owl-profile-certificate 1 block ending one-directional true: a certification proves membership, a violation does not prove non-membership. extract_module renders purrdf-module-extraction 1, whose conservative line says whether the module is minimal or a sound superset. justify renders purrdf-justification 1 and explain_conclusion renders purrdf-chase-proof 1; both re-check their own answers rather than restating them — sufficient and minimal are re-decided over the justification and over each one-axiom-smaller subset, and a proof's derived-* lines are what the checker re-derived from the proof term, not what the proof claims.

A tableau performs no derivation steps, so justify is a justification and deliberately not called a proof; explain_conclusion is the chase lane's genuinely derivational one. They are different kinds of thing rather than two spellings of one, which is why there is no single explain.

Nothing here re-implements the reasoner: every entry point routes through the same shared boundary the WebAssembly and C hosts call, checked against one committed golden-vector artifact, so the four hosts return byte-identical results for the same input.

rdflib compatibility layer

The package ships an rdflib-shaped API over the native engine:

from purrdf.compat.rdflib import Graph, URIRef

g = Graph()
g.parse(data=my_ntriples, format="nt")
print(len(g), g.serialize(format="turtle"))

For a literal, zero-change import rdflib, install the opt-in extra:

pip install purrdf[rdflib]

This pulls in the separate purrdf-rdflib distribution, whose top-level rdflib package re-exports the compat surface, so existing third-party code doing import rdflib / from rdflib.namespace import RDF transparently runs on purrdf. Caveat: that shadow claims the rdflib import name and must never be installed alongside the genuine rdflib — the two cannot co-inhabit one environment. It is a separate distribution (never bundled into the main purrdf wheel) precisely so environments that need the real rdflib simply omit it.

GTS graph transport and relational exports

GTS is PurRDF's single-file, content-addressed, append-only container for RDF 1.2 graphs. Build one from quads and export it straight to relational stores:

import purrdf

gts_bytes = purrdf.gts_from_quads(my_nquads_bytes, format=purrdf.RdfFormat.N_QUADS)

purrdf.gts_to_sqlite(gts_bytes, "graph.db")
purrdf.gts_to_duckdb(gts_bytes, "graph.duckdb")
files = purrdf.gts_to_parquet(gts_bytes, "out/")

The same entry points are grouped under purrdf.gts for discoverability.

Learn more

Licensed under MIT OR Apache-2.0, at your option.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

purrdf-0.10.0.tar.gz (5.8 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

purrdf-0.10.0-cp313-abi3-manylinux_2_34_x86_64.whl (8.1 MB view details)

Uploaded CPython 3.13+manylinux: glibc 2.34+ x86-64

File details

Details for the file purrdf-0.10.0.tar.gz.

File metadata

  • Download URL: purrdf-0.10.0.tar.gz
  • Upload date:
  • Size: 5.8 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for purrdf-0.10.0.tar.gz
Algorithm Hash digest
SHA256 1153c63f1740af180b070c8e8dd97f9898c35e8bf311e5f895b38068c4a3ee5e
MD5 e8b2fd4bf91d235627a0f21786a2c554
BLAKE2b-256 9d6df137f6ec46e9d5a63f9543fd792259a6c9c873c37913d66a3fb30b2d6378

See more details on using hashes here.

Provenance

The following attestation bundles were made for purrdf-0.10.0.tar.gz:

Publisher: release-pypi.yaml on Blackcat-Informatics/purrdf

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file purrdf-0.10.0-cp313-abi3-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for purrdf-0.10.0-cp313-abi3-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 7df789d30bc26c5f45efc6cdf4571cddec9132cb8278623056f747d587392ae7
MD5 ea0317073ce0f2c1ee4592e0cf96b48b
BLAKE2b-256 275ed01db203568a861145fe16343e0db98c340f83cf4dbce853b6fc5db70ae8

See more details on using hashes here.

Provenance

The following attestation bundles were made for purrdf-0.10.0-cp313-abi3-manylinux_2_34_x86_64.whl:

Publisher: release-pypi.yaml on Blackcat-Informatics/purrdf

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page