SchemaIR for Python
The maintained 0.4 runtime declines recursive inputs in type interpretation,
schema/term relations, and variable solving. Reachable selfRef or cyclic
preloaded refs produce incomplete (interpreter) or unknown (analysis), with
an unsupported.recursive diagnostic. Whole-root recursive payload validation
and reference loading remain available. Payload validation is bounded (depth
64, 10,000 work steps); exhaustion returns unknown. refine and general
intersect execution are disabled even for non-recursive inputs. Container
constraints require a provider; the variable solver is experimental. See the
legacy support boundary.
Types as data.
SchemaIR is a language-independent protocol and intermediate representation for portable data contracts, operation definitions, and type calculations.
The schemair package is the Python
implementation. It represents contracts, operations, and type expressions as
serializable data.
For the protocol's theory, IR model, design philosophy, standard vocabulary, cross-language boundaries, and conformance rules, see the SchemaIR project README and the protocol specification. This page focuses on using SchemaIR from Python.
About this implementation
SchemaIR declarations are ordinary JSON-compatible dictionaries. You can store a contract with a project, exchange it with another tool, and inspect or calculate types without executing the application behind the contract.
This binding implements the shared protocol and conformance rules rather than Python annotations or Pydantic models. See the implementation matrix for coverage and status.
Install
Use Python 3.11 or newer. Install the published package from PyPI:
python -m pip install schemair
To work on the implementation from this repository, install the local package instead:
python -m pip install -e ./python
The package is published on PyPI as schemair.
Quick start
import json
from schemair import authoring as s
from schemair import (
check_schema_relation,
evaluate_expression,
validate_payload,
validate_schema,
)
# A SchemaNode describes the concrete output shape as data.
user = s.record([
s.required_field("id", s.string()),
s.required_field("name", s.string()),
])
# Reload the declaration without importing its original application.
loaded = json.loads(json.dumps(user))
assert validate_schema(loaded).status == "accepted"
# Build a SchemaTypeExpr, then evaluate it before executing an operation.
selection = s.field_of(loaded, "name")
name_type = evaluate_expression(selection)
assert name_type.status == "evaluated"
assert check_schema_relation(name_type.value, s.string()).status == "assignable"
# Validate actual values independently of type calculation.
payload = validate_payload({"id": "123", "name": "Ada"}, loaded)
assert payload.status == "accepted"
Python exposes snake_case helpers directly from schemair.authoring and also
provides schemair.authoring.schema and schemair.authoring.expression
namespaces that mirror the TypeScript package's authoring API. Wire fields
retain the protocol's camelCase spelling, such as elementSchema, refPath, and
semanticPath.
API and results
| Task | Entry points | Results |
|---|---|---|
| Validate declarations | validate_schema, validate_operation, validate_operation_container, validate_structured_operation, validate_structured_operation_container, validate_schema_type_expr, validate_schema_type_term |
accepted, rejected |
| Load references | resolve_schema_references, resolve_operation_closure, and their _async variants (from schemair.resolve) |
resolved, incomplete, rejected |
| Validate data | validate_payload |
accepted, rejected, unknown |
| Compare contracts | check_schema_relation, check_operation_relation |
assignable, incompatible, unknown, rejected |
| Report compatibility | report_schema_compatibility, report_operation_compatibility |
compatible, breaking, unknown, invalid |
| Calculate types | evaluate_expression, evaluate_schema_type_term |
evaluated, incomplete, rejected |
| Compare type terms | check_type_term_relation |
Relation statuses |
| Infer variables | solve_type_variables |
solved, unknown, incompatible, rejected |
Calculation functions return a Result with status, issues, and value.
Reference loaders return SchemaReferencesResult with status, issues, and
graph.
Expression results put the evaluated term in value and expose unresolved
variable names through unresolved_type_vars; solver results put the binding
dictionary there. An evaluated term may still be an expression, so check its
shape before using it as a concrete schema. Diagnostics contain a code,
path, message, and optional details.
validate_expression(value, term=True) remains available for validating either
a schema or an expression. Definition validation does not execute operations
or resolve external references.
Compatibility reports run both substitution directions for a previous and next
declaration. The value contains previous_to_next and next_to_previous
checks, each with its mapped status and original relation result. The report
does not select migrations or a publication policy.
Solve a type variable
from schemair import authoring as s, evaluate_expression, solve_type_variables
solution = solve_type_variables(
["T"],
[{"source": s.number(), "target": s.type_var("T")}],
)
assert solution.status == "solved"
output = evaluate_expression(s.array_of(s.type_var("T")), type_vars=solution.value)
assert output.value == {"kind": "array", "elementSchema": s.number()}
The solver infers from supported source-to-target constraints. It is finite,
not a complete host-language generic type checker. Inspect unknown
and diagnostics when candidate relations cannot be proved.
Operations and host callbacks
An operation is a concrete declaration of input, output, errors, and emitted
channels. Structured payloads contain SchemaNode values. Functions and
unresolved expressions are not operation payload schemas. See the
operation model.
Calculations are synchronous; reference loading supports both execution modes.
The schemair.resolve module exposes resolve_schema_references and
resolve_operation_closure for synchronous resolvers, plus
resolve_schema_references_async and resolve_operation_closure_async for
synchronous or asynchronous resolvers. Resolvers return schemas or resolution
decisions and receive a tuple of path segments:
from schemair import authoring as s, validate_payload
from schemair.resolve import resolve_schema_references
def resolve_ref(path):
return s.string() if path == ("UserName",) else None
entry = s.ref(["UserName"])
loaded = resolve_schema_references(entry, resolve_ref=resolve_ref)
if loaded.status != "resolved":
raise ValueError(loaded.issues)
result = validate_payload("Ada", entry, refs=loaded.graph.refs)
assert result.status == "accepted"
For database or network access, use an async loader inside an async host workflow:
from schemair.resolve import resolve_schema_references_async
async def validate_from_catalog(entry, value, fetch_schema):
loaded = await resolve_schema_references_async(entry, resolve_ref=fetch_schema)
if loaded.status != "resolved":
raise ValueError(loaded.issues)
return validate_payload(value, entry, refs=loaded.graph.refs)
Both entry points share traversal, caching, limits, and diagnostics. The sync
loader refuses awaitable results with an incomplete diagnostic; the async
loader accepts both immediate and awaitable results. Only reference loading
awaits host I/O; calculations consume the resulting in-memory snapshot.
The loader recursively follows references inside resolved schemas. Its
ResolvedSchemaGraph contains refs entries shaped as
{"refPath": ["catalog", "User"], "schema": schema} and cycles as path
arrays. It preserves reference edges rather than inlining schemas. Inspect
resolved, incomplete, or rejected before using the snapshot; loader
limits bound traversal depth, steps, and newly loaded schemas.
Pass the same refs snapshot to payload validation, schema and operation
relations, expression evaluation, type-term relations, compatibility reports,
and the solver's context. These APIs are synchronous. The expression
interpreter accepts preloaded references only and does not invoke a resolver.
Payload and schema relation APIs also accept synchronous resolve_ref,
semantic_provider, and constraint_provider policies. Modern payload
providers receive (semantic_path, value, context) or
(constraint_path, args, value, context); the legacy two-argument forms remain
available. Relation providers receive source and target paths or constraint
lists, with an optional context. Term relations additionally accept
host_type_relation(source_path, target_path) for opaque host types.
Prepare external policy data before invoking the core. Awaitable policy results do not prove validation or assignability. Database uniqueness, permissions, and other business checks belong in the host workflow.
Standard vocabulary
schemair.standard exposes STANDARD_SEMANTIC_PATHS,
STANDARD_CONSTRAINT_PATHS, definition maps, standard_vocabulary_registry,
and classify_semantic_path. Standard validation is opt-in:
from schemair import authoring as s, standard, validate_payload
email = s.string(semantic_path=standard.STANDARD_SEMANTIC_PATHS["string"]["email"])
result = validate_payload("ada@example.com", email, **standard.standard_payload_context)
assert result.status == "accepted"
standard_schema_satisfiability, refine_standard_schema, and
intersect_standard_schemas expose the shared conservative standard constraint
profile, including binary64 interval
boundaries, exact float-derived multipleOf checks, pattern handling, and
standard argument domains. The relation context implements the shared
semantic narrowing and constraint implication rules; host regex/date behavior
still follows Python's libraries.
Host policies matter: Python uses its date/time, IP, and regex libraries, and
Python integers are unbounded, but
standard multipleOf converts numbers to binary64 and compares exact rational
representations. Do not assume arbitrary-precision decimal behavior or
identical regex/date acceptance across hosts. See
standard vocabulary host policies.
Integrations
from schemair import authoring as s, project_json_schema
schema = s.record([s.required_field("name", s.string())])
projection = project_json_schema(schema, target="draft-2020-12")
assert projection.fidelity == "exact"
print(projection.schema, projection.diagnostics)
Targets are draft-2020-12, draft-07, and openapi-3.0.
schema_node_to_json_schema and schema_node_to_openapi_schema return a
diagnostic-bearing projection. to_json_schema returns only the dictionary;
prefer the projection API when fidelity matters. Supported options include
both Python snake_case names and protocol-compatible camelCase aliases
for ref and mapper strategies. Fidelity results are exact, lossy, or
unsupported.
to_standard_schema(schema, **context) returns a ~standard adapter whose
validate callable is synchronous and preserves accepted input values. Load
external references before constructing the adapter.
schemaIR_to_standard_json_schema(schema) exposes jsonSchema.input and
jsonSchema.output callables; each takes a target string and raises ValueError
for unsupported projections. These are Python callable surfaces.
Development and verification scope
From python/, run:
python -m unittest discover -s tests -v
python -m compileall -q schemair
Tests read shared schema, operation, expression, solver, and vocabulary vectors
from ../conformance/, alongside local integration tests. Shared expression
vectors assert evaluated term contents and unresolved variables; solver vectors
assert bindings; rejected cases assert diagnostics; local tests cover provider
decision shapes, reference loading, resolver cycles, and interpreter budgets.
The binding does not execute operations, supply a scheduler or reference catalog, or provide a compile-time inference utility. See the roadmap for remaining parity and ecosystem work.
Metadata
Release files for schemair 0.4.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| schemair-0.4.1.tar.gz | 57.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| schemair-0.4.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 109.1 kB
Release files / schemair-0.4.1.tar.gz
| Download URL | schemair-0.4.1.tar.gz |
|---|---|
| Size | 57.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fa7e21236e43bd93a2011c1076e769d9276e92123cb67de20c5a9902de425b36
|
|
BLAKE2b-256 checksum How to use checksums |
f2dbf779ed265a64f04fd8b5d7beac7257d34f31283761bdc3247fd3d4013bca
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.10.11
|
Release files / schemair-0.4.1-py3-none-any.whl
| Download URL | schemair-0.4.1-py3-none-any.whl |
|---|---|
| Size | 51.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
71b7738e80d5bc032ac66ec3c4110bc79a9cf3e0269e25885b78bed8f224d5f6
|
|
BLAKE2b-256 checksum How to use checksums |
3a9e9e0e8b6265719c5892d3340f88dd2b48d305087a235849a3479ae4b2d100
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.10.11
|