Skip to main content

relflow

Python 3.12+ Apache-2.0 license Documentation Discord channel invite

RelFlow builds PyTorch/Lightning models directly from JSON-like schemas. It is meant for predictive modeling on records that are not naturally flat: customers with transactions, orders with line items, sessions with clickstream events, devices recurring across histories, and mixed datatypes at every level.

Most ML pipelines flatten that shape first, then train on one fixed feature row. relflow takes the opposite path: describe the structured record, and the schema becomes the model.

Core Idea

A relflow schema is both a data contract and an architecture blueprint.

  • Leaf fields such as Number, Category, Set, Entity, Text, and Vector become datatype-specific tensorfields.
  • Branch nodes define shared contexts for child fields, with optional local attention and pooling before the representation flows upward.
  • Targets, masks, pruning, and embeddings are configured on the same schema tree.
  • Prediction output is keyed by schema address, so decoded values and embeddings remain attached to the part of the record that produced them.

That gives one model surface for supervised prediction, masked reconstruction, unsupervised embedding workflows, schema mutation, field importance, batch inference, and serving.

A Model From A Nested Record

import relflow as rf

model = rf.Model(
    name="order",
    d_model=64,
    n_layers=2,
    n_heads=4,
    embed=True,
    customer_tier=rf.Category(size=16),
    line_items=rf.Branch(
        length=32,
        embed=True,
        sku=rf.Category(size=2048),
        quantity=rf.Number,
        price=rf.Number,
    ),
    returned=rf.Category(target=True, size=2),
)

This model reads records shaped like:

{
    "customer_tier": "gold",
    "line_items": [
        {"sku": "A12", "quantity": 2, "price": 19.99},
        {"sku": "B07", "quantity": 1, "price": 45.50},
    ],
    "returned": "false",
}

The line_items branch has its own repeated context, returned is withheld from input and decoded as a supervised target, and embed=True asks prediction to emit embeddings at configured addresses.

Train With Lightning

rf.Model is a LightningModule. rf.PolarsDataModule and rf.StreamingDataModule are LightningDataModule implementations. The schema defines the model tree, typed losses, prediction outputs, and embeddings; Lightning runs fit, validate, test, and predict.

import lightning.pytorch as lit
import polars as pl
import torch

import relflow as rf

records = pl.read_ndjson("docs/data/iris.jsonl").head(36).with_row_index()
train_records = records.filter((pl.col("index") % 3) != 2).drop("index")
validate_records = records.filter((pl.col("index") % 3) == 2).drop("index")

model = rf.Model(
    d_model=16,
    n_layers=1,
    n_heads=4,
    batch_size=8,
    embed=True,
    optimizer=lambda module: torch.optim.AdamW(module.parameters(), lr=1e-2),
    sepal_length=rf.Number,
    petal_length=rf.Number,
    species=rf.Category(target=True, size=3, topk=[2]),
)

datamodule = rf.PolarsDataModule(
    model=model,
    train=train_records,
    validate=validate_records,
    num_workers=0,
    persistent_workers=False,
    pin_memory=False,
    observation_buffer_size=32,
    sample_rate=1.0,
)

trainer = lit.Trainer(
    accelerator="cpu",
    max_epochs=1,
    logger=False,
    enable_progress_bar=False,
    enable_model_summary=False,
    enable_checkpointing=False,
    limit_train_batches=1,
    limit_val_batches=1,
)

trainer.fit(model=model, datamodule=datamodule)

This tiny deterministic split is only a wiring example. Use a representative, leakage-safe validation design before interpreting the metrics as model quality.

For larger jobs, the same model can run through normal Lightning callbacks, checkpointing, precision settings, device placement, and distributed strategies. See Training With Lightning.

Predict And Embed

For small interactive batches, call model.predict(...) with raw dictionaries.

requests = validate_records.drop("species").head(3).to_dicts()
predictions = model.predict(requests)

species = predictions[rf.Address("record", "species")]
record = predictions[rf.Address("record")]

print(species["content"]["value"])
print(species["content"]["probability"])
print(record["embedding"])

For larger offline jobs, configure a predict split on a data module and attach rf.Writer to Lightning's prediction loop.

writer = rf.Writer("predictions")

trainer = lit.Trainer(
    accelerator="cpu",
    callbacks=[writer],
    logger=False,
)

predict_datamodule = rf.PolarsDataModule(
    model=model,
    predict=validate_records.drop("species"),
    num_workers=0,
    persistent_workers=False,
    pin_memory=False,
)

trainer.predict(
    model=model,
    datamodule=predict_datamodule,
    return_predictions=False,
)

Writer creates rank-partitioned Parquet files such as predictions/rank-0.parquet. Use a postprocessor when downstream systems need flat columns, renamed addresses, redacted payloads, or fewer fields. See Batch Inference and Postprocessors.

Learning Modes

relflow does not maintain separate supervised and self-supervised code paths. Supervised learning is the special case where a target field is hidden from the input 100% of the time and decoded from the remaining context.

Setting What the model sees What prediction can emit
plain input value is visible no decoded output unless otherwise configured
target=True value is hidden decoded supervised output
p_mask sampled configured leaf positions are hidden in train, validation, and test decoded reconstruction
p_prune whole leaf instances are hidden in train, validation, and test decoded reconstruction
embed=True does not hide the value embedding at that address

target=True is exact shorthand for p_prune=1.0. The current p_mask implementation samples configured positions, including null and padded positions, at rates lower than 1.0; see Dynamic Masking for exact selection behavior. Use embed=True when you want a representation returned from prediction.

Data Modules

Data modules load raw records, apply optional preprocessing, batch observations, tensorize values from the model schema, apply configured masking and target pruning in non-predict loops, and hand encoded batches to Lightning.

Choose the data module by where the records live:

Use case Module
Tutorials, tests, notebooks, in-memory Polars frames PolarsDataModule
Many local files StreamingDataModule
S3-backed datasets StreamingDataModule
Distributed training or prediction over large inputs StreamingDataModule

StreamingDataModule supports local ndjson, parquet, feather, csv, orc, and json inputs. S3 roots use the PyArrow-backed formats—parquet, feather, csv, orc, and json; ndjson is local-only. Avro is not supported by the current reader. Split arguments are compiled regular expressions matched against discovered file paths.

See Data Modules for split configuration, sharding, sampling, buffers, and preprocessors.

What Makes This Different

  • Hierarchical context encoding: child records interact locally before their representation flows upward.
  • Typed datatype architecture: each built-in field owns validation, tensorization, missing-state handling, masking, decoding, loss, metrics, and output writing. The external registration surface is experimental and same-process only in the current release.
  • Unified training roles: target=True, p_prune, and p_mask all use the same reconstruction path.
  • Embedding trees: embeddings can come from the root, branches, or selected leaves.
  • Schema evolution: fields can be added, removed, updated, reset, or temporarily overridden after construction.
  • Production missingness semantics: valued, null, padded, masked, and reserved other are distinct tensorfield states.
  • Training-serving parity: queries, preprocessors, tensorization, model execution, prediction writing, and postprocessors stay on the same configured path.

Where It Fits

Use relflow when relationships inside the record matter: account histories, fraud or risk snapshots, order and fulfillment events, flight itineraries, operations telemetry, user sessions, repeated measurements, or mixed datatype objects where flattening would discard useful structure.

Use a simpler tabular model when flattening loses no meaningful context. The point is not to replace every table. The point is to model nested business data without making a feature table the only representation the model can see.

What It Does Not Do

relflow stops at the representation and typed prediction layer. It is not a feature store, governance system, rule engine, authorization layer, decision-capture system, or audit platform. Those systems can consume relflow embeddings and predictions, but their policies and operational controls remain separate concerns.

The open-source layer is the reusable encoder and runtime infrastructure. It does not require users to publish data, schemas, checkpoints, or model parameters.

Install

RelFlow requires Python >=3.12. The currently documented user installation is directly from the GitHub repository:

python -m pip install "relflow @ git+https://github.com/relflow/relflow.git"

This follows the repository's default branch. Pin a tag or commit in the Git URL for reproducible environments; the published documentation currently tracks main and does not retain versioned snapshots.

Install optional functionality from the same source:

python -m pip install "relflow[text] @ git+https://github.com/relflow/relflow.git"
python -m pip install "relflow[serving] @ git+https://github.com/relflow/relflow.git"

Verify the environment:

python -c "import importlib.metadata; import relflow; print(importlib.metadata.version('relflow'))"

For a contributor checkout, use the locked development environment instead:

uv sync

Contributor extras:

uv sync --extra text
uv sync --extra serving
uv sync --extra docs

The text extra installs Hugging Face transformers. The serving extra installs FastAPI-backed deployment dependencies. The docs extra installs the Python packages used by the Quarto docs.

Documentation Map

Start with:

Tutorials and guides:

Build the docs locally with:

make render
uv run pytest tests/examples/test_e2e_examples.py

Repository Layout

  • src/relflow/architecture: model assembly, attention, pooling, and routing
  • src/relflow/data: dataset fetch/read/process/batch/encode pipeline and preprocessor exports
  • src/relflow/inference: serving and prediction callbacks
  • src/relflow/logging: runtime logging callbacks
  • src/relflow/structs: pydantic config models, enums, and tree nodes
  • src/relflow/tensorfields: tensorfield plugin system and built-in fields
  • tests/: package test suite
  • docs/: Quarto project, pages, guides, stylesheets, and sample data

Development

Run tests:

uv run pytest

Run type and lint checks:

uv run ty check src/relflow --output-format concise
uv run ruff check

Community

Join the relflow Discord for questions, design discussion, and release notes.

License

Licensed under the Apache License, Version 2.0. See LICENSE and NOTICE.

References

  • BIBLIOGRAPHY.md
  • CITATION.bib

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

relflow-0.1.1.tar.gz (115.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

relflow-0.1.1-py3-none-any.whl (142.1 kB view details)

Uploaded Python 3

File details

Details for the file relflow-0.1.1.tar.gz.

File metadata

  • Download URL: relflow-0.1.1.tar.gz
  • Upload date:
  • Size: 115.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for relflow-0.1.1.tar.gz
Algorithm Hash digest
SHA256 e78895ba263b014be9bf6f95cfa50b5c40df334ee1a3449f6b131c26e1aa7103
MD5 ea4ae6d677a05a0488cad382c2dfc713
BLAKE2b-256 68acd782cbb4885638562b851f2d15e99b88330f1ca9112088158896e5c7aadd

See more details on using hashes here.

Provenance

The following attestation bundles were made for relflow-0.1.1.tar.gz:

Publisher: release.yml on relflow/relflow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file relflow-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: relflow-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 142.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for relflow-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 5e6744a74f34cd8cd7b342e97fb66ffadfef82165a03f209043220eb51d4065a
MD5 4436410cca61965906fb63fb8d6e8b57
BLAKE2b-256 78c86c90398aaacb4cab7f7e6c14b1db76c4ef72cc66592f4e8d000f741eaeb8

See more details on using hashes here.

Provenance

The following attestation bundles were made for relflow-0.1.1-py3-none-any.whl:

Publisher: release.yml on relflow/relflow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

This release

0.1.1 This release

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page