Skip to main content

toolkit-data-contracts

License: Apache-2.0

Contracts for LLM-shaped data. A small, dependency-free Python CLI that checks chat transcripts, tool-call arguments, SFT and preference (DPO) rows and RAG chunk records against a JSON Schema contract, and flags drift at every nested path. Contracts interoperate with the Open Data Contract Standard (ODCS). It is meant as a cheap CI gate in front of training, evaluation or retrieval data.

toolkit-contracts check --input train.jsonl --preset openai-chat --tools tools.json

It reads JSONL (and, with the parquet extra, Parquet or CSV) in one streaming pass. If you need warehouse connectors or a full suite of statistical drift tests, look at larger projects such as datacontract-cli or Evidently.

Status

Capability Status Notes
Contract format Working {"version": 2, "schema": <JSON Schema>}. A bare JSON Schema file and version 1 contracts (fields/types/required) are also accepted. See Contracts.
JSON Schema validation Working Draft 2020-12 subset: type, enum, const, numeric ranges and multipleOf, string lengths and pattern, items/prefixItems/contains, array sizes and uniqueItems, properties/required/additionalProperties/patternProperties/propertyNames, dependentRequired/dependentSchemas, allOf/anyOf/oneOf/not/if-then-else, same-document $ref. Checked against the official JSON-Schema-Test-Suite. Unsupported keywords make check fail with exit 2 instead of passing.
Contract inference (infer) Working Nested objects (properties, required), arrays (items, merged over all elements), nullability (null in type), optional enum for low-cardinality strings (--enum-max N).
Validation (check) Working Issues are counted per kind and path (meta.source, messages[].role). Integer and number are one JSON numeric type unless --strict-integer is set.
Baseline profile (profile) Working Per declared path, including nested keys and array items: missing rate (relative to the parent object), JSON type counts, numeric mean/std/min/max, and value counts for string/boolean values with up to 50 distinct values.
Drift: missing rate Working Flags a path whose missing rate is above --max-missing (default 0.01) and higher than the baseline's.
Drift: null rate Working Flags a rise of more than --max-null-increase (default 0.10, absolute) in the share of present values that are null.
Drift: categorical distribution Working Population Stability Index (PSI) on string/boolean value counts, threshold --max-psi (default 0.25). Paths with more than 50 distinct baseline values are skipped as high-cardinality.
Drift: text length and token count Working Every string path: character and token counts in power-of-two bins (0, 1, 2-3, 4-7, ...), compared with PSI (--max-length-psi, default 0.25). Tokens are whitespace-split by default or counted with profile --tokenizer tiktoken:<encoding> (extra tiktoken).
Drift: language distribution Working (opt-in) profile --language PATH detects each text's language with langdetect (extra lang; texts under 20 non-space characters count as und); PSI against --max-psi.
Drift: embedding centroid Working (opt-in) profile --embedding PATH stores the mean vector of that path; check flags a Euclidean (or --centroid-metric cosine) distance above --max-centroid-distance (default 0.2) or a dimension change. Pure Python; tune the threshold for your embedding model.
Drift: numeric mean Partial Flags abs(mean - baseline_mean) / baseline_std > --max-mean-shift-sigma (default 3.0). This is a heuristic, not a statistical test: it ignores batch size and distribution shape.
Drift on nested keys and array items Working Every path the contract declares is profiled and drift-checked.
Drift on keys the contract does not declare Planned Undeclared keys are not profiled. Set additionalProperties: false (infer --disallow-extra) to reject them in validation.
Statistical tests (KS, chi-squared) Planned Not implemented.
Large files / streaming Working infer, profile and check read the input once, in constant memory apart from the schema, the profile and at most 51 category counts per path (the rag-chunks duplicate-id check keeps the ids).
Parquet / CSV input Working --input-format auto (by file extension), jsonl, parquet or csv; needs the parquet extra (pyarrow). Rows are read batch by batch; dates become ISO 8601 strings, decimals numbers, binary base64.
Contract diff / compatibility (diff) Working Classifies each change as tightening, loosening, both or annotation, and as breaking under --mode backward (default), forward or full. Changes inside anyOf/oneOf/not/if are reported as "not analysed" (both directions). GitHub Action included. See Reviewing contract changes.
LLM presets Working openai-chat, anthropic-messages, tool-calls, sft, dpo, rag-chunks. Tool-call arguments are validated against each tool's JSON Schema. See LLM presets.
ODCS v3 import / export Working odcs export writes ODCS v3.1.0 (validated against the official ODCS JSON Schema in tests); odcs import reads ODCS v3.0-v3.2 including v3.2 enum, map and vector. See ODCS.
Report envelope v1 (check and diff JSON output) Working in-toto Statement v1, canonical JSON, input digests; see Report format.
PyPI package Planned Not published yet; install from source. The release workflow is ready and waits on the one-time PyPI setup in RELEASING.md.

Install

Requires Python 3.10+. The core has no runtime dependencies. Optional extras: odcs (PyYAML, for ODCS YAML files), tiktoken (tiktoken token counts), lang (langdetect, for language drift), parquet (pyarrow, for Parquet and CSV input).

git clone https://github.com/AKIVA-AI/toolkit-data-contracts.git
cd toolkit-data-contracts
pip install .            # or ".[odcs,tiktoken,lang,parquet]"; pip install -e ".[dev]" for development
toolkit-contracts --help

Five-minute example

examples/data/ultrachat_200k_test_sft_40.jsonl holds 40 chat rows from the public UltraChat 200k dataset (MIT license; see examples/data/README.md). From the repository root:

# 1. Does the file have the shape trainers expect for SFT?
toolkit-contracts check --input examples/data/ultrachat_200k_test_sft_40.jsonl --preset sft --format table
Status: PASS

Validation: OK (no issues)

Drift: OK (no issues)
# 2. Freeze its shape as a contract (roles become an enum) and profile it as the baseline
toolkit-contracts infer --input examples/data/ultrachat_200k_test_sft_40.jsonl \
  --out chat.contract.json --enum-max 3
toolkit-contracts profile --input examples/data/ultrachat_200k_test_sft_40.jsonl \
  --contract chat.contract.json --out baseline.profile.json

# 3. Make a bad batch: assistant answers cut to 20 characters, a system turn added
python - <<'EOF'
import json
with open("examples/data/ultrachat_200k_test_sft_40.jsonl", encoding="utf-8") as src, \
        open("bad.jsonl", "w", encoding="utf-8") as out:
    for line in src:
        row = json.loads(line)
        for m in row["messages"]:
            if m["role"] == "assistant":
                m["content"] = m["content"][:20]
        row["messages"].insert(0, {"role": "system", "content": "You are terse."})
        out.write(json.dumps(row) + "\n")
EOF

# 4. Gate it (exit code 4)
toolkit-contracts check --input bad.jsonl --contract chat.contract.json \
  --baseline baseline.profile.json --format table
Status: FAIL

Validation Issues:
  Kind                      Field                 Count  Message
  ------------------------- -------------------- ------  ------------------------------
  enum_mismatch             messages[].role          40  enum

Drift Issues:
  Kind                      Field                 Count  Message
  ------------------------- -------------------- ------  ------------------------------
  drift_text_length         messages[].content        1  psi=5.707 exceeds 0.250 (mean 902.6 -> 154.5)
  drift_token_count         messages[].content        1  psi=5.642 exceeds 0.250 (mean 144.3 -> 25.4)
  drift_categorical         messages[].role           1  psi=0.941 exceeds 0.250
# 5. Share the contract as an Open Data Contract Standard document (needs the odcs extra)
toolkit-contracts odcs export --contract chat.contract.json --out chat.odcs.yaml --name ultrachat_sft

tests/test_readme_walkthrough.py runs these steps in CI.

Usage

# 1. Infer a contract from known-good records (reads the whole file; --limit N to cap)
toolkit-contracts infer --input samples.jsonl --out contract.json

# 2. Profile known-good records as the drift baseline
toolkit-contracts profile --input baseline.jsonl --contract contract.json --out baseline.profile.json

# 3. Validate a new batch, and drift-check it against the baseline
toolkit-contracts check --input new_batch.jsonl --contract contract.json \
  --baseline baseline.profile.json

check prints a JSON report, or writes it with --out FILE. Other formats: --format table and --format markdown (for a PR comment) are for people; --format json-legacy prints the pre-1.0 {ok, validation_issues, drift_issues} object and will be removed in the next minor version.

Report format

The JSON report is the shared toolkit report envelope v1: an in-toto Statement v1 in canonical JSON (sorted keys, no extra whitespace, trailing newline), so its SHA-256 is stable and it can be signed. See docs/report-envelope.md and the JSON Schema schemas/report-envelope.v1.json.

  • subject: the checked batch file, with its SHA-256.

  • predicate.kind: data.check.

  • predicate.inputs: the contract (or preset:NAME), baseline profile and --tools files, with their SHA-256.

  • predicate.verdict / exit_code: pass/0, fail/4, or error/2 (the input was readable but could not be judged, for example a malformed line). If the input file itself cannot be read, no report is written.

  • predicate.summary (what a CI gate reads):

    Key Meaning
    ok true only when there are no validation or drift issues
    records records read from the batch
    validation_issues distinct (kind, field, message) validation issues
    invalid_occurrences total occurrences across those issues
    drift_checked whether --baseline was given
    drift_issues number of drift issues
    invalid_records records with at least one validation issue
    tool_calls_checked, tool_calls_unchecked with the openai-chat, anthropic-messages and tool-calls presets

    diff reports use kind: data.diff with summary keys mode, changes, breaking, non_breaking and ok, and a details.changes list.

  • predicate.details: validation_issues and drift_issues lists (kind, field, message, count), the drift thresholds used, and the preset name.

Set SOURCE_DATE_EPOCH to fix created_at and make a report byte-for-byte reproducible. Signing is optional and not built in; for example, with toolkit-ml-provenance installed:

toolkit-mlsbom sign-file report.json    # and later: toolkit-mlsbom verify-file report.json

Useful check options:

Option Default Meaning
--strict-integer off Reject fractional values on integer-only fields.
--max-missing 0.01 Missing-rate ceiling.
--max-null-increase 0.10 Allowed absolute rise in null rate.
--max-psi 0.25 Allowed categorical PSI (below 0.1 is usually read as stable, above 0.25 as a significant shift).
--max-mean-shift-sigma 3.0 Allowed mean shift in baseline standard deviations.
--max-length-psi 0.25 Allowed PSI of text length and token-count histograms.
--max-centroid-distance 0.2 Allowed distance between embedding centroids (--centroid-metric euclidean or cosine).
--metrics-out FILE none Write pass/fail counters as JSON.

Global options: --verbose, --log-format json, --version.

Text and embedding drift

profile records text statistics for every string path automatically. The tokenizer, the paths with language detection and the embedding paths are stored in the profile's config, and check profiles the new batch the same way, so they are chosen once:

toolkit-contracts profile --input base.jsonl --preset rag-chunks --out base.profile.json   --tokenizer tiktoken:cl100k_base --language text --embedding embedding
toolkit-contracts check --input new.jsonl --preset rag-chunks --baseline base.profile.json

Comparing profiles made with different tokenizers is an error (exit 2).

LLM presets

Built-in contracts for common LLM data shapes. Use --preset NAME instead of --contract with check and profile; toolkit-contracts presets list lists them and presets show NAME prints the JSON Schema so you can copy and tighten it.

Preset Row shape Extra record checks
openai-chat {"messages": [...], "tools"?: [...]}; roles system/developer/user/assistant/tool, text or content-part arrays, assistant tool_calls, weight tool-call arguments (a JSON string) parsed and validated against the tool's parameters; unknown tools; tool messages whose tool_call_id answers no earlier call
anthropic-messages {"system"?, "messages": [...], "tools"?}; first message from user; text, image, document, tool_use, tool_result, thinking, redacted_thinking blocks (other block types pass with a type) tool_use.input validated against the tool's input_schema; unknown tools; tool_result blocks with no matching tool_use
tool-calls one call per row: {"name", "arguments"}, {"type": "function", "function": {...}} or a tool_use block arguments validated against --tools (required)
sft messages with at least one assistant turn, prompt + non-empty completion, or non-empty text
dpo prompt? + chosen + rejected (strings or message lists), or OpenAI input + preferred_output + non_preferred_output identical chosen/rejected; chosen and rejected of different types
rag-chunks id, non-empty text, optional doc_id, source, title, chunk_index, start_char/end_char, metadata, embedding duplicate ids; embeddings of different lengths; start_char > end_char

--tools FILE takes OpenAI tools ([{"type": "function", "function": {"name", "parameters"}}]), legacy OpenAI functions or Anthropic tools ([{"name", "input_schema"}]), as a list or as {"tools": [...]}. Tools inside a row take precedence. Without any tool definitions, tool calls are counted as tool_calls_unchecked in the report summary instead of failing. Tool arguments use strict JSON Schema integer rules (2.0 is an integer, 2.5 is not). Issue paths look like tool[get_weather].location. The duplicate-id check keeps every chunk id in memory.

Contracts

infer --enum-max 2 on a chat file writes a contract like this (the real file has its keys sorted):

{
  "version": 2,
  "schema": {
    "$schema": "https://json-schema.org/draft/2020-12/schema",
    "type": "object",
    "properties": {
      "id": {"type": "string"},
      "messages": {
        "type": "array",
        "items": {
          "type": "object",
          "properties": {
            "role": {"type": "string", "enum": ["assistant", "user"]},
            "content": {"type": ["null", "string"]}
          },
          "required": ["content", "role"],
          "additionalProperties": true
        }
      }
    },
    "required": ["id", "messages"],
    "additionalProperties": true
  }
}

Edit it by hand to tighten it (pattern, minimum/maximum, maxLength, minItems, additionalProperties: false, and so on), or pass any JSON Schema file that stays within the supported subset as --contract. Issue kinds: missing_required, type_mismatch, unexpected_field, enum_mismatch, const_mismatch, range_violation, length_violation, pattern_mismatch, unique_violation, contains_violation, unexpected_item, invalid_property_name, combinator_mismatch, false_schema. When no branch of an anyOf/oneOf matches, the issues of the closest branch are reported. pattern uses Python regular expressions, which differ from ECMA-262 in rare cases (for example \p{...} is rejected).

Input must be strict JSON: NaN, Infinity and numbers that overflow are rejected with exit code 2. Profiles written by older versions still work; regenerate the baseline to get nested paths and categorical counts.

ODCS (Open Data Contract Standard)

ODCS is the Bitol (Linux Foundation) standard for data contracts, also used by tools such as datacontract-cli. Convert in either direction:

pip install ".[odcs]"    # PyYAML; .json output and input work without it
toolkit-contracts odcs export --contract contract.json --out contract.odcs.yaml --name chat_sft
toolkit-contracts odcs import --input contract.odcs.yaml --out contract.json [--object NAME]

How the mapping works:

JSON Schema ODCS property
one non-null type logicalType (date/timestamp/time import as strings with a format)
type without null required: true (ODCS required means "not null")
nested key in required parent's logicalTypeOptions.required
length, range, pattern, item-count keywords logicalTypeOptions
items, nested properties items, properties
enum / const quality rule metric: invalidValues, arguments.validValues, mustBe: 0 (and v3.2 enum on import)
(import) map, vector additionalProperties: <value schema>; array of numbers with minItems = maxItems = dimensions

ODCS has no presence flag for top-level columns, so a top-level key exports as required: true only when it is both required and non-null. Constructs ODCS cannot express (anyOf, additionalProperties: false, $ref, a required-but-nullable top-level key, and so on) are listed as a warning, and the exact JSON Schema is embedded in the schema object as the custom property jsonSchema. odcs import uses that exact schema only while the ODCS fields still describe it; if someone edited the fields, it uses the fields and warns. Quality rules other than a strict valid-values list are reported, not enforced.

Library use

from toolkit_data_contracts_drift.contract import (
    drift_check, infer_contract, profile_records, validate_records,
)

contract = infer_contract(training_rows)
issues = validate_records(contract=contract, records=new_rows)
drift = drift_check(
    baseline=profile_records(contract=contract, records=training_rows),
    current=profile_records(contract=contract, records=new_rows),
)

See examples/ml_pipeline_example.py for a runnable walkthrough.

Exit codes

Code Meaning
0 Success: validation passed, no drift detected
2 CLI error: invalid arguments, missing files, bad input
3 Unexpected error, or interrupted
4 Check failed: validation issues or drift detected

Reviewing contract changes

diff compares two contracts and classifies every change:

Direction Meaning Examples
tightens the new contract rejects data the old one accepted: existing data and producers may start failing new required field, narrower type, removed enum value, lower maxLength, additionalProperties: false
loosens the new contract accepts data the old one rejected: consumers may see values they do not expect nullable field, new enum value, wider range, field no longer required
both either can happen changed pattern, changes inside anyOf/oneOf/not/if (not analysed further)
annotation no effect on validity description, title, examples, format

--mode backward (default, as in schema registries) treats tightens and both as breaking; forward treats loosens and both as breaking; full treats all three as breaking. The exit code is 4 when anything is breaking.

toolkit-contracts diff --old main-contract.json --new contract.json --format table

Formats: json (report envelope, kind: data.diff), table, markdown (for a PR comment) and github (workflow annotations).

GitHub Action

on: pull_request
jobs:
  contract:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0          # the action reads the base version with git show
      - uses: AKIVA-AI/toolkit-data-contracts@main   # pin a release tag once one exists
        with:
          contract: contracts/chat.contract.json
          mode: backward

The action writes a Markdown summary to the job page, annotates breaking changes on the PR, fails the step when a change is breaking, and exposes the JSON report path as the report output. Pass old and new instead of contract to compare two files directly.

CI example

- name: Data contract gate
  run: |
    pip install git+https://github.com/AKIVA-AI/toolkit-data-contracts.git
    toolkit-contracts check --input data/batch.jsonl \
      --contract contracts/contract.json --baseline profiles/baseline.profile.json

Docker

docker build -t toolkit-data-contracts .
docker run --rm -v "$PWD/data:/app/data" toolkit-data-contracts \
  toolkit-contracts check --input /app/data/batch.jsonl --contract /app/data/contract.json

docker-compose.yml has example infer and check services under the tools profile.

Development

See CONTRIBUTING.md. In short: pytest, ruff check src/ tests/, pyright.

Contributing and security

Contributions are welcome: see CONTRIBUTING.md and the Code of Conduct. Please report security problems privately, as described in SECURITY.md.

Releasing

Releases are cut by pushing a vX.Y.Z tag. CI runs the tests, builds the sdist and wheel, checks them, attaches them to a GitHub Release and publishes them to PyPI with Trusted Publishing. RELEASING.md describes the process and how to verify a release.

License

Apache License 2.0. See LICENSE and NOTICE. Releases before the relicensing remain available under the MIT License.

Metadata

Release files for toolkit-data-contracts 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for toolkit-data-contracts 1.0.0
File Size Uploaded
toolkit_data_contracts-1.0.0.tar.gz 100.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for toolkit-data-contracts 1.0.0
File Interpreter ABI Platform
toolkit_data_contracts-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 164.3 kB

Release files / toolkit_data_contracts-1.0.0.tar.gz

Download URL toolkit_data_contracts-1.0.0.tar.gz
Size 100.8 kB
Tags Source
SHA-256 checksum
How to use checksums
6d19fe94971df9f6aee274d541c0b06f7417e0d539089e03f99b5531e8d44378
BLAKE2b-256 checksum
How to use checksums
98c6170d80df0ed18595b379348d21aeb7c9d3f3579273020553f2cdc5a361e2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / toolkit_data_contracts-1.0.0-py3-none-any.whl

Download URL toolkit_data_contracts-1.0.0-py3-none-any.whl
Size 63.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
000807fab3813850c77e842ab43f5e1af966bd05d71258c025f4ce7e5b16ee21
BLAKE2b-256 checksum
How to use checksums
10fe199240368d9da671f7e8f2216eba8b98e996624b9a49266a6f95b2967a0b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page