Skip to main content

CPDataKit

CI Latest release PyPI License

Crystal Plasticity Data Quality Toolkit is a solver-independent Python toolkit for validating, normalizing, summarizing, and visualizing crystal-plasticity simulation datasets.

Alpha software: CPDataKit verifies conformance to an explicit data contract. It does not certify that a simulation, constitutive model, or physical interpretation is correct. Bundled datasets are deterministic, wholly synthetic examples for demonstration and tests.

Why and for whom

CPDataKit helps researchers, simulation engineers, and data stewards make tabular materials simulation data explicit and traceable before analysis or exchange. The core package is not a finite-element solver, DAMASK post-processor, Abaqus plug-in, UMAT runner, or ODB reader. It has no official affiliation with DAMASK, Abaqus, or Dassault Systèmes.

Supported contracts and formats

The open CPDataKit schema v1.0 defines three profiles:

  • curve: ordered macroscopic steps such as time, strain, stress, and load curves;
  • point: material-point, integration-point, element, or sample records;
  • field2d: scalar samples with two-dimensional Cartesian coordinates.

Inputs are UTF-8 CSV, JSON arrays of records, and CPDataKit HDF5 (.h5/.hdf5). CSV and JSON take units and semantics from the selected schema. HDF5 embeds units, mapping, validation summary, source filename and SHA-256, UTC conversion time, Python/CPDataKit versions, and an operation log. CPDataKit HDF5 is not DAMASK DADF5 or Abaqus ODB.

Schemas declare standard names, aliases, requiredness, dtype, per-record shape, role, unit, missing-value policy, index constraints, ranges, and scientific conventions. Custom fields must be declared or use user_. CPDataKit never guesses stress/strain measures, tensor order, orientation representation, units, or identifier semantics. See the data format.

Install

Install the current v0.2.0 release from PyPI:

python -m pip install cpdatakit

For a pinned GitHub release wheel, use:

python -m pip install "https://github.com/17636365690/cpdatakit/releases/download/v0.2.0/cpdatakit-0.2.0-py3-none-any.whl"

Then follow the five-minute quickstart to validate, summarize, convert, and plot a deterministic example.

Installing from the source checkout is intended for contributors:

git clone https://github.com/17636365690/cpdatakit.git
cd cpdatakit
python -m venv .venv

Activate on Windows PowerShell:

.venv\Scripts\Activate.ps1
python -m pip install -e ".[dev]"

Activate on POSIX shells:

source .venv/bin/activate
python -m pip install -e ".[dev]"

Regenerate the fixed-seed examples at any time:

python examples/generate_sample_data.py --output sample_data

Supported workflows

CPDataKit supports these concrete, repository-demonstrated workflows:

  • validate exported curve, point, or two-dimensional field records against an explicit contract;
  • normalize exporter-specific column names and units with a reviewable JSON mapping file;
  • preserve validated vectors and tensors in JSON/HDF5 with declared shapes and component order;
  • convert records into auditable HDF5 with units, mapping, provenance, and validation metadata;
  • run deterministic synthetic fixtures in notebooks, CI, and documentation examples.

These are supported workflows rather than claims of external adoption. No unverified downstream project or laboratory is presented as a CPDataKit user.

Project and integration links

Command line

Validate and write a JSON report:

cpdatakit validate sample_data/synthetic_curve.csv --schema curve --json-output validation.json

Summarize, convert, and create both image formats:

cpdatakit summary sample_data/synthetic_curve.csv --schema curve --json-output summary.json
cpdatakit convert sample_data/synthetic_curve.csv --schema curve --output curve.h5 --source-description "Synthetic README example"
cpdatakit plot curve.h5 --schema curve --kind stress-strain --output stress-strain.png
cpdatakit plot curve.h5 --schema curve --kind stress-strain --output stress-strain.svg

For an exporter with different names or units, provide an explicit mapping file:

cpdatakit convert raw.csv --schema curve --mapping mapping.json --output curve.h5

See the schema authoring and mapping guide for the JSON format and no-inference rules.

Use --force to replace an output. Expected errors are concise and have no traceback; put the global --debug option before the subcommand to debug unexpected failures. A validation or summary command returns 0 for conforming data and 1 for findings; usage/read/output failures return 2. Run cpdatakit --help or cpdatakit <command> --help for details.

Python API

from cpdatakit import (
    FieldMapping,
    load_dataset,
    normalize_dataset,
    summarize_dataset,
    validate_dataset,
)

raw = load_dataset("raw.csv")
normalized = normalize_dataset(
    raw,
    "curve",
    [
        FieldMapping("increment", "step", "1", "dimensionless"),
        FieldMapping("eps", "strain", "1", "dimensionless"),
        FieldMapping("sigma_pa", "stress", "Pa", "MPa", "export specification"),
    ],
)
report = validate_dataset(normalized, "curve")
summary = summarize_dataset(normalized, "curve", validation=report)
print(report.valid, summary["record_count"])

Mapping conflicts, unknown fields, and incompatible units raise documented subclasses of CPDataKitError. Normalization returns a copy and preserves unmapped columns unless drop_unmapped=True.

Example outputs

After running the commands above, stress-strain.png and stress-strain.svg contain a titled, unit-labeled synthetic curve with a legend. Plotting functions in cpdatakit.plotting return Matplotlib (Figure, Axes) for further editing and use the non-interactive Agg backend.

Development

pytest --cov=cpdatakit
ruff check .
ruff format --check .
python -m build

Architecture, extension boundaries, and maintainer checks are in the architecture documentation. Contributions follow CONTRIBUTING.md and the Code of Conduct.

If CPDataKit helps your workflow, star the repository to improve discoverability and open an issue describing the data contract or solver-neutral workflow you need. Concrete research use cases guide the roadmap more than raw popularity metrics.

Known limitations and roadmap

Version 0.2.0 handles in-memory tabular data, explicit vectors/tensors, and an explicit two-dimensional scalar sample representation. It does not perform solver integration, constitutive integration, 3D interactive graphics, automatic scientific inference, streaming, or distributed processing. DAMASK and Abaqus are only reserved adapter boundaries; no unverified adapter is shipped. See the roadmap for the next three versions.

Citation and license

Use CITATION.cff to cite the software. CPDataKit is licensed under Apache-2.0; see LICENSE. Direct runtime dependency licenses and review notes are in NOTICE. No real experimental or commercial-solver data are redistributed.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cpdatakit-0.2.0.tar.gz (255.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cpdatakit-0.2.0-py3-none-any.whl (31.7 kB view details)

Uploaded Python 3

File details

Details for the file cpdatakit-0.2.0.tar.gz.

File metadata

  • Download URL: cpdatakit-0.2.0.tar.gz
  • Upload date:
  • Size: 255.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for cpdatakit-0.2.0.tar.gz
Algorithm Hash digest
SHA256 41964bdd4d1ac0ee1d27379af11f84414dad233c6d0b19d00e1e908c018a5699
MD5 0a3e4f19a261e7b7d84828f4d6162552
BLAKE2b-256 19929e85c31ce30bc5f90c3c3ffe4d2885d11c9a6a55d675bc8a87d0cd8c7baa

See more details on using hashes here.

Provenance

The following attestation bundles were made for cpdatakit-0.2.0.tar.gz:

Publisher: publish-pypi.yml on 17636365690/cpdatakit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file cpdatakit-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: cpdatakit-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 31.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for cpdatakit-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 76f74d70f22dd643e434ee96ec6a2c79098be010c9669367e6b27c5ac4c50a4e
MD5 35640649b1d0fef88c9f9d6cb5bd625c
BLAKE2b-256 8767bfeed51b96a0ececefaecdc2efdf8706ccd2d462db8b8379030a63778f03

See more details on using hashes here.

Provenance

The following attestation bundles were made for cpdatakit-0.2.0-py3-none-any.whl:

Publisher: publish-pypi.yml on 17636365690/cpdatakit

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

This release

0.2.0 This release

2 files

0.1.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page