Skip to main content

Lightweight, readable, and extensible validation for pandas DataFrames.

Project description

Project Logo

codecov

Lightweight, flexible, and intuitive validation for pandas DataFrames.
Define expectations for your data, validate them cleanly, and surface friendly errors or warnings — no configuration files, ever.


📦 Installation

pip install framecheck

Main Features

  • Designed for pandas users
  • Simple, fluent API
  • Supports error or warning-level assertions
  • Validates both column-level and DataFrame-level rules
  • No config files, decorators, or boilerplate

Table of Contents


🔥 Example: Catch data issues before they cause bugs

Example dataframe:

import pandas as pd
from framecheck import FrameCheck

df = pd.DataFrame({
    'a': [0, 1, 0, 1, 2],
    'b': [1, 1, 0, 0, 3],
    'timestamp': ['2022-01-01', '2022-01-02', '2019-12-31', '2021-01-01', '2023-05-01'],
    'email': ['a@example.com', 'bad', 'b@example.com', 'not-an-email', 'c@example.com'],
    'extra': ['x'] * 5
})

With FrameCheck:

validator = (
    FrameCheck()
    .columns(['a', 'b'], type='int', in_set=[0, 1])
    .column('timestamp', type='datetime', after='2020-01-01')
    .column('email', type='string', regex=r'.+@.+\..+', warn_only=True)
    .only_defined_columns()
    .row_count(min=5, max=100)
    .not_empty()
    .raise_on_error()
)

result = validator.validate(df)
  • .warn_only=True allows specific checks to issue warnings instead of failing validation.
  • .raise_on_error() makes the entire validation raise an exception if any non-warning failure occurs.
  • Together, they let you enforce hard rules while still being lenient on others.

For example, in the code above:

  • Invalid email formats will trigger a warning but not block execution.
  • Everything else (bad types, missing columns, out-of-bound values, etc.) will raise an error.

🧾 Output

If the data is invalid, you'll get warning ...

FrameCheckWarning: FrameCheck validation warnings:
- Column 'email' has values not matching regex '.+@.+\..+': ['bad'].
  result = validator.validate(df)

... and because you used .raise_on_error(), it'll raise a clean exception:

ValueError: FrameCheck validation failed:
Column 'a' is missing.
Column 'b' is missing.
Column 'timestamp' is missing.
DataFrame must have at least 5 rows (found 3).
Unexpected columns in DataFrame: ['good_credit', 'home_owner', 'id', 'promo_eligible', 'score']

Comparison with Other Approaches

Equivalent code in great_expectations (which is a fantastic package with a much broader focus than FrameCheck).

import great_expectations as gx

df['timestamp'] = pd.to_datetime(df['timestamp'])

context = gx.get_context(mode="ephemeral")
datasource = context.data_sources.add_pandas(name="pandas_src")
asset = datasource.add_dataframe_asset(name="df_asset")
batch_def = asset.add_batch_definition_whole_dataframe(name="df_batch")
batch = batch_def.get_batch({"dataframe": df})

batch.expect_column_values_to_be_in_set("a", [0, 1])
batch.expect_column_values_to_be_of_type("a", "int64")
batch.expect_column_values_to_be_in_set("b", [0, 1])
batch.expect_column_values_to_be_of_type("b", "int64")
batch.expect_column_values_to_be_of_type("timestamp", "datetime64[ns]")
batch.expect_column_values_to_be_between("timestamp", min_value="2020-01-01")
batch.expect_column_values_to_match_regex("email", r".+@.+\..+")
batch.expect_table_row_count_to_be_between(min_value=5, max_value=100)
batch.expect_table_row_count_to_be_greater_than(0)
batch.expect_table_columns_to_match_ordered_list(expected_column_names=["a", "b", "timestamp", "email"])
results = batch.validate()

if not results["success"]:
    raise ValueError("Validation failed")

Equivalent code without a package:

import pandas as pd
import re

errors = []

if df.empty:
    errors.append("DataFrame is empty.")

row_count = len(df)
if row_count < 5:
    errors.append("DataFrame has fewer than 5 rows.")
if row_count > 100:
    errors.append("DataFrame has more than 100 rows.")

for col in ['a', 'b']:
    if col not in df.columns:
        errors.append(f"Missing column: {col}")
    else:
        if not pd.api.types.is_integer_dtype(df[col]):
            errors.append(f"Column '{col}' is not of integer type.")
        if not df[col].isin([0, 1]).all():
            errors.append(f"Column '{col}' contains values outside [0, 1].")

if 'timestamp' not in df.columns:
    errors.append("Missing column: 'timestamp'")
else:
    try:
        ts = pd.to_datetime(df['timestamp'], errors='coerce')
        if ts.isna().any():
            errors.append("Column 'timestamp' contains non-datetime values.")
        elif (ts < pd.Timestamp('2020-01-01')).any():
            errors.append("Column 'timestamp' has values before 2020-01-01.")
    except Exception:
        errors.append("Could not convert 'timestamp' to datetime.")

if 'email' in df.columns:
    invalid_emails = df[~df['email'].astype(str).str.match(r'.+@.+\..+')]
    if not invalid_emails.empty:
        print("Warning: Some emails don't match expected pattern.")

expected_cols = {'a', 'b', 'timestamp', 'email'}
actual_cols = set(df.columns)
unexpected = actual_cols - expected_cols
if unexpected:
    errors.append(f"Unexpected columns in DataFrame: {sorted(unexpected)}")

# Final decision
if errors:
    raise ValueError("Validation failed:\n" + "\n".join(errors))

FrameCheck Methods

import pandas as pd
from framecheck import FrameCheck

column(...) – Core Behaviors

✅ Ensures column exists

df = pd.DataFrame({'x': [1, 2, 3]})

schema = FrameCheck().column('x')
result = schema.validate(df)
FrameCheck validation passed.

✅ Type enforcement

df = pd.DataFrame({'x': [1, 2, 'bad']})

schema = FrameCheck().column('x', type='int')
result = schema.validate(df)
FrameCheck validation errors:
- Column 'x' contains values that are not integer-like (e.g., decimals or strings): ['bad'].

.column(..., in_set=...) – Allowed values

df = pd.DataFrame({'status': ['new', 'active', 'archived']})

schema = FrameCheck().column('status', in_set=['new', 'active'])
result = schema.validate(df)
FrameCheck validation errors:
- Column 'status' contains values not in allowed set: ['archived'].

.column(..., equals=...) – All values must equal one thing

df = pd.DataFrame({'is_active': [True, False, True]})

schema = FrameCheck().column('is_active', type='bool', equals=True)
result = schema.validate(df)
FrameCheck validation errors:
- Column 'is_active' must equal True, but found values: [False].

.column(..., not_null=...) – All values non-null if set to True

df = pd.DataFrame({'is_active': [True, False, None]})

schema = FrameCheck().column('is_active', type='bool', not_null=True)
result = schema.validate(df)
FrameCheck validation errors:
- Column 'is_active' contains missing values.
  result = schema.validate(df)

.column(..., regex=...) – Pattern matching (for strings)

df = pd.DataFrame({'email': ['x@example.com', 'bademail']})

schema = FrameCheck().column('email', type='string', regex=r'.+@.+\..+')
result = schema.validate(df)
FrameCheck validation errors:
- Column 'email' has values not matching regex '.+@.+\..+': ['bademail'].

.column(..., min=..., max=..., after=..., before=...) – Range & bound checks

df = pd.DataFrame({
    'age': [25, 17, 101],
    'score': [0.9, 0.5, 1.2],
    'signup_date': ['2021-01-01', '2019-12-31', '2023-05-01'],
    'last_login': ['2020-01-01', '2026-01-01', '2023-06-15']
})

schema = (
    FrameCheck()
    .column('age', type='int', min=18, max=99)
    .column('score', type='float', min=0.0, max=1.0)
    .column('signup_date', type='datetime', after='2020-01-01', before='2025-01-01')
    .column('last_login', type='datetime', min='2020-01-01', max='2025-01-01')
)
result = schema.validate(df)
FrameCheck validation errors:
- Column 'age' has values less than 18.
- Column 'age' has values greater than 99.
- Column 'score' has values greater than 1.0.
- Column 'signup_date' violates 'after' constraint: 2020-01-01.
- Column 'last_login' violates 'max' constraint: 2025-01-01.

columns(...)

Any .column() operation can be applied to multiple columns of the same type.

df = pd.DataFrame({
    'a': [0, 1, 2],
    'b': [1, 0, 3],
    'c': [1, 1, 1]
})

schema = (
    FrameCheck()
    .columns(['a', 'b'], type='int', in_set=[0, 1])
)

result = schema.validate(df)
FrameCheck validation errors:
- Column 'a' contains values not in allowed set: [2].
- Column 'b' contains values not in allowed set: [3].

columns_are(...) – Exact column names and order

df = pd.DataFrame({'b': [1], 'a': [2]})

schema = FrameCheck().columns_are(['a', 'b'])
result = schema.validate(df)
FrameCheck validation errors:

Expected columns in order: ['a', 'b']

Found columns in order: ['b', 'a']

Go to Top

custom_check(...)

df = pd.DataFrame({
    'score': [0.2, 0.95, 0.6],
    'flagged': [False, False, True]
})

schema = (
FrameCheck()
.column('score', type='float')
.column('flagged', type='bool')
.custom_check(
    lambda row: row['score'] <= 0.9 or row['flagged'] is True,
    description="flagged must be True when score > 0.9"
)
)
result = schema.validate(df)
FrameCheck validation errors:

flagged must be True when score > 0.9 (failed on 1 row(s))

empty() – Ensure the DataFrame is empty

df = pd.DataFrame({'x': [1, 2]})

schema = FrameCheck().empty()
result = schema.validate(df)
FrameCheck validation errors:

DataFrame is expected to be empty but contains rows.

not_empty() – Ensure the DataFrame is not empty

df = pd.DataFrame(columns=['a', 'b'])

schema = FrameCheck().not_empty()
result = schema.validate(df)
FrameCheck validation errors:

DataFrame is unexpectedly empty.

only_defined_columns() – No extra/unexpected columns allowed

df = pd.DataFrame({'a': [1], 'b': [2], 'extra': [999]})

schema = (
FrameCheck()
.column('a')
.column('b')
.only_defined_columns()
)
result = schema.validate(df)
FrameCheck validation errors:

Unexpected columns in DataFrame: ['extra']

row_count(...) – Validate the number of rows

✅ Minimum rows

df = pd.DataFrame({'x': [1, 2]})

schema = FrameCheck().row_count(min=5)
result = schema.validate(df)
FrameCheck validation errors:

DataFrame must have at least 5 rows (found 2).

✅ Exact rows

df = pd.DataFrame({'x': [1, 2, 3]})

schema = FrameCheck().row_count(exact=2)
result = schema.validate(df)
FrameCheck validation errors:

DataFrame must have exactly 2 rows (found 3).

unique(...) – Rows must be unique

✅ All rows must be entirely unique

df = pd.DataFrame({
'user_id': [1, 2, 2],
'email': ['a@example.com', 'b@example.com', 'b@example.com']
})

schema = FrameCheck().unique()
result = schema.validate(df)
FrameCheck validation errors:

Rows are not unique.

✅ Rows must be unique based on specific columns

df = pd.DataFrame({
'user_id': [1, 2, 2],
'email': ['a@example.com', 'b@example.com', 'c@example.com']
})

schema = FrameCheck().unique(columns=['user_id'])
result = schema.validate(df)
FrameCheck validation errors:

Rows are not unique based on columns: ['user_id']

Go to Top


License

MIT


Contact

LinkedIn Badge


Go to Top

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

framecheck-0.4.1.tar.gz (23.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

framecheck-0.4.1-py3-none-any.whl (23.1 kB view details)

Uploaded Python 3

File details

Details for the file framecheck-0.4.1.tar.gz.

File metadata

  • Download URL: framecheck-0.4.1.tar.gz
  • Upload date:
  • Size: 23.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.8.20

File hashes

Hashes for framecheck-0.4.1.tar.gz
Algorithm Hash digest
SHA256 a05c652fac082267857e8c6523a14615f6907b49c17c1828d15b396d390808bf
MD5 c7eb97a7951be421eea9c6fe6d230f69
BLAKE2b-256 2268b63b6da84d0ff6b32a21ec8ce11999ef366e02458088097351c3fea3c818

See more details on using hashes here.

File details

Details for the file framecheck-0.4.1-py3-none-any.whl.

File metadata

  • Download URL: framecheck-0.4.1-py3-none-any.whl
  • Upload date:
  • Size: 23.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.8.20

File hashes

Hashes for framecheck-0.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 d0b67e5dbeb1e959ba9691ffd494cc83714e7b07344aee886f025f0c2e68e6da
MD5 8b9a6170e4f2995ffac9862e70f07012
BLAKE2b-256 5efbc2a9ba0dad14485b782c56504531b9b9c3e4784a2daa42ab08d0981121e1

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page