Skip to main content
This is a pre-production deployment of Warehouse. Changes made here affect the production instance of PyPI (pypi.python.org).
Help us improve Python packaging - Donate today!

Validate the pandas objects such as DataFrame and Series.

Project Description

Validates the pandas object such as DataFrame and Series. And this can define validator like django form class.

Why bugs occur in Data Wrangling with pandas

When we wrangle our data with pandas, We use DataFrame frequently. DataFrame is very powerfull and easy to handle. But DataFrame has no it’s schema, so It allows irregular values without being aware of it. We are confused by these values and affect the results of data wrangling.

pandas-schema offers the functions for validating DataFrame or Series objects and generating factory data.

Overview

import pandas as pd
import pandas_validator as pv

class SampleDataFrameValidator(pv.DataFrameValidator):
    row_num = 5
    column_num = 2
    label1 = pv.IntegerColumnValidator('label1', min_value=0, max_value=10)
    label2 = pv.FloatColumnValidator('label2', min_value=0, max_value=10)

validator = SampleDataFrameValidator()

df = pd.DataFrame({'label1': [0, 1, 2, 3, 4], 'label2': [5.0, 6.0, 7.0, 8.0, 9.0]})
validator.is_valid(df)  # True.

df = pd.DataFrame({'label1': [11, 12, 13, 14, 15], 'label2': [5.0, 6.0, 7.0, 8.0, 9.0]})
validator.is_valid(df)  # False.

df = pd.DataFrame({'label1': [0, 1, 2], 'label2': [5.0, 6.0, 7.0]})
validator.is_valid(df)  # False

Getting Started

Requirements

  • Support python version: 2.7, 3.4, 3.5, 3.6
  • Support pandas version: 0.18, 0.19

Installation

$ pip install pandas_validator

Usage

Please see the following demo written by ipython notebook.

License

This software is licensed under the MIT License.

Resources

CHANGES

0.5.0 (2017-01-06)

  • Add LambdaColumnValidator
  • Add IndexValidator
  • .validate(df) method is deprecated. Please use .is_valid(df, raise_exception=True)

0.4.0 (2015-10-28)

  • Hot fix: cannot include source file

0.3.2 (2015-10-28)

  • Python 2.7, 3.2, 3.3, 3.4, 3.5 support
  • pandas 0.14, 0.15, 0.16, 0.17 support

0.3.1 (2015-10-28)

  • Update support python version
  • Update dependencies library version

0.3.0 (2015-07-15)

  • Critical bug fix

0.2.0 (2015-05-24)

  • Support char type validation
  • flake8 testing

0.1.0 (2015-05-22)

Initial release.

  • Support integer series validator
  • Support float series validator
  • Support dataframe validator
  • Testing on python2.7 and python 3.4

0.0.0 (2015-05-17)

Create this project.

Release History

Release History

This version
History Node

0.5.0

History Node

0.4.0

History Node

0.3.2

History Node

0.3.1

History Node

0.3.0

History Node

0.2.0

History Node

0.1.0

Download Files

Download Files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

File Name & Checksum SHA256 Checksum Help Version File Type Upload Date
pandas_validator-0.5.0-py3-none-any.whl (11.1 kB) Copy SHA256 Checksum SHA256 py3 Wheel Jan 6, 2017
pandas_validator-0.5.0.tar.gz (7.1 kB) Copy SHA256 Checksum SHA256 Source Jan 6, 2017

Supported By

WebFaction WebFaction Technical Writing Elastic Elastic Search Pingdom Pingdom Monitoring Dyn Dyn DNS Sentry Sentry Error Logging CloudAMQP CloudAMQP RabbitMQ Heroku Heroku PaaS Kabu Creative Kabu Creative UX & Design Fastly Fastly CDN DigiCert DigiCert EV Certificate Rackspace Rackspace Cloud Servers DreamHost DreamHost Log Hosting