Skip to main content

Tiny, extensible pandas utilities for aliasing and basic data cleaning

Project description

Cleanroom 🧹

Tiny, extensible pandas utilities for aliasing and basic data cleaning

Cleanroom provides a simple set of pure functions you can compose to clean messy real-world data. It's designed to be lightweight, extensible, and work seamlessly with pandas DataFrames.

🚀 Features

  • Flexible Column Matching: Automatically detects column variations (first_name, First Name, f name, fname, etc.)
  • Email Cleaning: Normalize email addresses to lowercase
  • Name Formatting: Smart title case with proper handling of particles (van, de, etc.)
  • Phone Standardization: Convert phone numbers to international format
  • Address Normalization: Clean ZIP codes, standardize states and countries
  • Schema-Based Cleaning: Apply custom cleaning rules with flexible schemas
  • Auto-Clean: Automatically clean common columns with zero configuration

📦 Installation

pip install cleanroom

🔧 Quick Start

import pandas as pd
import cleanroom

# Sample messy data
df = pd.DataFrame({
    'f name': ['  john ', 'JANE'],
    'e-mail': ['JOHN@EXAMPLE.COM', 'jane@TEST.org'],
    'pnumber': ['(555) 123-4567', '555.987.6543'],
    'zip code': ['12345-6789', '98765'],
    'st': ['California', 'TX']
})

# One-line cleaning with flexible column matching
cleaned_df = cleanroom.auto_clean(df)
print(cleaned_df)

Output:

  f name e-mail              pnumber    zip code st
0   John  john@example.com   +15551234567  123456789  CA
1   Jane  jane@test.org      +15559876543     98765  TX

📚 API Reference

Individual Cleaning Functions

clean_email(series)

Normalize email addresses to lowercase and strip whitespace.

emails = pd.Series(['  JOHN@EXAMPLE.COM  ', 'jane@TEST.org'])
cleaned = cleanroom.clean_email(emails)
# Result: ['john@example.com', 'jane@test.org']

clean_name(series, alias_map=None, case="title")

Clean and format names with smart title casing.

names = pd.Series(['  john smith  ', 'JANE DOE', 'bob o\'connor'])
cleaned = cleanroom.clean_name(names)
# Result: ['John Smith', 'Jane Doe', 'Bob O\'Connor']

# With custom aliases
aliases = {'johnny': 'John', 'bobby': 'Robert'}
cleaned = cleanroom.clean_name(names, alias_map=aliases)

clean_phone(series, default_country="US")

Standardize phone numbers to international format.

phones = pd.Series(['(555) 123-4567', '555.987.6543', '+1-800-555-0199'])
cleaned = cleanroom.clean_phone(phones)
# Result: ['+15551234567', '+15559876543', '+18005550199']

clean_number(series)

Extract only digits (useful for ZIP codes, IDs).

zips = pd.Series(['12345-6789', 'ABC 98765', '  54321  '])
cleaned = cleanroom.clean_number(zips)
# Result: ['123456789', '98765', '54321']

clean_state(series, state_map=None)

Standardize US state names to abbreviations.

states = pd.Series(['California', 'TX', 'new york'])
cleaned = cleanroom.clean_state(states)
# Result: ['CA', 'TX', 'NY']

clean_country(series, country_map=None)

Standardize country names to ISO codes.

countries = pd.Series(['United States', 'USA', 'United Kingdom'])
cleaned = cleanroom.clean_country(countries)
# Result: ['US', 'US', 'GB']

DataFrame Operations

auto_clean(df)

Automatically clean common columns with flexible name matching.

Recognizes these column patterns:

  • Email: email, e-mail, mail, email_address, e_mail
  • Names: first_name, First Name, fname, f_name, given_name
  • Phone: phone, telephone, pnumber, phone_number, mobile
  • Address: zip, zip_code, postal_code, state, country
# Works with any column naming convention
df_messy = pd.DataFrame({
    'First Name': ['john', 'JANE'],
    'family name': ['smith', 'DOE'],  
    'E-Mail Address': ['JOHN@EXAMPLE.COM', 'jane@test.org'],
    'telephone': ['(555) 123-4567', '555.987.6543']
})

cleaned = cleanroom.auto_clean(df_messy)

apply_schema(df, schema)

Apply custom cleaning rules with a flexible schema.

schema = {
    'clean_email': {
        'func': cleanroom.clean_email,
        'source': ['email', 'email_address', 'e_mail']  # Try multiple columns
    },
    'full_name': {
        'func': cleanroom.clean_name,
        'source': ['first_name', 'last_name'],  # Combine columns
        'kwargs': {'case': 'title'}
    }
}

cleaned_df = cleanroom.apply_schema(df, schema)

🌍 Flexible Column Matching

Cleanroom automatically handles various column naming conventions:

Data Type Recognized Patterns
First Name first_name, First Name, fname, f_name, f name, firstname, given_name
Last Name last_name, Last Name, lname, surname, family_name, lastname
Email email, e-mail, mail, email_address, e_mail
Phone phone, telephone, pnumber, phone_number, mobile, cell
ZIP Code zip, zip_code, zipcode, postal_code, postcode
State state, st, province, region
Country country, nation, nationality

🏗️ Design Philosophy

  • Simple: Small set of pure functions you can compose
  • Extensible: Pass custom alias maps and cleaning rules
  • Flexible I/O: Works on pd.Series, DataFrame columns, or entire DataFrames
  • Non-destructive: Always returns new objects, never modifies input data
  • Pandas-native: Leverages pandas' powerful string operations

🧪 Examples

Real-world messy data

import pandas as pd
import cleanroom

# Typical messy customer data
customers = pd.DataFrame({
    'f name': ['  alice  ', 'BOB', 'Charlie Brown'],
    'surname': ['SMITH', 'jones', 'o\'connor'],
    'e-mail': ['ALICE@GMAIL.COM', 'bob@YAHOO.com', '  charlie@test.org  '],
    'pnumber': ['555-1234', '(800) 555-0199', '+1.212.555.9876'],
    'zip code': ['12345-6789', 'ABC 90210', '10001'],
    'st': ['California', 'TX', 'new york'],
    'nation': ['USA', 'United States', 'US']
})

# Clean everything with one function call
clean_customers = cleanroom.auto_clean(customers)

Custom cleaning pipeline

# Build your own cleaning pipeline
def clean_customer_data(df):
    result = df.copy()
    
    # Clean emails
    if 'email' in result.columns:
        result['email'] = cleanroom.clean_email(result['email'])
    
    # Standardize names with custom aliases
    name_aliases = {'bobby': 'Robert', 'mike': 'Michael'}
    if 'first_name' in result.columns:
        result['first_name'] = cleanroom.clean_name(
            result['first_name'], 
            alias_map=name_aliases
        )
    
    return result

📋 Requirements

  • Python 3.8+
  • pandas >= 1.3.0
  • numpy >= 1.21.0

🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

📄 License

MIT License - see LICENSE file for details.

🔗 Links

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pandas_cleanroom-1.3.1.tar.gz (10.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pandas_cleanroom-1.3.1-py3-none-any.whl (8.9 kB view details)

Uploaded Python 3

File details

Details for the file pandas_cleanroom-1.3.1.tar.gz.

File metadata

  • Download URL: pandas_cleanroom-1.3.1.tar.gz
  • Upload date:
  • Size: 10.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.6

File hashes

Hashes for pandas_cleanroom-1.3.1.tar.gz
Algorithm Hash digest
SHA256 df8a5661405e9cfe66295ce09650ac2eb35f4ef85222772f1de184f7545a7edd
MD5 198134eb1780be466d89ec7ed67bc57f
BLAKE2b-256 1b59d505d12ff2f8eaa86a46991a4f9078c5dca4a634e9b7b27ac0e025b74fcd

See more details on using hashes here.

File details

Details for the file pandas_cleanroom-1.3.1-py3-none-any.whl.

File metadata

File hashes

Hashes for pandas_cleanroom-1.3.1-py3-none-any.whl
Algorithm Hash digest
SHA256 49f65a3553a0d5be8251137c253e98a92e9111f6d9e8bf79004977458624889e
MD5 01d29fc808f2b76fa3811faccc264821
BLAKE2b-256 7c29aedcf923d53972fae2368bb622d4662ee791fe95de79555d1188c03f6242

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page