Skip to main content

lochan-eda

Exploratory data analysis and behaviour-driven preprocessing for tabular machine learning.

Documentation Case Study PyPI version Python versions License GitHub

lochan-eda is a Python package for exploratory data analysis and preprocessing of tabular datasets.

It provides reusable components for inspecting numerical and categorical features, analysing missing values, handling common preprocessing tasks, and constructing a consistent train/test preprocessing workflow.

The package is designed around a simple principle:

Understand the behaviour of the data before choosing how to preprocess it.


Installation

Install the latest release from PyPI:

pip install lochan-eda

For development:

git clone https://github.com/LochanJangid/lochan-eda.git
cd lochan-eda

pip install -e .

Quick Start

Automated workflow

import pandas as pd

from lochan_eda import AutomatedEDA

df = pd.read_csv("data.csv")

eda = AutomatedEDA()

X_train, X_test, y_train, y_test = eda.prepare(
    df,
    target="target"
)

prepare() provides a high-level interface for the common tabular preprocessing workflow.


Dataset Profiling

For explicit dataset inspection, use Profiler:

from lochan_eda import Profiler

profile = Profiler(df)

profile.overview()

The profiler provides access to dataset-level analysis and the numerical and categorical analysis components.


Numerical Analysis

Numerical features are handled through the Numerical component.

from lochan_eda import Numerical

numerical = Numerical(df)

numerical.summary()
numerical.plot()

The numerical interface includes operations for:

  • Missing-value handling
  • Outlier management
  • Scaling
  • Statistical summaries
  • Visualization

Available methods include:

imputer()
outlier_manager()
scaler()
summary()
plot()

Categorical Analysis

Categorical features are handled separately through Categorical.

from lochan_eda import Categorical

categorical = Categorical(df)

categorical.summary()
categorical.plot()

The categorical interface provides operations for:

  • Missing-value handling
  • Rare-category management
  • Encoding
  • Statistical summaries
  • Visualization

Available methods include:

imputer()
rare_manager()
encoder()
summary()
plot()

Missing Values

Missing-value analysis is available independently:

# Example interface

missing.plot()

This allows missing-value patterns to be inspected before deciding how they should be handled.


Preprocessing Philosophy

A common preprocessing workflow can be written directly with scikit-learn:

from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

preprocessor = ColumnTransformer([
    (
        "numeric",
        Pipeline([
            ("imputer", SimpleImputer(strategy="median")),
            ("scaler", StandardScaler())
        ]),
        numeric_columns
    ),
    (
        "categorical",
        Pipeline([
            ("imputer", SimpleImputer(strategy="most_frequent")),
            ("encoder", OneHotEncoder(
                handle_unknown="ignore"
            ))
        ]),
        categorical_columns
    )
])

This is a valid and useful approach.

The problem is not the pipeline itself.

The problem is deciding whether the selected transformations are appropriate for the dataset.

For example:

median imputation
       ↓
Why?

standard scaling
       ↓
Why?

most-frequent imputation
       ↓
Why?

one-hot encoding
       ↓
Why?

lochan-eda provides a layer for analysing the dataset before those decisions are applied.

                  Dataset
                     │
                     ▼
              ┌─────────────┐
              │   Profile   │
              └──────┬──────┘
                     │
          ┌──────────┴──────────┐
          ▼                     ▼
     Numerical             Categorical
     behaviour               behaviour
          │                     │
          └──────────┬──────────┘
                     ▼
             Preprocessing
               workflow
                     │
                     ▼
              Model-ready data

The package therefore complements preprocessing libraries rather than attempting to replace them.


Train / Test Workflow

Preprocessing should be learned from training data and then reused when transforming other data.

AutomatedEDA supports this separation:

eda.fit(X_train)

X_train = eda.transform(X_train)
X_test = eda.transform(X_test)

This keeps the fitting of preprocessing separate from its application.

For the common workflow, prepare() provides a single entry point:

X_train, X_test, y_train, y_test = eda.prepare(
    df,
    target="target"
)

API

AutomatedEDA

High-level interface for the complete preprocessing workflow.

prepare()
fit()
transform()

Profiler

Dataset-level analysis interface.

overview()
numerical
categorical

Numerical

Numerical feature analysis and preprocessing.

imputer()
outlier_manager()
scaler()
summary()
plot()

Categorical

Categorical feature analysis and preprocessing.

imputer()
rare_manager()
encoder()
summary()
plot()

Missing

Missing-value visualization.

plot()

Report

Generate a shareable analysis report.

save()

Design

lochan-eda separates the workflow into two levels.

High-level API

Use AutomatedEDA when the goal is to move efficiently from a DataFrame to model-ready data.

eda = AutomatedEDA()

X_train, X_test, y_train, y_test = eda.prepare(
    df,
    target="target"
)

Component-level API

Use Profiler, Numerical, and Categorical when explicit control or inspection is required.

profile = Profiler(df)

profile.overview()

profile.numerical
profile.categorical

This keeps the common workflow simple without removing access to the underlying analysis components.


Package Structure

The project is organized around the major stages of tabular data analysis and preprocessing:

lochan_eda/
│
├── automated_eda/
├── profiler/
├── numerical/
├── categorical/
├── missing/
└── report/

The internal implementation may evolve independently from the public API.


Dependencies

lochan-eda is built around the Python data-science ecosystem and integrates with commonly used tools for tabular machine learning.

The package is intended to work alongside libraries such as:

  • pandas
  • NumPy
  • scikit-learn
  • matplotlib

Rather than replacing these libraries, lochan-eda provides a higher-level workflow for analysis and preprocessing.


Example

A complete workflow can be as small as:

import pandas as pd

from lochan_eda import Profiler
from lochan_eda import AutomatedEDA

df = pd.read_csv("data.csv")

# Inspect
profile = Profiler(df)
profile.overview()

# Prepare
eda = AutomatedEDA()

X_train, X_test, y_train, y_test = eda.prepare(
    df,
    target="target"
)

# Continue with model training
model.fit(X_train, y_train)

predictions = model.predict(X_test)

Development

Clone the repository:

git clone https://github.com/LochanJangid/lochan-eda.git
cd lochan-eda

Install in editable mode:

pip install -e .

If you are developing new functionality, install the development dependencies defined by the project.

Run the test suite with the project's configured test command.


Contributing

Contributions are welcome.

Before submitting a change:

  1. Keep the public API consistent with the existing design.
  2. Add or update tests for behavioural changes.
  3. Keep preprocessing decisions explicit and reproducible.
  4. Avoid introducing unnecessary dependencies.
  5. Document new public functionality.

For larger changes, open an issue first to discuss the proposed design.


License

See the repository's license file for the applicable license.


Links


Status

lochan-eda is an actively developed project. The API may evolve as additional preprocessing and analysis capabilities are introduced.

Metadata

Release files for lochan-eda 0.2.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lochan-eda 0.2.6
File Size Uploaded
lochan_eda-0.2.6.tar.gz 18.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lochan-eda 0.2.6
File Interpreter ABI Platform
lochan_eda-0.2.6-py3-none-any.whl Python 3 none any Details

Total release size: 35.7 kB

Release files / lochan_eda-0.2.6.tar.gz

Download URL lochan_eda-0.2.6.tar.gz
Size 18.9 kB
Tags Source
SHA-256 checksum
How to use checksums
0bfec5b785eb29aaddcc88d0e7557bb825834f9906622ba9d4c6a81520c7b3b1
BLAKE2b-256 checksum
How to use checksums
7662fa63e58e2b5d35acf30eab00f890148b68e012f6fa11dab074c5213b49ef
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / lochan_eda-0.2.6-py3-none-any.whl

Download URL lochan_eda-0.2.6-py3-none-any.whl
Size 16.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d7f11579e246e8269327c5646dea7be9436fa5f1f8622b2ff5d50f4ed8986e93
BLAKE2b-256 checksum
How to use checksums
40d1aeb7e00711acc78930cc058f809598a002244b3c7084028a064c7ec870cf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.7

2 release files

This release

0.2.6 This release

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page