lochan-eda
A practical toolkit for exploring and preparing tabular data for machine learning.
lochan-eda provides two simple workflows:
- Explore a dataset with summaries, plots, missing-value analysis, and a PDF report.
- Prepare tabular data with automatic numerical and categorical preprocessing.
Ready to Use
Installation
pip install lochan-eda
1. Automatically Prepare Data for Machine Learning
Use AutomatedEDA when you want to prepare a tabular dataset before training a machine-learning model.
import pandas as pd
from lochan_eda import AutomatedEDA
# Load your data
df = pd.read_csv("data.csv")
eda = AutomatedEDA()
X_train, X_test, y_train, y_test = eda.prepare(
df,
target="target"
)
By default, prepare():
- separates the target column
- creates an 80/20 train-test split
- learns preprocessing from the training data
- applies the same learned preprocessing to the test data
- processes numerical and categorical columns separately
You can then pass the processed data directly to your model.
from sklearn.ensemble import RandomForestClassifier
model = RandomForestClassifier(random_state=42)
model.fit(X_train, y_train)
Exclude columns
Keep columns such as IDs out of automatic preprocessing:
eda = AutomatedEDA()
X_train, X_test, y_train, y_test = eda.prepare(
df,
target="target",
exclude=["customer_id"]
)
Use your own target Series
X = df.drop(columns="target")
y = df["target"]
eda = AutomatedEDA()
X_train, X_test, y_train, y_test = eda.prepare(
X,
target=y
)
Prepare without a train-test split
X, y = eda.prepare(
df,
target="target",
split=False
)
If there is no target, prepare() can also process feature data only:
X = eda.prepare(
df,
split=False
)
2. Profile a Dataset
Use Profiler when you want to understand a dataset before modeling.
import pandas as pd
from lochan_eda import Profiler
df = pd.read_csv("data.csv")
profile = Profiler(df, target="target")
Dataset overview
profile.overview()
This returns information about:
- rows and columns
- memory usage
- duplicate rows
- duplicate percentage
- missing cells
- missing percentage
- numerical columns
- categorical columns
- datetime columns
Numerical analysis
profile.numerical.summary()
profile.numerical.plot()
Categorical analysis
profile.categorical.summary()
profile.categorical.plot()
Missing-value analysis
profile.missing.plot()
Generate a PDF report
profile.report.save("eda_report.pdf")
The report contains dataset overview information, numerical and categorical summaries, and generated analysis plots.
3. Analyze Only Selected Columns
Numerical and categorical plots can be limited to selected columns.
profile.numerical.plot(
columns=["age", "income"]
)
profile.categorical.plot(
columns=["city", "education"],
top_n=10
)
top_n keeps the most frequent categories visible and groups the remaining categories as Others.
API Overview
AutomatedEDA
from lochan_eda import AutomatedEDA
The main interface for automatic tabular-data preprocessing.
AutomatedEDA()
AutomatedEDA()
Creates an automatic preprocessing object.
prepare()
prepare(
X,
target=None,
exclude=None,
split=True,
test_size=0.2,
random_state=42,
stratify=None
)
The high-level method for preparing data.
| Parameter | Description |
|---|---|
X |
Input pandas.DataFrame. |
target |
Target column name or pandas.Series. |
exclude |
Column name or list of columns to exclude from preprocessing. |
split |
Whether to create train/test data. Default: True. |
test_size |
Proportion used for the test set. Default: 0.2. |
random_state |
Random state for reproducibility. Default: 42. |
stratify |
Values used for stratified splitting. |
Return values
The returned values depend on target and split:
target |
split |
Returns |
|---|---|---|
| Not provided | False |
X |
| Not provided | True |
X_train, X_test |
| Provided | False |
X, y |
| Provided | True |
X_train, X_test, y_train, y_test |
fit()
fit(X, y=None)
Learns preprocessing rules from the supplied data.
fit() separates numerical and categorical columns and learns the required transformations for each type.
transform()
transform(X, y=None)
Applies preprocessing rules learned by fit().
transform() must be called after fit().
Profiler
from lochan_eda import Profiler
Provides dataset-level exploratory analysis.
Profiler()
Profiler(df, target=None)
| Parameter | Description |
|---|---|
df |
Input pandas.DataFrame. |
target |
Target column name. |
When target is a column name, the target is separated from the profiling data.
overview()
profile.overview()
Prints and returns a dictionary containing:
rowscolumnsmemory_usageduplicate_rowsduplicate_percentagemissing_cellsmissing_percentagenumerical_columnscategorical_columnsdatetime_columns
Numerical
from lochan_eda import Numerical
Handles numerical-column profiling and preprocessing.
Numerical()
Numerical(profiler_df=None)
profiler_df is used by the profiling methods such as summary() and plot().
fit()
fit(data, exclude=None)
Learns numerical preprocessing rules from the supplied data.
The numerical workflow can learn rules for:
- missing-value imputation
- column removal based on missingness
- outlier handling
- scaling
transform()
transform(data, exclude=None)
Applies the rules learned by fit().
summary()
profile.numerical.summary()
Returns a DataFrame containing descriptive numerical statistics together with missing-value, skewness, zero, and outlier information.
plot()
profile.numerical.plot(columns=None)
Creates numerical analysis plots for each selected numerical column:
- distribution
- box plot
- Q-Q plot
The figure is also saved as numerical_plots.png.
imputer()
imputer(learn=False)
Internal numerical preprocessing step for learning or applying missing-value handling.
outlier_manager()
outlier_manager(learn=False)
Internal numerical preprocessing step for learning or applying outlier rules.
scaler()
scaler(learn=False)
Internal numerical preprocessing step for learning or applying feature scaling.
imputer(),outlier_manager(), andscaler()are lower-level methods. For normal usage, preferAutomatedEDAorNumerical.fit()/Numerical.transform().
Categorical
from lochan_eda import Categorical
Handles categorical-column profiling and preprocessing.
Categorical()
Categorical(profiler_df=None)
profiler_df is used by the profiling methods.
fit()
fit(data, exclude=None)
Learns categorical preprocessing rules.
The categorical workflow can learn rules for:
- missing-value handling
- rare-category handling
- binary encoding
- one-hot encoding
- frequency encoding
transform()
transform(data, exclude=None)
Applies the categorical preprocessing rules learned by fit().
summary()
profile.categorical.summary()
Returns descriptive statistics for categorical columns.
plot()
profile.categorical.plot(columns=None, top_n=10)
Creates category-distribution bar charts.
The figure is also saved as categorical_plots.png.
imputer()
imputer(learn=False)
Internal categorical preprocessing step for learning or applying missing-value handling.
rare_manager()
rare_manager(learn=False)
Internal preprocessing step that learns or applies rare-category grouping.
encoder()
encoder(learn=False)
Internal preprocessing step that selects and applies categorical encoding based on category cardinality.
imputer(),rare_manager(), andencoder()are lower-level methods. For normal usage, preferAutomatedEDAorCategorical.fit()/Categorical.transform().
Missing
from lochan_eda.missing import Missing
Provides missing-value visualization.
Missing()
Missing(profiler_df=None)
Creates a missing-value analysis object.
plot()
profile.missing.plot()
Creates a horizontal bar chart showing missing-value percentages by column.
The figure is also saved as missing_plot.png.
Report
from lochan_eda.report import Report
Generates an EDA PDF report from a Profiler instance.
Report()
Report(profiler)
save()
profile.report.save("report.pdf")
Generates and saves the exploratory data analysis report.
Default path:
profile.report.save()
which creates:
report.pdf
Utility Functions
The package also contains helper functions in lochan_eda.utils.
get_iqr_bounds()
from lochan_eda.utils import get_iqr_bounds
lower, upper = get_iqr_bounds(series)
Returns lower and upper bounds calculated using the IQR method.
get_active_cols()
from lochan_eda.utils import get_active_cols
columns = get_active_cols(all_cols, exclude=["id"])
Returns columns after removing excluded columns.
is_matching()
from lochan_eda.utils import is_matching
is_matching(fit_df, transform_df)
Checks whether the two DataFrames contain matching column sets.
Typical Workflow
A common machine-learning workflow with lochan-eda looks like this:
import pandas as pd
from lochan_eda import AutomatedEDA
from sklearn.ensemble import RandomForestClassifier
# 1. Load data
df = pd.read_csv("data.csv")
# 2. Prepare data
eda = AutomatedEDA()
X_train, X_test, y_train, y_test = eda.prepare(
df,
target="target"
)
# 3. Train your model
model = RandomForestClassifier(random_state=42)
model.fit(X_train, y_train)
# 4. Predict
predictions = model.predict(X_test)
For exploration before modeling:
from lochan_eda import Profiler
profile = Profiler(df, target="target")
profile.overview()
profile.numerical.summary()
profile.categorical.summary()
profile.missing.plot()
profile.report.save("eda_report.pdf")
Package
- Package:
lochan-eda - Author: Lochan Jangid
- Version:
0.2.0
Metadata
Release files for lochan-eda 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| lochan_eda-0.2.1.tar.gz | 18.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| lochan_eda-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 35.8 kB
Release files / lochan_eda-0.2.1.tar.gz
| Download URL | lochan_eda-0.2.1.tar.gz |
|---|---|
| Size | 18.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8da1c9164177712f91c89034be65e2796f72b347ecf248efc7856f8b6d7dcc36
|
|
BLAKE2b-256 checksum How to use checksums |
86924a6227a8abbd069b79464df09d81043bfd5a36494f5241f72a913884a747
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency logRelease files / lochan_eda-0.2.1-py3-none-any.whl
| Download URL | lochan_eda-0.2.1-py3-none-any.whl |
|---|---|
| Size | 16.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
545381c82807e98cfd9294703cdf8e75ecff5475931661b1f6f5770529aca642
|
|
BLAKE2b-256 checksum How to use checksums |
4cd509a39a3600c899629da1b5632ea2b4c8bb46bc6643103287dc9eaaf58dd4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency log