Skip to main content

VisEDA — Visual Exploratory Data Analysis

VisEDA is a Python toolkit for exploratory data analysis of image, video, hyperspectral, point-cloud, and text/NLP datasets. It provides numerical summaries, dataset-level visualisations, per-sample diagnostics, duplicate or similarity analysis where applicable, command-line workflows, and self-contained HTML reports.

Modules

  • ImageEDA — spatial, pixel, quality, colour, texture, frequency, duplicate, class-balance, and normalisation analysis.
  • VideoEDA — spatial, temporal, motion, blur, scene-change, colour, and video-similarity analysis.
  • HyperspectralEDA — per-band statistics, SNR/noise, spectral quality, vegetation/water indices, PCA, false-colour, texture, and spectral-diversity analysis.
  • PointCloudEDA — geometry, density, height, duplicate/outlier, nearest-neighbour, PCA shape-descriptor, and attribute analysis.
  • TextEDA — length, vocabulary, lexical diversity, symbols, readability, writing scripts, duplicates, TF-IDF document distances, and label analysis.

Installation

pip install viseda

Optional file-format support:

# MATLAB/ENVI/GeoTIFF hyperspectral files
pip install "viseda[hyperspectral]"

# LAS/LAZ point clouds
pip install "viseda[pointcloud]"

# All optional file-format dependencies
pip install "viseda[all]"

Python 3.9 or later is required.

Quick start

ImageEDA

from viseda import ImageEDA

eda = ImageEDA(verbose=True)
eda.load("path/to/images", label_from_parent=True)

summary = eda.summary()
print(summary["inventory"])
print(summary["quality"])

eda.plot(save_path="image_dashboard.png")
eda.report("image_report.html")

In-memory images are supported through load_arrays().

VideoEDA

from viseda import VideoEDA

eda = VideoEDA(verbose=True, frame_sample_rate=5)
eda.load("path/to/videos", label_from_parent=True)

print(eda.summary()["temporal"])
eda.plot_dataset(save_path="video_dashboard.png")
eda.report("video_report.html")

NumPy video arrays can be loaded with load_arrays().

HyperspectralEDA

import numpy as np
from viseda import HyperspectralEDA

cube = np.random.default_rng(0).random((128, 128, 103)).astype("float32")
wavelengths = np.linspace(400, 1500, cube.shape[2])

np.save("scene.npy", cube)

eda = HyperspectralEDA(
    wavelengths=wavelengths,
    compute_glcm=False,
    compute_pca=True,
)
eda.load("scene.npy")

print(eda.summary()["spectral_quality"])
ndvi = eda.compute_index(cube_index=0, index_name="ndvi")
scores, variance_ratio = eda.pca_scores(cube_index=0, n_components=3)

eda.plot(save_path="hyper_dashboard.png")
eda.report("hyper_report.html")

The hyperspectral extra adds support for MATLAB .mat, ENVI, and multi-band GeoTIFF files.

PointCloudEDA

import numpy as np
from viseda import PointCloudEDA

points = np.random.default_rng(0).random((10000, 3)).astype("float32")

eda = PointCloudEDA(
    max_points_per_cloud=200000,
    compute_neighbors=True,
    compute_geometry=True,
)
eda.load_arrays([points], labels=["sample"])

print(eda.summary()["geometry"])
eda.plot_dataset(save_path="pointcloud_dashboard.png")
eda.report("pointcloud_report.html")

The pointcloud extra adds LAS/LAZ support. NPY, NPZ, TXT, CSV, XYZ, PTS, and ASCII PLY are supported by the base installation.

TextEDA

from viseda import TextEDA

eda = TextEDA()
eda.load_texts(
    [
        "Exploratory data analysis is useful before model training.",
        "TextEDA summarises vocabulary, length, readability and duplicates.",
    ],
    labels=["eda", "eda"],
)

print(eda.summary()["lexical"])
print(eda.vocabulary(top_n=10))

eda.plot_dataset(save_path="text_dashboard.png")
eda.report("text_report.html")

TextEDA also loads TXT, Markdown, HTML, CSV, TSV, JSON, JSONL and NDJSON datasets through load().

Command line

The package installs the viseda command:

viseda --help

Available subcommands are:

viseda image ...
viseda hyper ...
viseda cloud ...
viseda video ...
viseda text ...

Examples:

viseda image "C:/datasets/images" --label-from-parent --plot --report image_report.html

viseda video "C:/datasets/videos" --label-from-parent --plot --report video_report.html

viseda hyper "C:/datasets/hyper" --label-from-parent --plot --dataset-plot \
  --report hyper_report.html

viseda cloud "C:/datasets/pointclouds" --label-from-parent --plot \
  --report pointcloud_report.html

viseda text "C:/datasets/text" --label-from-parent --plot \
  --report text_report.html

Use viseda <subcommand> --help for the options of a particular modality.

Supported input formats

Module Main file inputs
ImageEDA JPG/JPEG, PNG, BMP, TIFF and other formats accepted by the current image loader
VideoEDA MP4, AVI, MOV, MKV, WEBM, MPEG/MPG, M4V
HyperspectralEDA MAT, NPY, NPZ, ENVI HDR/BIL/BIP/BSQ/ENVI, TIF/TIFF
PointCloudEDA NPY, NPZ, TXT, CSV, XYZ, PTS, ASCII PLY, LAS/LAZ
TextEDA TXT/TEXT, MD, RST, LOG, HTML/HTM, CSV/TSV, JSON, JSONL/NDJSON

Some file formats require the optional extras described above.

Development

Install the project in editable mode with development tools:

python -m pip install -e ".[dev]"
pytest

Build and validate a release:

python -m pip install --upgrade build twine
python -m build
python -m twine check dist/*

License

MIT. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

viseda-1.0.0.tar.gz (91.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

viseda-1.0.0-py3-none-any.whl (91.9 kB view details)

Uploaded Python 3

File details

Details for the file viseda-1.0.0.tar.gz.

File metadata

  • Download URL: viseda-1.0.0.tar.gz
  • Upload date:
  • Size: 91.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.13

File hashes

Hashes for viseda-1.0.0.tar.gz
Algorithm Hash digest
SHA256 8cc62962e84a95169469fb269ef83ae92faf5938ee85760aa6d6d750b84e5e99
MD5 908d1ee3a1326f00a99b20c9df482d77
BLAKE2b-256 c7ba54a41c2f6547856b785b6eae9c3df106fc35f99501646fad71298bc114a1

See more details on using hashes here.

File details

Details for the file viseda-1.0.0-py3-none-any.whl.

File metadata

  • Download URL: viseda-1.0.0-py3-none-any.whl
  • Upload date:
  • Size: 91.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.10.13

File hashes

Hashes for viseda-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 123d12a1781f295b12096398167eb447122c4d9c541f34aa6a5028262317c7f7
MD5 dfd50d48436951723ac3953523a43c18
BLAKE2b-256 873258b3680d5d53d6253a17b3fb456faae308c2c6634a344fa220a30fe71cae

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page