Skip to main content

A scikit-learn compatible feature selector using the Interval Chi-Square Score (ICSS) algorithm.

Project description

Interval Chi-Square Score (ICSS) Feature Selection

A Python implementation of the Interval Chi-Square Score (ICSS) feature selection algorithm for interval-valued data. This library provides a scikit-learn compatible feature selector that ranks and selects features based on their statistical independence from class labels.

Overview

The Interval Chi-Square Score (ICSS) is a feature selection technique designed specifically for interval-valued data, where each feature is represented as an interval $[a_i, b_i]$ rather than a single continuous value. This approach is particularly useful in domains where data naturally comes as ranges or intervals, such as:

  • Temperature ranges in climate studies
  • Price ranges in e-commerce
  • Uncertainty quantification in measurements
  • Survey response ranges

The ICSS algorithm evaluates the degree of independence between each interval-valued feature and the target class label using a chi-square based approach adapted for interval data.

Features

  • ✨ Scikit-learn compatible API (fit, transform, fit_transform)
  • 📊 Works with interval-valued features (shape: (n_samples, n_features, 2))
  • 🎯 Selects top-k features based on ICSS scores
  • 📈 Integrates seamlessly with scikit-learn pipelines
  • 🧬 Based on peer-reviewed research

Installation

Via PyPI (Recommended)

pip install icss-feature-selection

See the PyPI package page for more information.

From Source

git clone https://github.com/vinaykumarngitub/icss-feature-selection.git
cd icss-feature-selection
pip install -e .

Requirements

  • Python >= 3.8
  • NumPy
  • scikit-learn

Quick Start

Basic Usage

import numpy as np
from icss_feature_selection import ICSSFeatureSelector

# Create dummy interval data: (n_samples, n_features, 2)
# Each feature is represented as [lower_bound, upper_bound]
n_samples = 100
X = np.random.rand(n_samples, 4, 2)
X = np.sort(X, axis=2)  # Ensure [min, max] order

# Create labels correlated with some features
y = np.zeros(n_samples)
y[(X[:, 0, 0] > 0.1) & (X[:, 0, 1] < 0.2)] = 1
y[(X[:, 2, 0] > 0.7) & (X[:, 2, 1] < 0.9)] = 1

# Select top 2 features
selector = ICSSFeatureSelector(k=2)
selector.fit(X, y)

# Check ICSS scores for each feature
print("ICSS Scores:", selector.scores_)

# Transform data to keep only selected features
X_selected = selector.transform(X)
print("Selected features shape:", X_selected.shape)  # (100, 2, 2)

Using with Scikit-learn Pipelines

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

# Create a pipeline
pipeline = Pipeline([
    ('icss_select', ICSSFeatureSelector(k=2)),
    # Add your own preprocessing/classification steps here
])

# Fit and use the pipeline
pipeline.fit(X, y)
X_transformed = pipeline.transform(X)

Input Format

The feature selector expects input in the following format:

  • X: 3D NumPy array of shape (n_samples, n_features, 2)

    • n_samples: Number of samples in the dataset
    • n_features: Number of interval-valued features
    • 2: Each feature is represented as [lower_bound, upper_bound]
  • y: 1D NumPy array of shape (n_samples,)

    • Target class labels (discrete values for classification)

Example Data Creation

import numpy as np

# Method 1: Random intervals
n_samples, n_features = 100, 5
X = np.random.rand(n_samples, n_features, 2)
X = np.sort(X, axis=2)  # Ensure min <= max

# Method 2: From existing data with uncertainty
lower_bounds = np.random.rand(n_samples, n_features)
upper_bounds = lower_bounds + np.random.rand(n_samples, n_features) * 0.5
X = np.stack([lower_bounds, upper_bounds], axis=2)

# Class labels
y = np.random.randint(0, 3, n_samples)  # 3 classes

API Reference

ICSSFeatureSelector

class ICSSFeatureSelector(BaseEstimator, SelectorMixin)

Parameters

  • k (int, default=10): Number of top features to select based on ICSS scores.

Attributes

  • scores_ (ndarray of shape (n_features,)): ICSS score for each feature after fitting.

Methods

  • fit(X, y): Compute ICSS scores for all features.
  • transform(X): Select features with highest ICSS scores.
  • fit_transform(X, y): Fit and transform in one step.
  • get_support(indices=False): Get a boolean mask or indices of selected features.

Complete Example

Here's a complete example demonstrating the feature selection workflow:

import numpy as np
from icss_feature_selection import ICSSFeatureSelector

# 1. Create dummy interval data
n_samples = 100
X = np.random.rand(n_samples, 4, 2)
X = np.sort(X, axis=2)  # Ensure [min, max] order

# 2. Create labels correlated with features 0 and 2
y = np.zeros(n_samples)
y[(X[:, 0, 0] > 0.1) & (X[:, 0, 1] < 0.2)] = 1
y[(X[:, 2, 0] > 0.7) & (X[:, 2, 1] < 0.9)] = 1

print(f"Original X shape: {X.shape}")  # (100, 4, 2)

# 3. Use the ICSSFeatureSelector to select top 2 features
selector = ICSSFeatureSelector(k=2)

# 4. Fit the selector
selector.fit(X, y)

# 5. Check the scores (higher is better)
print(f"ICSS Scores: {selector.scores_}")

# 6. Transform the data
X_selected = selector.transform(X)
print(f"Transformed X shape: {X_selected.shape}")  # (100, 2, 2)

# 7. Get selected feature indices
support_indices = selector.get_support(indices=True)
print(f"Selected feature indices: {support_indices}")

print("\n✓ Successfully performed ICSS feature selection!")

Algorithm Details

The ICSS algorithm works as follows:

  1. Reference Intervals: For each feature, extract all unique interval values from the dataset.

  2. Similarity Kernel: Define a similarity measure between intervals: $$\text{sim}(I^q, I^r) = \begin{cases} 1 & \text{if } I^q \text{ contains } I^r \ 0 & \text{otherwise} \end{cases}$$

  3. Contingency Table: Build a contingency table using the similarity kernel to count occurrences of each reference interval in each class.

  4. Chi-Square Calculation: Compute the chi-square statistic adapted for interval data: $$I\chi^2 = \sum_{i,j} \frac{(O_{ij} - E_{ij})^2}{E_{ij}}$$ where $O_{ij}$ are observed frequencies and $E_{ij}$ are expected frequencies.

  5. Feature Ranking: Rank features by their ICSS scores and select the top-k.

For more details, please refer to the research paper.

Research Paper

This implementation is based on the following peer-reviewed research:

Interval Chi-Square Score (ICSS): Feature Selection of Interval Valued Data

  • Authors: Guru, D.S.; Vinay Kumar, N.
  • Conference: Intelligent Systems Design and Applications (ISDA 2018)
  • Published: April 14, 2019
  • Publisher: Springer, Cham
  • Book: Advances in Intelligent Systems and Computing, Vol. 941
  • Pages: 579–590
  • Print ISBN: 978-3-030-16659-5
  • Online ISBN: 978-3-030-16660-1
  • DOI: https://doi.org/10.1007/978-3-030-16660-1_67

Links

Citation

If you use this implementation in your research, please cite the original paper:

@incollection{guru2020interval,
  title={Interval Chi-Square Score (ICSS): Feature Selection of Interval Valued Data},
  author={Guru, D.S. and Vinay Kumar, N.},
  booktitle={Intelligent Systems Design and Applications},
  pages={579--590},
  year={2020},
  publisher={Springer, Cham},
  isbn={978-3-030-16660-1},
  doi={10.1007/978-3-030-16660-1_67}
}

Testing

Run the test example to verify the installation:

python tests/test.py

Expected output:

Original X shape: (100, 4, 2)
ICSS Scores: [...]
Transformed X shape: (100, 2, 2)

✓ Successfully created and tested ICSSFeatureSelector!

Use Cases

  • Climate Data: Select important temperature/humidity ranges for weather prediction
  • Medical Data: Identify key biometric ranges that correlate with disease diagnosis
  • Financial Data: Select relevant price ranges for stock trading systems
  • Sensor Data: Choose important measurement ranges from multi-sensor systems
  • Survey Analysis: Identify key response ranges in survey data

Implementation Notes

  • The algorithm uses a similarity kernel where an interval $I^q$ "contains" interval $I^r$ if $q_- \leq r_- \text{ and } q_+ \geq r_+$
  • ICSS scores are computed using the chi-square formula adapted for interval data
  • Features with higher ICSS scores have stronger association with the class label
  • The implementation is compatible with scikit-learn's feature selection interface

License

This project is licensed under the Apache License Version 2.0 - see the LICENSE file for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  • Related Work

For other feature selection techniques and interval data methods, see:


Last Updated: November 2025

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

icss_feature_selection-0.1.4.tar.gz (20.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

icss_feature_selection-0.1.4-py3-none-any.whl (15.9 kB view details)

Uploaded Python 3

File details

Details for the file icss_feature_selection-0.1.4.tar.gz.

File metadata

  • Download URL: icss_feature_selection-0.1.4.tar.gz
  • Upload date:
  • Size: 20.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for icss_feature_selection-0.1.4.tar.gz
Algorithm Hash digest
SHA256 d177b6222ae56a48f51f2d713310d90c6f0c0e135f30e07bd847d387ff8c39b6
MD5 be88766774b414e10fdeb3f5b134de3e
BLAKE2b-256 5dd85b8bff489bfd2624464a6b8fdd6d77cc77a2b573888cd793e60c78fc48f6

See more details on using hashes here.

Provenance

The following attestation bundles were made for icss_feature_selection-0.1.4.tar.gz:

Publisher: publish-to-pypi.yml on vinaykumarngitub/icss-feature-selection

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file icss_feature_selection-0.1.4-py3-none-any.whl.

File metadata

File hashes

Hashes for icss_feature_selection-0.1.4-py3-none-any.whl
Algorithm Hash digest
SHA256 fee336781922425b8e608015fef12458d2e25f0dc4e63c05b60833a19e73e37e
MD5 c8f07439391f34de81fcda9d60f88a55
BLAKE2b-256 c2307e7cb9016aa2f873da45a53a81b96d1a3ad5d4d3e47dd5d90fb08ffe7dbe

See more details on using hashes here.

Provenance

The following attestation bundles were made for icss_feature_selection-0.1.4-py3-none-any.whl:

Publisher: publish-to-pypi.yml on vinaykumarngitub/icss-feature-selection

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page